Decision record 0004
0004 — The product does not run on a personal AI subscription
Status: accepted · Date: 2026-09
Context
During development, the fastest way to get a capable model into a prototype is a personal chat subscription that the developer already pays for — some tools even let you route requests through it. It is tempting to ship the first version that way: no separate billing, no keys to manage, and a flat monthly price.
Options
- Route product traffic through the developer's personal subscription.
- Use the provider's API under commercial terms, with keys per environment.
- Self-host an open model.
Decision
Option 2 for answer generation; embeddings are already local (no data leaves the server for indexing).
Why not 1: - Terms. Consumer subscriptions are licensed to an individual for their own use. Serving other people's requests through them is outside those terms; a product built on that can be switched off without notice, and the customer would be the one who notices. - Data. Customer documents would flow through an account set up for personal use, with whatever retention and training settings that account has. That is not a position you can defend to a customer's security review. - Operations. No per-request cost, no rate limits you control, no separation between development and production traffic.
Why not 3 (for now): good option for strict data-residency customers, but it moves the hard problem to GPU operations; the architecture keeps the model behind one function (app/llm.py) so this stays a swap, not a rewrite.
Cost
- Real per-call cost, visible from day one. That is a feature: it forces caching and the "don't call the model without evidence" rule (decision 0003).
- Tests never call the model: retrieval, relevance, citations and permissions are pure functions tested without it, so CI makes no paid calls and stays deterministic. Development against the live model uses its own key, separate from production.
- Key management per environment (local, staging, production), never committed.