Decision records
Decision records
Short records of decisions that shaped this code: the context, the options considered, what was chosen, and what it costs. Written so a reviewer can see why, not only what. A decision that was later replaced stays here, marked superseded — the history is part of the record.
ADR 0001Permissions are enforced before the model, in SQLAnswers are written by a model over documents with different audiences; the question is where access control lives. Decision: in SQL — one visibility rule applied to every query that reads chunks (vectors, words, word statistics, typo vocabulary), with no exception even for the company owner, proven against a real Postgres with a forced-fail run. Cost: an access join on every retrieval query, and every new way of reading chunks has to use the same function.ADR 0002Vector search alone is not enough: hybrid retrieval with a relevance filterPure vector search missed a short record asked about in another word form — relevant and unrelated chunks sat in the same narrow distance band. Decision: hybrid retrieval, vector plus fulltext fused by Reciprocal Rank Fusion, then a relevance filter that demands evidence. Cost: more queries per question and thresholds tied to one embedding model; the word matching was later superseded by 0006.ADR 0003Citations are the source of truth; no evidence means "I don't know"Listing every retrieved chunk as a source overstated the evidence, and weak context still produced fluent answers. Decision: the model cites document titles in [brackets]; only exactly cited documents become sources, one per document, and with no evidence the model is not called. Cost: a forgotten bracket means no sources — the safe failure — titles must be distinctive, and “I don't know” is visible.ADR 0004The product does not run on a personal AI subscriptionRouting product traffic through a personal chat subscription is fast, but outside its terms, puts customer data in a personal account and gives no operational control. Decision: the provider's API under commercial terms with keys per environment; embeddings stay local and the model sits behind one function, so self-hosting stays a swap. Cost: a real price per call from day one and key management per environment.ADR 0005Deploy gates fail closedA smoke check that could not reach the service, and a check that only verified that a page rendered, both let broken deploys through as “no failure reported”. Decision: gates fail closed — every expected check must pass, a check that did not run has failed, checks assert on content, and a deliberately failing case runs first; in this code scripts/gate.py does that in CI. Cost: more false alarms and slower checks, and a correct deploy sometimes waits for a fixed gate.ADR 0006Lemma + stem keys replace prefix stemsFour-letter prefixes failed both ways: false matches (Pavlovi ~ pavlač, pesto searchable while pes was not) and missed alternations (psa → pes, lišce, dne). Decision: lemma + stem keys computed once in Python and stored in the index, the prefix only as a half-weight fallback, idf-weighted relevance, typo correction against the asker's own vocabulary, negative tests as first-class. Cost: dictionary and stemmer dependencies with checked licences, keys per configured language, and lemmatisation at ingest.