Decision record 0003
0003 — Citations are the source of truth; no evidence means "I don't know"
Status: accepted · Date: 2026-10
Context
Every answer shows its sources. The first version listed every retrieved chunk. Users read that list as "this is what the answer is based on" — and it wasn't: retrieval returns candidates, the model uses a few. A source list that includes unused chunks is a false claim, and it trains people to stop checking.
The other failure was the opposite: when retrieval found nothing good, the model still produced a fluent, plausible answer from weak context.
Options
- List all retrieved chunks (status quo).
- Ask the model to return structured source IDs.
- The model cites document titles inline; only cited chunks become sources. No evidence → no model call.
Decision
Option 3.
- The system prompt requires citing the document title for every claim, in square brackets:
[Release process]. After the answer is generated, sources are the retrieved chunks whose document title is cited — the exact title inside a bracket marker (case and spacing normalised; a breadcrumb or part suffix such as[Release process › Rollback]is accepted). One source per document (the best-ranked chunk), shown with a readable excerpt (app/sources.py). - If the relevance filter (decision 0002) leaves nothing, the endpoint returns a fixed "not in the knowledge base" message, with no sources and without calling the model.
- If the model is unavailable, the user still gets the evidence chunks, clearly marked as
not-an-answer (
grounded=false), instead of an error.
Why not 2: structured IDs are cleaner on paper, but a model that hallucinates an answer will hallucinate a plausible ID as well. Title citation is human-checkable in the answer text itself, and the mapping back to chunks is deterministic code, not model output.
Amendment (2026-10): the first version treated a title as cited whenever it occurred anywhere in the answer as a substring. Short, common titles ("Release process", "Onboarding", "FAQ") then became sources whenever the answer used those words in ordinary prose. Only bracket markers with an exact title count now; the false-positive cases are tests (tests/test_sources.py).
Cost
- If the model forgets the brackets, the answer has no sources. That is the safe failure (an unsupported claim of support is worse than a missing one), and it is visible.
- Title matching depends on titles being distinctive. Two documents with the same title would both be shown; acceptable at this scale, and a reason to enforce unique titles later.
- "I don't know" is visible failure. It is the intended behaviour, but it makes coverage gaps obvious — which is useful feedback, and occasionally uncomfortable in a demo.