Decision record 0002
0002 — Vector search alone is not enough: hybrid retrieval with a relevance filter
Status: superseded by 0006 · Date: 2026-10
Superseded by 0006. Hybrid retrieval, RRF fusion and the relevance filter described here are still in place. The word matching is not: 4-character prefix stems produced false positives ("Pavlovi" ~ "pavlač", "pesto" searchable while "pes" is not) and missed alternations ("psa" → "pes"). Words are now matched by lemma + stem keys computed in Python and stored in the database, with the prefix kept only as a half-weight fallback. The record below is kept unchanged as the history of the decision.
Context
Retrieval started as pure vector search (multilingual sentence embeddings, pgvector, top-k by cosine distance). It worked on demos with long, well-written documents. It failed on a real case:
- A short record (two sentences) mentioned a person in an inflected form — in Czech, "Novákovi" ("to Novák").
- The question was short and used the base form: "Who is Novák?"
- The short record's cosine distance was worse than that of several long documents that were only vaguely on topic. With this embedding model the distances of "relevant" and "unrelated" chunks sit in a narrow band (roughly 0.10–0.17), so a threshold that admits the record also admits noise, and one that rejects noise rejects the record.
We wrote that case as a test first. It failed with vector-only retrieval.
Options
- Tune the threshold. Rejected: the bands overlap; there is no threshold that separates them.
- A bigger / different embedding model. Helps at the margin, costs latency and memory, and short-text-vs-long-document is a structural weakness of dense retrieval, not a model bug.
- Language-specific stemming (Snowball/Hunspell for Czech). Better recall, but a dictionary dependency per language and a mismatch risk between the Python side and the database side.
- Hybrid: vector + fulltext with prefix stems, fused by RRF, then a relevance filter.
Decision
Option 4.
- Fulltext: a generated
tsvectorcolumn over unaccented text with thesimpleconfiguration (keeps every word form), queried with prefix termsstem:*. A stem is the first four characters of a word without diacritics — crude, dictionary-free, and identical in Python and SQL, so the two halves agree on what a word is. For an inflected language this matches "liška / lišku / lišky" (one noun, three cases) without knowing any grammar. - Fusion: Reciprocal Rank Fusion (
1/(60 + rank)summed across the two lists). No score normalisation, no weights to tune. - Relevance filter: a chunk becomes a source only with evidence — either a close vector (within a threshold and within a small margin of the best candidate) or rare question words (present in ≤ 2 % of the user's chunks) that together cover more than half of the question. The coverage rule is what makes "I don't know" possible: a question about something the corpus does not contain ("green hummingbird") can share two rare words with an unrelated chunk, but the unknown word keeps coverage below half.
If nothing passes, the answer is "not in the knowledge base" and the model is not called.
Cost
- Two queries plus a per-stem count query per question. Fine at thousands of chunks; would need caching of document frequencies at millions.
- The stemmer has known limits (consonant alternation: "lišce" → a different stem). Accepted because fulltext is the second retriever — the vector side still covers most paraphrases.
- Thresholds are calibrated to one embedding model. Changing the model means recalibrating against the regression set (tests/test_hybrid.py), which is why the real failure case lives there as a test.