Notebook / Essay
Why vector search alone failed us
October 2026
Most retrieval-augmented systems start the same way: chunk the documents, embed the chunks, embed the question, take the nearest few, hand them to a model. Ours did too. It looked great on demo data — long, well-written documents and well-formed questions. It failed on the first case that looked like real life.
The case
Simplified, with the name changed: a user recorded a short note — two sentences, one of them saying they had sent an invoice to a supplier. In Czech, "to Novák" is Novákovi — the name is inflected. Later they asked the obvious question: "Who is Novák?"
The note did not come back.
Not because the embedding model is bad. It is a good multilingual model. The problem was structural:
- Short loses to long. The cosine distance from a short question to a two-sentence note was worse than its distance to several long documents that were vaguely about suppliers and invoices. Long documents cover more ground, so they are a little bit close to everything.
- The band is narrow. With this model, the distances of relevant and unrelated chunks sat in a narrow band — roughly 0.10 to 0.17. A threshold that admits the note also admits the noise; a threshold that rejects the noise rejects the note.
- The word form differs. Novák and Novákovi are the same person and different tokens. An exact-word index would not help either.
Write it down as a test first
Before changing anything, the case became a test: given these candidate chunks and their real distances, the note must be returned and the long document must not. With vector-only retrieval the test failed — nothing passed the threshold. That test is still in the repository; any future change to the retriever has to keep it green.
What we tried, and why we stopped
Tuning the threshold. There is no threshold. The bands overlap.
A bigger embedding model. It moves the numbers a little and costs latency and memory. Short text against long documents is a property of dense retrieval, not a bug in one model.
What we built — twice
A second retriever: fulltext. Vector search stays; next to it runs a fulltext search, and the
two candidate lists are merged with Reciprocal Rank Fusion: every chunk scores 1/(60 + rank) for
each list it appears in. No score normalisation, no weights to tune, and a chunk that both
retrievers like rises to the top.
First version: crude prefixes. We had rejected a real stemmer because it has to behave
identically in Python (where we filter) and in the database (where we search). So a word became its
first four letters without diacritics: liška, lišku, lišky → lisk, Novák, Novákovi → nova.
Crude on purpose, and it fixed the original case.
Then the fix needed fixing. Prefixes fail in both directions. They miss words whose stem
changes: Czech alternates consonants and drops vowels, so liška / lišce (fox) do not share four
letters, and short words such as pes / psa (dog) or den / dne (day) were too short to index at
all — pesto could be found, pes could not. And unrelated words do share four letters: a first
name in the dative, Pavlovi, and pavlač (a gallery walkway) both become pavl; mechu (moss)
and mechanika both become mech. We wrote both directions down as tests, and the prefix failed
them.
Second version: lemma and stem, computed once. Every word is now stored as two keys: its dictionary lemma (which knows psa is pes) and its Snowball stem (which handles names and words no dictionary has). The worry that sank the stemmer the first time — Python and the database disagreeing — disappeared once the keys are computed in Python and stored in the index: the database no longer has an opinion about words, it only compares keys. The old prefix stays as a weak, half-weight fallback.
Matching more forms keeps the opposite risk alive: merging words that merely look alike, now that short words count too. So next to every "must find" test there is a "must not find" one — pes must not find pesto, Pavlovi must not find pavlač, ruka (hand) must not find rukavice (gloves).
Typos, within your own vocabulary. A misspelled or accent-less word in the question is corrected to the closest word that actually occurs in the documents — but only in the documents the asker may see, so the correction cannot reveal words from anyone else's.
Then demand evidence. Fusion improves recall; it does not tell you whether anything actually answers the question. So a chunk becomes a source only if it has one of two kinds of evidence:
- a close vector — within the threshold and within a small margin of the best candidate; or
- rare question words matched by lemma or stem, which together cover more than half of the question, weighted by rarity — common words do not count, and a word the corpus has never seen weighs the most.
That last rule is what lets the system say "I don't know". Ask "What does the green hummingbird drink in Ostrava?" and an unrelated chunk about green vegetables near Ostrava shares two rare words. But hummingbird is not in the corpus at all, it carries the most weight, coverage stays under half, and the answer is honest: this is not in the knowledge base. The model is not called.
One detail that matters: "rare" is computed only over documents the user is allowed to see. Otherwise the ranking would behave differently because of documents they cannot see — a side channel. (More in Permissions before the model.)
What it cost
- More queries per question: vector, fulltext and a count per question word. Cheap at thousands of chunks; at millions the word counts would need caching.
- A lemma dictionary and a stemmer as dependencies, with their licences checked, and a one-time backfill of the keys for existing documents.
- Thresholds are calibrated to one embedding model. Changing the model means recalibrating against the regression tests — which is exactly why the real failures live there, in both directions.
- A visible "not found" more often than before. That is the point, and it is also honest feedback about what the knowledge base is missing.
The lesson I keep: a demo set made of long, clean documents will tell you vector search works. The first short, messy, inflected record will tell you whether it does.