ENCSRequest code access

Decision record 0006

0006 — Lemma + stem keys replace prefix stems

Status: accepted · Date: 2026-10 · Supersedes: 0002 (the word-matching part; hybrid retrieval, RRF and the relevance filter stay)

Context

Decision 0002 added a fulltext retriever next to the vector one and matched words by their first four characters without diacritics. It rejected real stemming because of "a dictionary dependency per language and a mismatch risk between the Python side and the database side".

In use, the prefix turned out to be wrong in both directions, and the failures are exactly the ones a reviewer tries first. Czech is the example language:

  • False positives. Unrelated words share a prefix: "Pavlovi" (a first name, dative) and "pavlač" (a gallery walkway) both become pavl, so a question about a person pulled in the courtyard repair notes; "mechu" (moss) and "mechanika" both become mech. Short words had no prefix at all, so "pes" (a dog) could not be searched, while "pesto" could.
  • False negatives. Inflection is not only suffixes. "pes / psa / psovi" (dog: nominative, genitive, dative) share no 4-character prefix; "liška / lišce" (fox) differ by a consonant alternation; "den / dne" (day) lose a vowel.
  • The mismatch worry was about computing keys in two places (Python for the relevance filter, a database dictionary for the index). It disappears if the keys are computed once, in Python, and only stored in the database.

Options

  1. Keep the prefix, tune its length. Longer prefixes lose recall on short words, shorter ones add collisions. Does not address alternations at all.
  2. Postgres text-search dictionaries (ispell / snowball configurations). Keys computed in SQL, the relevance filter in Python would need the same dictionary — the mismatch 0002 feared.
  3. Lemma + stem computed in Python, stored as tsvector keys. One function produces the keys for the stored chunks (at ingest) and for the question (at query time).
  4. Morphological analysers with neural or large statistical models (e.g. full taggers trained on annotated corpora). Best quality, but the models for Czech are licensed for non-commercial use only, which rules them out for a product.

Decision

Option 3.

  • Every content word gets l0<lemma> (simplemma, dictionary-based: psa → pes, lišce → liška, dne → den) and s0<stem> (Snowball, rule-based: covers names and words the dictionary does not know, Pavlovi → pavl). Keys are lower-case, accent-free, alphanumeric; the tag keeps a key one token. They are written to chunks.lex (a simple tsvector + GIN index) at ingest, and the question is turned into the same keys. Python and SQL cannot disagree, because there is only one implementation (app/textmatch.py).
  • A word matches strongly when a lemma or stem key is shared. The old accent-free prefix (chunks.tsv, pref:*) stays only as a weak fallback: half weight in the RRF fusion and half a match in the relevance filter, which requires strong matches of the rare question words. "Pavlovi" ~ "pavlač" is now at most a weak match and never evidence on its own.
  • Relevance weights words by idf (rarity in the asker's corpus); common words do not count toward coverage; an unknown word weighs the most, but finitely.
  • Typo correction. A word the asker's corpus does not contain at all would carry the highest weight and sink an otherwise good match — the same as a typo or a missing diacritic. Before ranking, such a word is corrected to the nearest word of the asker's own vocabulary (stem within edit distance 1 for stems of at least 5 characters, or a long common prefix with an accent-free form). Short stems are never corrected ("lesem" must not become "lesnatý"). The vocabulary is built only from chunks the asker may see (decision 0001) and cached per asker only while that set of chunks is unchanged.
  • Index guard. Lemma/stem statistics are used only when every chunk has keys; a half-filled index (interrupted migration, chunks written by an older version) would make every word look rare. Until then retrieval stays on the prefix path. Migration 0004 fills existing chunks; python -m app.lex_backfill --apply completes an interrupted fill.
  • Negative tests are first-class. The regression set has a "must match" list (inflection, alternation, names) and a "must not match" list (Pavlovi/pavlač, pes/pesto, mech/mechanika, ruka/rukavice), plus "unrelated words give zero results" end to end on Postgres (tests/test_textmatch.py, tests/integration/test_search_pg.py).

Cost

  • Dependencies and licences. simplemma (code MIT; lemma data under ODbL, CC BY-SA and CC BY) and snowballstemmer (BSD-3-Clause). Attribution is in NOTICE. Non-commercially licensed models were rejected on purpose.
  • Configured languages. Keys are generated for each language in SEARCH_LANGUAGES (default en,cs). More languages mean more keys and a small chance of a cross-language collision (the Czech lemma of English "ran" is a Czech noun); a deployment should list only the languages of its corpus. Changing the list means recomputing the keys (lex_backfill --all).
  • Stemmers over-stem. The English Snowball stemmer maps "policy" and "police" to one stem. Accepted: a strong match still has to be a rare word covering most of the question, and the vector side ranks alongside.
  • Write-time work. Every chunk is lemmatized at ingest, and existing chunks once in the migration. Cheap compared to computing the embedding.
  • One more column (chunks.lex) next to chunks.tsv, which is kept for the fallback and the typo vocabulary.