ENCSRequest code access

Decision record 0001

0001 — Permissions are enforced before the model, in SQL

Status: accepted · Date: 2026-08

Context

The knowledge base answers questions over company documents. Some documents are for everyone; some (internal notes, figures, contracts) are for the owner or for specific people. An LLM sits at the end of the pipeline and writes the answer.

The question is where access control lives.

Options

  1. Tell the model. Retrieve broadly, pass the user's role, and instruct the model not to reveal restricted content.
  2. Filter after retrieval, in Python. Retrieve top-k, then drop chunks the user may not see.
  3. Filter inside the retrieval query. Every SQL statement that touches chunks joins the document and applies the user's access clause.

Decision

Option 3. accessible_document_clause(user) (app/access.py) returns an SQL condition — shared with the whole company / owned by me / ACL by user / ACL by role — and it is applied to every query that reads chunks: the vector search, the word search, the word-frequency counts used by the relevance filter, and the vocabulary used to correct typos.

There is no "no filter" value, not even for the company owner. Visibility has one source of truth (sharing + owner_id + ACL); all_company is derived from sharing and a CHECK constraint keeps them equal. Someone else's personal item stays theirs — a promise that is only true if nobody, the owner included, is exempt. (Earlier versions let the owner bypass the filter and decided visibility from all_company, while sharing existed but was not read: two models that could drift apart.)

Why not 1: an instruction is not a control. A prompt-injected document, a paraphrased question, or a model update can all turn "please don't" into a leak, and you cannot test your way to certainty. Worse, models told to "be careful" over-correct and refuse the people who are entitled — we saw the owner get refused non-deterministically when the prompt mentioned confidentiality.

Why not 2: filtering after top-k silently shrinks the result. If 6 of the top 8 chunks are restricted, the user gets 2 weak chunks and a worse answer, with no hint why. Filtering in the query ranks only what the user may see.

The word-frequency counts were the non-obvious part: if "rare word" statistics were computed over the whole corpus, the relevance filter would behave differently depending on documents the user cannot see — a side channel. So they are scoped too, and so is the typo-correction vocabulary: correcting "astrolabbe" to a word that exists only in someone else's document would reveal that the word exists.

Cost

  • Every retrieval query carries the access join; with an index on document_acl(document_id) it is cheap at this scale, but it is not free.
  • The model prompt states the opposite of option 1: everything you were given, the user may see — answer fully. That is only safe because the filter is upstream. The two decisions are coupled; changing one without the other is a bug.
  • Access logic lives in one function. Any new way of reading chunks must go through it.
  • The claim is tested where it matters: against a real Postgres + pgvector, through the real queries (tests/integration/test_acl_pg.py) — vector search, word search, df and corpus size, vocabulary, the answer pipeline and the document endpoints, for five members with different rights. A forced-FAIL run (scripts/forced_fail.py) makes the filter allow everything, and separately drops the scope from single queries, and requires those tests to fail.