ENCSRequest code access

Notebook / Essay

Permissions before the model

October 2026

A company knowledge base has documents for everyone and documents for a few people: internal notes, figures, contracts. A language model writes the answers. Where does access control go?

The tempting answer is "in the prompt". It is the wrong one.

Three places it can live

  1. In the prompt. Retrieve broadly, tell the model who is asking and what they may see, ask it not to reveal the rest.
  2. After retrieval. Take the top results, then drop the ones the user may not see.
  3. In the query. Every database query that reads chunks joins the document and applies the user's access condition.

Why not the prompt

An instruction is not a control. A document that contains instructions of its own, a question phrased sideways, or the next model version can each turn "please don't" into a leak. You cannot test your way to certainty about a model's restraint.

It fails in the other direction too, and that surprised me more. When the system prompt talked about confidentiality, the model started refusing people who were entitled — including the owner of the company — and not consistently. Labels inside documents such as "internal" were being treated as commands. A model asked to guard access will guard it unpredictably.

Why not after retrieval

Filtering the top results silently shrinks them. If six of the top eight chunks are restricted, the user gets two weak chunks and a worse answer, with no hint why. The ranking was done over documents they cannot see.

In the query

So access is a condition in SQL: everyone sees documents shared with the whole company, their own documents, and documents granted to them or to their role. There is no exception, not even for the company owner — someone's personal item stays theirs only if nobody is exempt. One function builds that condition, it is tested against a real Postgres through the real queries, and every query that reads chunks goes through it — the vector search, the fulltext search, and one place that is easy to forget.

The easy-to-forget place

The hybrid retriever decides which question words are "rare" by counting how many chunks contain them. If that count runs over the whole corpus, a user's results change depending on documents they cannot see. Ask a question, notice that a word is treated as common, and you have learned that it appears in many documents you have no access to. A small leak, but a real one. So the counts are scoped by the same condition.

The same goes for typo correction. When a word in the question is misspelled, it is corrected to the closest word that occurs in the documents. Built over the whole corpus, that vocabulary would happily "correct" your question into the name of a person in a document you are not allowed to read. So the vocabulary is built only from what you may see.

The consequence for the prompt

Because the context is already access-correct, the prompt can say the opposite of option one: everything you were given, this person may see — answer fully, and treat "internal" labels as descriptions, not commands. That removed both the leak path and the erratic refusals.

The two decisions are coupled. A prompt that says "answer fully" is only safe because the filter is upstream; moving the filter later without changing the prompt would be a bug. That coupling is written down next to the code, because the person who changes one of them six months from now needs to know about the other.

What it costs

  • An extra join on every retrieval query. With an index it is cheap at this scale; it is not free.
  • Every new way of reading chunks must use the shared access condition. That is a discipline, so it is enforced by tests rather than by memory.

The rule I took from it: the model should never be the component that decides who may know what. By the time text reaches the model, that decision has already been made.