ENCSRequest code access
EST. 2025

Lab notebook — retrieval, permissions, governance

Research on what an AI may see, when it must say “I don't know”, and how an answer proves its source.

Two working systems and the notes behind them. Every decision is written down with the options rejected and what it costs.

FIG. 1 — FIG. 2

Every answer passes the same gates.

Access is decided in SQL before anything reaches the model. Words are matched by lemma and stem. Without evidence there is no model call — the answer says so.

QUESTION ACCESS FILTERSQL · BEFORE MODEL HYBRID RETRIEVALVECTOR + LEMMA/STEMRRF · k=60 EVIDENCE GATERARE WORDS · IDF MODEL CITEDANSWER “NOT IN THEKNOWLEDGE BASE” NO EVIDENCE · NO CALL VOCABULARY, DF: SCOPED
FIG. 1Answer pipeline — knowledge base →
$ curl -s /qa/ask -d '{"question": "Kdo je Novák?"}'
{ "grounded": true,
  "answer": "… [Invoice note]",
  "sources": [{ "document_title": "Invoice note",
                "excerpt": "Poslal jsem fakturu Novákovi, …" }] }

$ curl -s /qa/ask -d '{"question": "What does the green hummingbird drink?"}'
{ "grounded": false,
  "answer": "This is not in the knowledge base. …",
  "sources": [] }            # no model call
Must find · pes / psa · den / dne · liška / lišce · Novák / Novákovi
Must not find · pes / pesto · Pavlovi / pavlač · mech / mechanika · ruka / rukavice
Forced fail · every fix undone once — its tests must fail

Systems

Two working systems.

System 01

Knowledge base →

Permission-scoped answers over a company's own documents.

  • Access filtered in SQL before the model — vectors, words, statistics, typo vocabulary
  • Hybrid retrieval: vector + lemma / stem keys, fused by RRF
  • No evidence → “not in the knowledge base”, no model call
  • Sources are only what the answer cites, in [brackets]
  • Linked-service tokens encrypted at rest, used only server-side
Decision records 0001–0006 →

System 02

Governance runtime →

Rules that sit between AI tools and a company.

  • Rules as YAML playbooks: triggers, checks, severity
  • Checks: regular expressions, NLI, small sandboxed expressions
  • Advise / warn / block — stricter allowed, never silently weaker
  • A drift anchor per session, with a re-anchor prompt
  • A gate for outbound actions: severity from typed data, an audit record
Request code access →

Problems

What the work is about.

How I work

Habits, not slogans.

Writing & decisions

The notebook.

Frankie Research

The parts of LLM systems that are easy to skip and expensive to get wrong: what the model is allowed to see, when it should say "I don't know", how an answer proves where it came from, and how a deploy proves it works before it goes out.

Systems

A permission-scoped knowledge base. A company asks questions over its own documents. Some documents are for everyone, some only for the owner or for specific people. Access is enforced in the database query, before anything reaches the model — including the word statistics used for ranking, because computing them over documents a user cannot see would leak their vocabulary.

Retrieval is hybrid: vector similarity plus a fulltext index in which every word is stored as its dictionary lemma and its stem. In an inflected language that is what makes one word match all of its forms — in Czech pes / psa (dog), den / dne (day), liška / lišce (fox) — while words that merely look alike stay apart (pes is not pesto); both directions are regression tests. A misspelled or accent-less word in the question is corrected against a vocabulary built only from the documents the asker may see.

A relevance filter then decides whether there is evidence at all. If there is not, the answer is "this is not in the knowledge base" and the model is never called. Only the documents the answer actually cites are shown as sources. Secrets for linked services are encrypted at rest and used only by a server-side proxy, so a token never reaches the browser.

Governance runtime

The second system sits between AI tools and a company and decides what they may do.

  • Rules as data. Rules are YAML files ("playbooks") with triggers, checks and a severity. A check can be a regular expression, a natural-language-inference test or a small sandboxed expression. Changing the policy does not need a redeploy.
  • Three levels: advise, warn, block. A rule declares the minimum level it needs; a caller can make it stricter, never silently weaker.
  • An anchor against drift. A working session carries an anchor embedding of its topic; every call reports how far the conversation has drifted and asks to re-anchor when it wanders off.
  • A gate before anything goes out. It has a gate for outbound actions that derives the severity from typed data by rule — never from the caller's own description of what it is doing — renders the message from a template, and writes an audit record.
  • Fail-closed defaults. A rule without triggers never fires and an unknown operator evaluates to false.

Both systems are private. The code is shared by invitation — request access.

Problems I like

  • Making "I don't know" a feature instead of a failure.
  • Retrieval in inflected languages, where one word has many surface forms. My example language is Czech: liška, lišku, lišce — one fox, three grammatical cases.
  • Access control that holds even when the model is creative.
  • Checks and deploy gates that fail closed: a check that did not run is a failed check.

Writing

Decisions

Short records of decisions — the context, the options I rejected, what I chose and what it costs: decision records.

How I work

  • Write the failing case as a test before the fix.
  • Keep the model behind one function, so it can be mocked, measured or replaced.
  • Treat "no failure reported" as "not verified".
  • Write down why a decision was made and what it costs — the next person, often me, needs it.

Contact

Frankie Research · [email protected] · Telegram t.me/frankie_research · request code access