Frankie Research
The parts of LLM systems that are easy to skip and expensive to get wrong: what the model is
allowed to see, when it should say "I don't know", how an answer proves where it came from, and
how a deploy proves it works before it goes out.
Systems
A permission-scoped knowledge base. A company asks questions over its own documents.
Some documents are for everyone, some only for the owner or for specific people. Access is
enforced in the database query, before anything reaches the model — including the word
statistics used for ranking, because computing them over documents a user cannot see would leak
their vocabulary.
Retrieval is hybrid: vector similarity plus a fulltext index in which every word is stored as its
dictionary lemma and its stem. In an inflected language that is what makes one word match all of
its forms — in Czech pes / psa (dog), den / dne (day), liška / lišce (fox) — while words
that merely look alike stay apart (pes is not pesto); both directions are regression tests. A misspelled or accent-less word in the question is corrected against a
vocabulary built only from the documents the asker may see.
A relevance filter then decides whether there is evidence at all. If there is not, the answer is
"this is not in the knowledge base" and the model is never called. Only the documents the answer
actually cites are shown as sources. Secrets for linked services are encrypted at rest and used
only by a server-side proxy, so a token never reaches the browser.
Governance runtime
The second system sits between AI tools and a company and decides what they may do.
- Rules as data. Rules are YAML files ("playbooks") with triggers, checks and a severity.
A check can be a regular expression, a natural-language-inference test or a small sandboxed
expression. Changing the policy does not need a redeploy.
- Three levels: advise, warn, block. A rule declares the minimum level it needs; a caller can
make it stricter, never silently weaker.
- An anchor against drift. A working session carries an anchor embedding of its topic; every
call reports how far the conversation has drifted and asks to re-anchor when it wanders off.
- A gate before anything goes out. It has a gate for outbound actions that derives the
severity from typed data by rule — never from the caller's own description of what it is doing —
renders the message from a template, and writes an audit record.
- Fail-closed defaults. A rule without triggers never fires and an unknown operator evaluates
to false.
Both systems are private. The code is shared by invitation — request access.
Problems I like
- Making "I don't know" a feature instead of a failure.
- Retrieval in inflected languages, where one word has many surface forms. My example language is
Czech: liška, lišku, lišce — one fox, three grammatical cases.
- Access control that holds even when the model is creative.
- Checks and deploy gates that fail closed: a check that did not run is a failed check.
Writing
Decisions
Short records of decisions — the context, the options I rejected, what I chose and what it
costs: decision records.
How I work
- Write the failing case as a test before the fix.
- Keep the model behind one function, so it can be mocked, measured or replaced.
- Treat "no failure reported" as "not verified".
- Write down why a decision was made and what it costs — the next person, often me, needs it.
Frankie Research · [email protected] · Telegram t.me/frankie_research · request code access