Decision record 0005
0005 — Deploy gates fail closed
Status: accepted · Date: 2026-09
Context
Deploys go through a gate script: build, run checks against the new version, switch traffic only if the checks pass. Two incidents shaped how that gate behaves:
- A smoke check could not reach the service (wrong port in a test config). The check printed an error, returned nothing, and the gate treated "no failures reported" as success. The deploy went out broken.
- A check verified that a page rendered, not that it showed the right data. It passed while the numbers on the page were wrong.
Both are the same mistake: treating the absence of a failure as evidence of success.
Options
- Gate on "no check reported a failure".
- Gate on "every expected check reported a pass", and treat anything else — a check that did not run, timed out, could not connect, or returned something unexpected — as a failure.
Decision
Option 2. The gate is fail-closed:
- Each check returns an explicit exit status; the gate evaluates it rather than printing output and moving on.
- A missing, skipped, or unreachable check is a failure, not a neutral result.
- Checks assert on content (the value the user will see), not only on "the page loaded".
- Every gate includes one deliberately failing case run first, to prove the gate can say no. A gate that has never failed has not been tested.
- The gate is never bypassed "just this once". If it is wrong, the gate gets fixed; working around it and reproducing its side effects by hand (notifications, smoke runs) is how the next silent failure gets in.
- Before overwriting anything on a server: back up the target and diff against what is live. If the server has something the local copy does not, merge — never overwrite a newer version with an older one.
In this repository
The same rules run as scripts/gate.py, the only CI step:
- Deliberately failing cases first —
scripts/forced_fail.pyundoes each fix in a temporary copy (the access filter allows everything, a query loses its access scope, a prefix counts as a full match, the proxy allows every path, …) and requires that fix's tests to fail; then the same tests must pass on the untouched code. A test that cannot fail proves nothing. - Lint.
- Unit + integration tests against a real Postgres, with
KB_REQUIRE_DB=1: a missing database is a failure, not a skipped test, and the summary must be a clean pass.
A step that cannot run (missing tool, timeout) fails the gate. The deploy-specific rules above (backup and diff before overwriting a server) belong to the deployment pipeline of the system this code was extracted from; this repository ships no deployment.
Cost
- More false alarms: a flaky network now blocks a deploy instead of being ignored. That is the point — a deploy you cannot verify is not a deploy you should make.
- Checks take longer to write because they must assert real outcomes.
- Occasionally a correct deploy waits for a fixed gate. Accepted; the alternative was shipping broken builds with a green light.