ENCSRequest code access

Decision record 0005

0005 — Deploy gates fail closed

Status: accepted · Date: 2026-09

Context

Deploys go through a gate script: build, run checks against the new version, switch traffic only if the checks pass. Two incidents shaped how that gate behaves:

  1. A smoke check could not reach the service (wrong port in a test config). The check printed an error, returned nothing, and the gate treated "no failures reported" as success. The deploy went out broken.
  2. A check verified that a page rendered, not that it showed the right data. It passed while the numbers on the page were wrong.

Both are the same mistake: treating the absence of a failure as evidence of success.

Options

  1. Gate on "no check reported a failure".
  2. Gate on "every expected check reported a pass", and treat anything else — a check that did not run, timed out, could not connect, or returned something unexpected — as a failure.

Decision

Option 2. The gate is fail-closed:

  • Each check returns an explicit exit status; the gate evaluates it rather than printing output and moving on.
  • A missing, skipped, or unreachable check is a failure, not a neutral result.
  • Checks assert on content (the value the user will see), not only on "the page loaded".
  • Every gate includes one deliberately failing case run first, to prove the gate can say no. A gate that has never failed has not been tested.
  • The gate is never bypassed "just this once". If it is wrong, the gate gets fixed; working around it and reproducing its side effects by hand (notifications, smoke runs) is how the next silent failure gets in.
  • Before overwriting anything on a server: back up the target and diff against what is live. If the server has something the local copy does not, merge — never overwrite a newer version with an older one.

In this repository

The same rules run as scripts/gate.py, the only CI step:

  1. Deliberately failing cases first — scripts/forced_fail.py undoes each fix in a temporary copy (the access filter allows everything, a query loses its access scope, a prefix counts as a full match, the proxy allows every path, …) and requires that fix's tests to fail; then the same tests must pass on the untouched code. A test that cannot fail proves nothing.
  2. Lint.
  3. Unit + integration tests against a real Postgres, with KB_REQUIRE_DB=1: a missing database is a failure, not a skipped test, and the summary must be a clean pass.

A step that cannot run (missing tool, timeout) fails the gate. The deploy-specific rules above (backup and diff before overwriting a server) belong to the deployment pipeline of the system this code was extracted from; this repository ships no deployment.

Cost

  • More false alarms: a flaky network now blocks a deploy instead of being ignored. That is the point — a deploy you cannot verify is not a deploy you should make.
  • Checks take longer to write because they must assert real outcomes.
  • Occasionally a correct deploy waits for a fixed gate. Accepted; the alternative was shipping broken builds with a green light.