skip to content

Writing a Testable Rule

A rule can only be tested if it decides over a document you can construct, and the case that matters is the legitimate change it must not block. Interviewers probe what a green suite hides.

on this pageshow

explore

questions

12

What is a shadow run of a candidate policy rule, and what does it tell you before you enforce it?

level: juniorimportance: must knowfreq 63%

answer

  1. Measure before you block anyone
  2. Replay the rule over real history
  3. Verdicts recorded, nothing enforced
  4. Would-be denials, then manual triage

basics

~20 s

A shadow run evaluates a candidate rule offline against a corpus of real, already-completed changes and records the verdict it would have given each one, without blocking anything. It shows how often the rule fires and on what.

solid answer

~50 s

A shadow run takes a rule that is not yet wired to any gate, feeds it inputs that already exist, and records what it would have decided. For a rule requiring every workload image to descend from an approved base-image lineage, that means pulling the stored metadata for every image the pipelines built over the last ninety days and evaluating the rule once per image. Nothing is blocked, no message reaches a developer, and the engine is nowhere near anyone's build. What comes out is a table of would-be denials with the service that owns each one. That list is the input to triage, not a score: a would-be denial can be perfectly correct. The point is to replace an opinion about the rule with a measurement of what it does to the population it will actually meet.

go deeper

for a junior

Be ready to say what a shadow run produces: a recorded verdict for every real input and a list of the changes the rule would have denied, with nothing actually blocked while it runs.

for a middle

Explain where the inputs come from — stored build metadata rather than invented examples — and why the output still needs to be triaged by hand before any number is quoted from it.

for a senior

Show that you run it to decide ship or no-ship, and that you can name what the corpus does not contain: rare rebuilds, long-tail services, and anything nobody attempted.

for a principal

Own the framing that measurement is what makes a guardrail negotiable with the teams it will affect. You want to arrive at that conversation with a number and a named list, not a conviction.

## The move A shadow run — also called a dry evaluation or a backtest of a rule — is the step between writing a rule and letting it decide anything. You take the candidate rule, feed it real inputs that already exist somewhere, and record the verdict it would have produced for each one. Nothing is enforced, nothing is warned, no developer sees a message. The run is a batch job over stored data, not a deployment. ## What it looks like concretely Take a rule that says every workload image must descend from an approved base-image lineage. The surface it reads is image metadata: the labels the build stamped on the image, the recorded base reference, the layer history. That metadata already exists for every image your pipelines produced. So the corpus is, for example, ninety days of built images — a few thousand records — and the run evaluates the rule once per record. The output is a table with one row per image: the image, the service and team that owns it, the verdict, and the message the rule would have printed. Two aggregates fall out of it immediately: how many images the rule would have denied, and how those denials distribute across services and teams. ## Why review is not enough Reading a rule tells you whether it expresses the intent you had. It cannot tell you what the rule does to the population it will meet, because that population is a fact about your estate, not about your rule. Estates are full of history: a service that still builds from a base image approved two years ago, a build that renames labels, a mirrored registry path that is the same base under another name. Every one of those is a change that would be denied by a rule you would have signed off on in review. The asymmetry matters. A rule that misses a bad change costs you a risk you already had. A rule that denies a good change costs you a person's afternoon, an escalation, and — repeated a few times — the credibility of the whole gate. So the measurement you want before shipping is specifically about the good changes. ## What comes out is a list, not a verdict The most common misreading is to treat the shadow denial count as an error count. It is not. A denial can mean the rule is right and the image genuinely descends from something unapproved, in which case you have found real work for a team to do. It can also mean the image complies in substance and the rule read the wrong thing. Separating the two is manual triage, and the shadow run's job is only to hand you the shortlist. The allowed rows deserve a look too. Everything the rule let through is where its misses live, and no aggregate in the report will point at them; you only find them by sampling. ## What it costs and what it risks Almost nothing. The engine is not in anyone's request path, so there is no latency question and no availability question — an evaluation that crashes on a malformed record costs you a rerun. The real cost is the triage labour on the denial list, which is why a corpus is chosen to be representative rather than exhaustive. ## What it cannot tell you A shadow run measures history. It contains only changes people actually made, in a world where this rule did not exist. It does not contain the change nobody attempted, the emergency rebuild that happens twice a year, or the practice teams will adopt once the rule is real. That is a limit to state out loud when you present the result, not a reason to skip the measurement. ## The decision it feeds At the end you can say something a lead can act on: this rule would have denied N changes over ninety days, of which M were legitimate, concentrated in these teams. That sentence is the difference between arguing about a guardrail and deciding about one — and sometimes the honest conclusion it supports is that the rule should not ship at all.

  • What corpus would you assemble to shadow-run a rule about base images?
    The stored metadata for every image the pipelines produced over a recent window — say ninety days — plus the latest image from every service in the estate, however old that build is. The window alone is dominated by whatever rebuilds nightly; the per-service sample is what drags the long tail into the corpus.
  • How is a shadow run different from having the rule evaluate in production without blocking?
    A shadow run happens before the rule is wired to anything: stored documents are replayed offline, so there is no engine in a request path, nothing to be unavailable, and no message reaching a developer. It is a measurement you run to decide whether the rule should exist, not a way of operating one that already does.
  • The shadow run denied one service's images 400 times. Is that 400 problems?
    Almost certainly one. If a service rebuilds the same image nightly, every rebuild reproduces the same verdict. Deduplicate by service or by distinct image content before quoting any number, or a single misconfigured pipeline will dominate the result and hide everything else.

It is the same move as replaying a new fraud rule over last quarter's card transactions before it is allowed to decline anybody: same rule, same real data, no consequences.

saying these in an interview costs you the question

  • Treats every shadow denial as a confirmed violation
  • Assumes a clean code review means the rule is safe
  • Builds the corpus only from the busiest repositories
  • Says the run proves nothing because nothing was blocked
  • Never looks at what the rule allowed

context

open as a page

Why does a policy rule's test suite need near-miss allow cases, not just deny cases?

level: juniorimportance: must knowfreq 68%

basics

~10 s

Near-miss cases are legitimate inputs sitting just inside the boundary — the changes the rule must let through. Deny cases only prove a rule can refuse something. They never prove it refuses nothing else.

open as a page

What is a policy rule's input contract, and what does it pin down?

level: juniorimportance: must knowfreq 66%

basics

~10 s

The input contract is the written agreement about the document a rule decides over: which step produces it, how it is wrapped, and which fields are guaranteed present. Agree it before writing the rule.

open as a page

A rule blocking secrets as container env vars missed a chart using envFrom — what went wrong?

level: middleimportance: must knowfreq 54%

basics

~20 s

The same fact has several representations in the document. A secret reaches the environment through env[].valueFrom.secretKeyRef and through envFrom[].secretRef, and the rule inspected only the first. The contract must enumerate every field the fact can appear in.

open as a page

A candidate rule shadow-run over 2,400 built images denied 118 of them — how do you compute its false-positive rate?

level: middleimportance: should knowfreq 52%

basics

~20 s

Triage all 118 denials first: a denial is a false positive only if the image actually complied. The rate is confirmed false positives over the images evaluated. A raw denial count is not a rate.

open as a page

Why build policy fixtures from real merged changes instead of hand-writing the input?

level: middleimportance: should knowfreq 52%

basics

~20 s

A hand-written fixture encodes the author's mental picture of the input document — the same picture the rule encodes. If that picture is wrong, both are wrong together and the test still passes. Captured documents carry shapes nobody imagined.

open as a page

Your shadow run over ninety days of real changes shows zero false positives — why is that not enough to enforce the rule?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because the corpus only holds changes people actually made and completed within that window. Rare events like a quarterly base refresh or an incident rebuild fall outside it, and a zero can also mean the rule never fired at all.

open as a page

Why should a policy test assert the returned message, not only that the verdict was deny?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A verdict-only assertion passes when the rule denies for the wrong reason, so a broken condition stays hidden behind a green test. The message is also the whole product for the developer who is blocked, so it deserves an assertion of its own.

open as a page

Why can a green policy test suite still miss what the real enforcer sends?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Because the fixtures were hand-written from documentation while the enforcer sends a defaulted, already-rewritten document. The tests prove the rule works on a shape nobody produces. Validate captured real output against the declared contract in CI.

open as a page

Your policy suite is green but no captured plan ever exercised one rule branch — how do you close that gap?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Enumerate the rule's branches by reading the rule, not the corpus, then hand-write a fixture for each branch reality has not yet produced. Captured plans only contain what teams have already done, so whole conditions can sit untested behind a green suite.

open as a page

Your rule requiring every image to descend from an approved base is correct, but a third of production images are vendor-built and carry no lineage metadata — do you ship it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Not as written. A rule is only shippable if every population it denies has a fix someone can perform, and nobody can add lineage metadata to a vendor's image. Narrow it to your own builds, or shelve it.

open as a page

The renderer producing your policy input is changing shape — how do you version the contract and rules together?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Stamp a schemaVersion on the input document and have each rule declare which versions it can decide over. On an unknown version the rule refuses to decide and the gate blocks, so drift is loud.

open as a page