skip to content

The Life of a Rule

A rule is software: it needs a test suite, an owner, a version consumers pin, and a decision record when it blocks the wrong change. Interviewers probe it because policy programmes die of neglect.

on this pageshow

explore

questions

page 1 of 2

What is a shadow run of a candidate policy rule, and what does it tell you before you enforce it?

level: juniorimportance: must knowfreq 63%

answer

  1. Measure before you block anyone
  2. Replay the rule over real history
  3. Verdicts recorded, nothing enforced
  4. Would-be denials, then manual triage

basics

~20 s

A shadow run evaluates a candidate rule offline against a corpus of real, already-completed changes and records the verdict it would have given each one, without blocking anything. It shows how often the rule fires and on what.

solid answer

~50 s

A shadow run takes a rule that is not yet wired to any gate, feeds it inputs that already exist, and records what it would have decided. For a rule requiring every workload image to descend from an approved base-image lineage, that means pulling the stored metadata for every image the pipelines built over the last ninety days and evaluating the rule once per image. Nothing is blocked, no message reaches a developer, and the engine is nowhere near anyone's build. What comes out is a table of would-be denials with the service that owns each one. That list is the input to triage, not a score: a would-be denial can be perfectly correct. The point is to replace an opinion about the rule with a measurement of what it does to the population it will actually meet.

go deeper

for a junior

Be ready to say what a shadow run produces: a recorded verdict for every real input and a list of the changes the rule would have denied, with nothing actually blocked while it runs.

for a middle

Explain where the inputs come from — stored build metadata rather than invented examples — and why the output still needs to be triaged by hand before any number is quoted from it.

for a senior

Show that you run it to decide ship or no-ship, and that you can name what the corpus does not contain: rare rebuilds, long-tail services, and anything nobody attempted.

for a principal

Own the framing that measurement is what makes a guardrail negotiable with the teams it will affect. You want to arrive at that conversation with a number and a named list, not a conviction.

## The move A shadow run — also called a dry evaluation or a backtest of a rule — is the step between writing a rule and letting it decide anything. You take the candidate rule, feed it real inputs that already exist somewhere, and record the verdict it would have produced for each one. Nothing is enforced, nothing is warned, no developer sees a message. The run is a batch job over stored data, not a deployment. ## What it looks like concretely Take a rule that says every workload image must descend from an approved base-image lineage. The surface it reads is image metadata: the labels the build stamped on the image, the recorded base reference, the layer history. That metadata already exists for every image your pipelines produced. So the corpus is, for example, ninety days of built images — a few thousand records — and the run evaluates the rule once per record. The output is a table with one row per image: the image, the service and team that owns it, the verdict, and the message the rule would have printed. Two aggregates fall out of it immediately: how many images the rule would have denied, and how those denials distribute across services and teams. ## Why review is not enough Reading a rule tells you whether it expresses the intent you had. It cannot tell you what the rule does to the population it will meet, because that population is a fact about your estate, not about your rule. Estates are full of history: a service that still builds from a base image approved two years ago, a build that renames labels, a mirrored registry path that is the same base under another name. Every one of those is a change that would be denied by a rule you would have signed off on in review. The asymmetry matters. A rule that misses a bad change costs you a risk you already had. A rule that denies a good change costs you a person's afternoon, an escalation, and — repeated a few times — the credibility of the whole gate. So the measurement you want before shipping is specifically about the good changes. ## What comes out is a list, not a verdict The most common misreading is to treat the shadow denial count as an error count. It is not. A denial can mean the rule is right and the image genuinely descends from something unapproved, in which case you have found real work for a team to do. It can also mean the image complies in substance and the rule read the wrong thing. Separating the two is manual triage, and the shadow run's job is only to hand you the shortlist. The allowed rows deserve a look too. Everything the rule let through is where its misses live, and no aggregate in the report will point at them; you only find them by sampling. ## What it costs and what it risks Almost nothing. The engine is not in anyone's request path, so there is no latency question and no availability question — an evaluation that crashes on a malformed record costs you a rerun. The real cost is the triage labour on the denial list, which is why a corpus is chosen to be representative rather than exhaustive. ## What it cannot tell you A shadow run measures history. It contains only changes people actually made, in a world where this rule did not exist. It does not contain the change nobody attempted, the emergency rebuild that happens twice a year, or the practice teams will adopt once the rule is real. That is a limit to state out loud when you present the result, not a reason to skip the measurement. ## The decision it feeds At the end you can say something a lead can act on: this rule would have denied N changes over ninety days, of which M were legitimate, concentrated in these teams. That sentence is the difference between arguing about a guardrail and deciding about one — and sometimes the honest conclusion it supports is that the rule should not ship at all.

  • What corpus would you assemble to shadow-run a rule about base images?
    The stored metadata for every image the pipelines produced over a recent window — say ninety days — plus the latest image from every service in the estate, however old that build is. The window alone is dominated by whatever rebuilds nightly; the per-service sample is what drags the long tail into the corpus.
  • How is a shadow run different from having the rule evaluate in production without blocking?
    A shadow run happens before the rule is wired to anything: stored documents are replayed offline, so there is no engine in a request path, nothing to be unavailable, and no message reaching a developer. It is a measurement you run to decide whether the rule should exist, not a way of operating one that already does.
  • The shadow run denied one service's images 400 times. Is that 400 problems?
    Almost certainly one. If a service rebuilds the same image nightly, every rebuild reproduces the same verdict. Deduplicate by service or by distinct image content before quoting any number, or a single misconfigured pipeline will dominate the result and hide everything else.

It is the same move as replaying a new fraud rule over last quarter's card transactions before it is allowed to decline anybody: same rule, same real data, no consequences.

saying these in an interview costs you the question

  • Treats every shadow denial as a confirmed violation
  • Assumes a clean code review means the rule is safe
  • Builds the corpus only from the busiest repositories
  • Says the run proves nothing because nothing was blocked
  • Never looks at what the rule allowed

context

open as a page

Why does a policy rule's test suite need near-miss allow cases, not just deny cases?

level: juniorimportance: must knowfreq 68%

basics

~10 s

Near-miss cases are legitimate inputs sitting just inside the boundary — the changes the rule must let through. Deny cases only prove a rule can refuse something. They never prove it refuses nothing else.

open as a page

What is a policy rule's input contract, and what does it pin down?

level: juniorimportance: must knowfreq 66%

basics

~10 s

The input contract is the written agreement about the document a rule decides over: which step produces it, how it is wrapped, and which fields are guaranteed present. Agree it before writing the rule.

open as a page

Your nightly encryption-at-rest sweep reports zero violations — what must you know before that means anything?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Which ruleset produced the number and what it covered. Zero violations only means the rules that actually loaded found nothing in the resources that were actually enumerated. An empty or replaced ruleset reports exactly the same green.

open as a page

Why is a policy rule repository reviewed, tested and released like application code?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A rule is production code: one bad rule blocks every team's builds at once. Review, tests and CI catch it before it reaches a gate, and give each change an author, a reviewer and a history.

open as a page

Why must a policy rule carry a stable id and an owner, not just a title?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A rule's id is the handle everything outside the rule keys on: waivers, suppressions, control maps, dashboards. Titles get reworded, so they cannot be that handle. The owner tells a blocked engineer who to ask.

open as a page

Why does a synchronous policy gate need a stated p99 latency budget and a hard timeout?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A synchronous gate runs inside someone else's request, so its time is added to theirs. The p99 target bounds the tail that callers actually feel, and the hard timeout caps the worst case so the caller's own deadline stays predictable.

open as a page

Your CI gate denied a build on policy. What is your first step to reproduce that denial?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Capture the exact input document the engine evaluated and the version of the rules it used, then replay that pair on your own machine. Reproduce the decision before you start editing the job definition.

open as a page

A policy rule has not denied a single change in twelve months - is it safe to delete?

level: juniorimportance: must knowfreq 56%

basics

~20 s

Not yet. Zero denials is ambiguous: it can mean nothing in scope violated the rule, or that the rule never matched anything at all. Compare how often it was evaluated with how often it denied before deciding.

open as a page

A rule blocking secrets as container env vars missed a chart using envFrom — what went wrong?

level: middleimportance: must knowfreq 54%

basics

~20 s

The same fact has several representations in the document. A secret reaches the environment through env[].valueFrom.secretKeyRef and through envFrom[].secretRef, and the rule inspected only the first. The contract must enumerate every field the fact can appear in.

open as a page

Which platform changes can make a working policy rule stop enforcing with no error at all?

level: middleimportance: must knowfreq 66%

basics

~20 s

Anything that moves the shape the rule reads: a field renamed or relocated between API versions, a kind served under a new group-version, a field that gains a default so it is never absent, or an engine change to how the input document is built.

open as a page

A policy rule you own blocks another team's deploy at 5pm: who answers, and what must the failure say?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Separate an engine failure from a working rule. The platform on-call owns the engine; the named rule owner owns the decision. The failure output must name the rule, the resource and property, the owning team and how to propose a change — otherwise every block routes to the platform team.

open as a page

Your CI check at rule library v2.3 blocks a manifest that a developer's pinned v1.9 pre-commit hook passed. How do you respond?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Confirm it is version drift rather than a false positive by running both versions on the same manifest and reading the changelog between them. Then unblock by fixing the manifest, and close the window by moving the lagging hooks forward. The gate stays current.

open as a page

One resource-limits rule runs in a pre-commit hook, CI and admission — why do the three copies drift?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Each decision point installs its own copy of the rule library and updates on its own schedule. Publishing a new version does not change what is already installed, so the three run different rule versions until each one is upgraded.

open as a page

Why keep a policy test fixture that your rule is expected to deny, and what does its sudden pass mean?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A rule that stops matching anything denies nothing, and that silence looks exactly like success. A fixture the rule must deny turns the silence into a failing test: if it suddenly passes, enforcement has quietly stopped.

open as a page

A candidate rule shadow-run over 2,400 built images denied 118 of them — how do you compute its false-positive rate?

level: middleimportance: should knowfreq 52%

basics

~20 s

Triage all 118 denials first: a denial is a false positive only if the image actually complied. The rate is confirmed false positives over the images evaluated. A raw denial count is not a rate.

open as a page

Why build policy fixtures from real merged changes instead of hand-writing the input?

level: middleimportance: should knowfreq 52%

basics

~20 s

A hand-written fixture encodes the author's mental picture of the input document — the same picture the rule encodes. If that picture is wrong, both are wrong together and the test still passes. Captured documents carry shapes nobody imagined.

open as a page

How does a policy enforcer establish that it loaded the intended ruleset and not a substituted one?

level: middleimportance: should knowfreq 46%

basics

~20 s

By computing a digest over the whole rule set it loaded and comparing it against an expected value obtained through a different path than the ruleset itself. A mismatch must fail the run loudly, and the digest should be stamped on every result.

open as a page

What must a policy rule's test suite assert beyond denying the obviously bad input?

level: middleimportance: should knowfreq 62%

basics

~20 s

It must assert what the rule allows, not only what it denies — otherwise a rule that blocks everything passes. It also needs the awkward cases: the property missing entirely, and a value that is only resolved at deploy time.

open as a page

When should a family of near-copy policy rules become one parameterised rule?

level: middleimportance: should knowfreq 47%

basics

~20 s

Collapse near-copies when the logic is identical and only a data value differs - the list of allowed licences, say. Keep them separate when the denial message or the remediation genuinely differ, because that is different logic wearing the same shape.

open as a page

In a shared policy rule library, what makes a change a major SemVer bump rather than a minor one?

level: middleimportance: should knowfreq 42%

basics

~20 s

Anything that turns a previously passing input into a failure: a new blocking rule, a tightened condition, or a stricter default. Changes that cannot newly fail anything — a fix that only loosens a check, a rule shipped switched off — are minor or patch.

open as a page

A gate evaluates 40 rules against an inventory of 5,000 listeners — where does the time go?

level: middleimportance: should knowfreq 51%

basics

~20 s

Mostly in two terms that multiply: marshalling the whole inventory once per call, then every rule walking every listener. Cost tracks rule count times input size, so one extra rule is paid on every call and on every item.

open as a page

Your local replay allows a CI job the gate denied. What could differ between the two evaluations?

level: middleimportance: should knowfreq 47%

basics

~20 s

One of the decision's inputs differs. Either you replayed a rebuilt document rather than the recorded one, a different rule revision, or reference data that has since changed. Chase the difference; do not call the gate flaky.

open as a page

Your shadow run over ninety days of real changes shows zero false positives — why is that not enough to enforce the rule?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because the corpus only holds changes people actually made and completed within that window. Rare events like a quarterly base refresh or an incident rebuild fall outside it, and a zero can also mean the rule never fired at all.

open as a page

Why should a policy test assert the returned message, not only that the verdict was deny?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A verdict-only assertion passes when the rule denies for the wrong reason, so a broken condition stays hidden behind a green test. The message is also the whole product for the developer who is blocked, so it deserves an assertion of its own.

open as a page

Why can a green policy test suite still miss what the real enforcer sends?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Because the fixtures were hand-written from documentation while the enforcer sends a defaulted, already-rewritten document. The tests prove the rule works on a shape nobody produces. Validate captured real output against the declared contract in CI.

open as a page

Your whole policy ruleset was swapped for one that allows everything and sweeps stay green — how do you detect it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Not from the results — they look perfect. Detect it out of band: alert on every publish to the rule distribution point, reconcile the enforcer's reported ruleset digest against what was actually published, and keep a deliberately non-compliant canary whose finding must appear in every sweep.

open as a page

You renamed a policy rule's id last sprint and nothing failed - what silently broke?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Everything keyed on the old id stopped matching, silently. Waivers now exempt nothing, suppressions in other repositories are dead text, control-map rows name a check nobody emits, and the rule's violation trend fell to zero - which reads like success.

open as a page

A policy gate caches allow decisions keyed on listener name and port. Why did a listener that dropped to TLS 1.0 still get an allow?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The cache key covered the object's identity, not the content the rule reads. Nothing in the name or port changed when the TLS minimum did, so the gate returned an old allow without evaluating anything.

open as a page

A policy gate denied every entry of a CI matrix job. How do you tell a rule defect from a real violation?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Diff the documents the engine actually evaluated against the shape the rule author assumed. If expansion produced entries where the field the rule reads is absent rather than wrong, the rule met an input it was never written for.

open as a page

showing 1–30 of 44