What is a shadow run of a candidate policy rule, and what does it tell you before you enforce it?
answer
- Measure before you block anyone
- Replay the rule over real history
- Verdicts recorded, nothing enforced
- Would-be denials, then manual triage
basics
~20 sA shadow run evaluates a candidate rule offline against a corpus of real, already-completed changes and records the verdict it would have given each one, without blocking anything. It shows how often the rule fires and on what.
solid answer
~50 sA shadow run takes a rule that is not yet wired to any gate, feeds it inputs that already exist, and records what it would have decided. For a rule requiring every workload image to descend from an approved base-image lineage, that means pulling the stored metadata for every image the pipelines built over the last ninety days and evaluating the rule once per image. Nothing is blocked, no message reaches a developer, and the engine is nowhere near anyone's build. What comes out is a table of would-be denials with the service that owns each one. That list is the input to triage, not a score: a would-be denial can be perfectly correct. The point is to replace an opinion about the rule with a measurement of what it does to the population it will actually meet.
go deeper
Be ready to say what a shadow run produces: a recorded verdict for every real input and a list of the changes the rule would have denied, with nothing actually blocked while it runs.
Explain where the inputs come from — stored build metadata rather than invented examples — and why the output still needs to be triaged by hand before any number is quoted from it.
Show that you run it to decide ship or no-ship, and that you can name what the corpus does not contain: rare rebuilds, long-tail services, and anything nobody attempted.
Own the framing that measurement is what makes a guardrail negotiable with the teams it will affect. You want to arrive at that conversation with a number and a named list, not a conviction.
## The move A shadow run — also called a dry evaluation or a backtest of a rule — is the step between writing a rule and letting it decide anything. You take the candidate rule, feed it real inputs that already exist somewhere, and record the verdict it would have produced for each one. Nothing is enforced, nothing is warned, no developer sees a message. The run is a batch job over stored data, not a deployment. ## What it looks like concretely Take a rule that says every workload image must descend from an approved base-image lineage. The surface it reads is image metadata: the labels the build stamped on the image, the recorded base reference, the layer history. That metadata already exists for every image your pipelines produced. So the corpus is, for example, ninety days of built images — a few thousand records — and the run evaluates the rule once per record. The output is a table with one row per image: the image, the service and team that owns it, the verdict, and the message the rule would have printed. Two aggregates fall out of it immediately: how many images the rule would have denied, and how those denials distribute across services and teams. ## Why review is not enough Reading a rule tells you whether it expresses the intent you had. It cannot tell you what the rule does to the population it will meet, because that population is a fact about your estate, not about your rule. Estates are full of history: a service that still builds from a base image approved two years ago, a build that renames labels, a mirrored registry path that is the same base under another name. Every one of those is a change that would be denied by a rule you would have signed off on in review. The asymmetry matters. A rule that misses a bad change costs you a risk you already had. A rule that denies a good change costs you a person's afternoon, an escalation, and — repeated a few times — the credibility of the whole gate. So the measurement you want before shipping is specifically about the good changes. ## What comes out is a list, not a verdict The most common misreading is to treat the shadow denial count as an error count. It is not. A denial can mean the rule is right and the image genuinely descends from something unapproved, in which case you have found real work for a team to do. It can also mean the image complies in substance and the rule read the wrong thing. Separating the two is manual triage, and the shadow run's job is only to hand you the shortlist. The allowed rows deserve a look too. Everything the rule let through is where its misses live, and no aggregate in the report will point at them; you only find them by sampling. ## What it costs and what it risks Almost nothing. The engine is not in anyone's request path, so there is no latency question and no availability question — an evaluation that crashes on a malformed record costs you a rerun. The real cost is the triage labour on the denial list, which is why a corpus is chosen to be representative rather than exhaustive. ## What it cannot tell you A shadow run measures history. It contains only changes people actually made, in a world where this rule did not exist. It does not contain the change nobody attempted, the emergency rebuild that happens twice a year, or the practice teams will adopt once the rule is real. That is a limit to state out loud when you present the result, not a reason to skip the measurement. ## The decision it feeds At the end you can say something a lead can act on: this rule would have denied N changes over ninety days, of which M were legitimate, concentrated in these teams. That sentence is the difference between arguing about a guardrail and deciding about one — and sometimes the honest conclusion it supports is that the rule should not ship at all.
- What corpus would you assemble to shadow-run a rule about base images?The stored metadata for every image the pipelines produced over a recent window — say ninety days — plus the latest image from every service in the estate, however old that build is. The window alone is dominated by whatever rebuilds nightly; the per-service sample is what drags the long tail into the corpus.
- How is a shadow run different from having the rule evaluate in production without blocking?A shadow run happens before the rule is wired to anything: stored documents are replayed offline, so there is no engine in a request path, nothing to be unavailable, and no message reaching a developer. It is a measurement you run to decide whether the rule should exist, not a way of operating one that already does.
- The shadow run denied one service's images 400 times. Is that 400 problems?Almost certainly one. If a service rebuilds the same image nightly, every rebuild reproduces the same verdict. Deduplicate by service or by distinct image content before quoting any number, or a single misconfigured pipeline will dominate the result and hide everything else.
It is the same move as replaying a new fraud rule over last quarter's card transactions before it is allowed to decline anybody: same rule, same real data, no consequences.
saying these in an interview costs you the question
- Treats every shadow denial as a confirmed violation
- Assumes a clean code review means the rule is safe
- Builds the corpus only from the busiest repositories
- Says the run proves nothing because nothing was blocked
- Never looks at what the rule allowed