Your shadow run over ninety days of real changes shows zero false positives — why is that not enough to enforce the rule?
answer
- History is not the future
- Completed changes are a survivorship sample
- Busy services crowd out the tail
- Zero denials can mean zero applicability
- Sample per service, not per build
basics
~20 sBecause the corpus only holds changes people actually made and completed within that window. Rare events like a quarterly base refresh or an incident rebuild fall outside it, and a zero can also mean the rule never fired at all.
solid answer
~50 sA clean result can fail in two ways that need separate checks. The population may be unrepresentative: ninety days of builds is dominated by whatever rebuilds nightly, while the services that ship twice a year, the annual vendor upgrade and the emergency rebuild during an incident are exactly the cases a base-image lineage rule will meet. Or the rule may never have applied: if the metadata it reads is absent on most images and absence is treated as acceptable, zero denials means zero applicability, which reads identically to a clean estate. So I count how many inputs actually reached the comparison, extend the window to cover one full cycle of the slowest recurring rebuild, sample one build per service rather than every build, and take the result to a few owning teams to ask whether a fix would have existed.
go deeper
Know that a shadow run covers only the changes that happened inside the window it looked at, so a clean result is bounded by what was in the corpus.
Be able to name what a build-frequency-weighted corpus over-represents and what it drops, and to explain why zero denials might mean the rule never had an opinion.
Demonstrate that you check applicability and impact separately, deliberately pull long-tail and off-normal builds into the corpus, and confirm findings with the teams that own them.
Own the habit of presenting a result with its sampling limits attached, so a clean run is read as bounded evidence rather than as clearance to enforce.
## Two different zeros A shadow run that produces no false positives is either good news or no news, and the report looks the same either way. Pulling them apart is the senior move. **Zero because the rule never applied.** A rule requiring every workload image to descend from an approved base-image lineage has to read something — a label the build stamped, a recorded base reference. If that field is missing on most of the corpus and the rule treats a missing field as nothing to say about, the run evaluated thousands of images and reached a verdict on a handful. Silence is not compliance. The check is to count inputs that reached the comparison, not inputs that produced a verdict: if 2,400 images went in and 180 carried the field, you measured 180 images and learned nothing about the other 2,220. Worse, a rule that is silent on absent data will be equally silent in production, which is a correctness problem hiding behind a clean measurement. **Zero because the corpus is the wrong sample.** This is the subtler one, and it is structural. ## What ninety days of completed changes systematically excludes - **The long tail.** A window of builds is weighted by build frequency. Services that rebuild nightly contribute hundreds of rows; a stable internal service that ships twice a year contributes none. The second group is precisely where old, unapproved base images survive. - **Recurring events longer than the window.** A quarterly base-image refresh, an annual vendor upgrade, a disaster-recovery rebuild from cold. If the cadence exceeds the window, the corpus contains zero examples of the change that will most reliably trip the rule. - **Off-normal work.** The rebuild done at 3am during an incident, the one-off image built by hand to test a hypothesis. These are made under time pressure, skip the usual templates, and are the changes people remember being blocked. - **Changes nobody made.** The corpus is history under a world where the rule did not exist. It cannot contain the change a team would have made but did not, and it does not predict how practice shifts once the rule is real. This is a sampling problem, not a data-volume problem. Doubling the window to 180 days makes the busy services louder without adding a single long-tail service. ## Firming the corpus up - **Sample by service, not by build.** Take the most recent image from every service in the estate, however old, and union it with the time window. That single change converts a corpus dominated by nightly rebuilds into one that at least touches every population the rule will meet. - **Cover one full cycle of the slowest recurring rebuild.** If base refreshes are quarterly, a ninety-day window covers roughly one; if they are annual, extend or deliberately reconstruct examples from older stored metadata. - **Go and find the off-normal cases.** Ask two or three teams for the last image they built outside the standard pipeline and evaluate the rule against those specifically. A handful of adversarially chosen inputs is worth more than another thousand routine ones. - **Measure applicability alongside impact.** Report how many inputs the rule actually had an opinion about. `The rule denied nothing` and `the rule had an opinion about 7 per cent of the estate` are two different sentences and only the second is honest. ## Closing the loop with the people affected Even a well-built corpus only tells you what the rule would have said, never whether the affected team could have done anything about it. Take the denial list — including an empty one — to a few owning teams and ask two questions: does this match your understanding of how you build, and if this had been denied, what would you have done? The answers surface remediation paths that do not exist, which is the failure a percentage will never show you. ## What you say when you present it State the result and its boundary in the same breath: this rule denied nothing across ninety days covering these services, it had an opinion about this fraction of the estate, and here is what the corpus does not contain. A clean shadow run is evidence, and evidence with its sampling limits attached is the only kind worth acting on.
- How would you get rarely-built services into the corpus?Sample by service rather than by build: take the latest image from every service in the estate, however old that build is, and union it with the time window. A corpus weighted by build frequency is dominated by whatever rebuilds nightly, and the stale bases you care about live in services that almost never rebuild.
- The rule fired on nothing at all. How do you tell a clean estate from a rule that never applied?Count the inputs that reached the comparison, not the ones that produced a verdict. If the metadata the rule reads was absent on most images and absence is treated as acceptable, the run measured only the images that carried the field. That is also a live defect: a rule silent on missing data will be silent in production.
- How much does an extra ninety days of window buy you?Mostly more of what you already have. Extending the window scales with build frequency, so it multiplies the busy services and adds no long-tail ones. It is worth doing only when a specific recurring event — a quarterly or annual base refresh — has a cadence the shorter window could not contain.
Testing a flood barrier against the last three months of rainfall tells you nothing about the storm that arrives every few years — and if the gauge was disconnected, it tells you nothing at all.
saying these in an interview costs you the question
- Treats a clean shadow run as a guarantee
- Samples only the busiest repositories
- Reads zero denials as an empty risk
- Ignores changes nobody attempted
- Fixes sampling bias by enlarging the window