A candidate rule shadow-run over 2,400 built images denied 118 of them — how do you compute its false-positive rate?
answer
- A count is not a rate
- Triage every denial by hand first
- Pick the denominator deliberately
- Deduplicate rebuilds by service
- The allowed set hides the misses
basics
~20 sTriage all 118 denials first: a denial is a false positive only if the image actually complied. The rate is confirmed false positives over the images evaluated. A raw denial count is not a rate.
solid answer
~50 s118 is a denial count, not a false-positive rate. I go through the 118 by hand and split them: images that genuinely descend from an unapproved base, where the rule is right and a fix exists, versus images that comply in substance but were denied anyway — usually because the lineage label was renamed, dropped by a later build stage, or points at an approved base mirrored under a different registry path. Only the second group are false positives. Then I report two numbers: false positives over denials, which says how much a red result can be trusted, and false positives over all 2,400, which converts into `this rule would have wrongly stopped N builds a week` — the number a lead can decide on. I also sample the 2,282 images the rule allowed, because misses never appear in a denial list.
go deeper
Know that a shadow run's denial count has to be checked case by case before any part of it can be called a false positive, because a denial is often the rule doing its job.
Be ready to do the arithmetic out loud: which denials were genuine, which denominator you chose and why, and what the resulting figure means in builds per week.
Show that you look past the headline: concentration across teams, repeated rebuilds of the same image, and a sample of the allowed set to catch what the rule is missing.
Own the standard that a ship decision needs one honest number plus the named population it lands on, measured on data the rule was not tuned against.
## The count is not the rate A shadow run over ninety days of image metadata that returns 118 denials out of 2,400 evaluations has told you one thing: the rule fires on about five per cent of the estate. It has not told you whether it is wrong about any of them. A rule that says every workload image must descend from an approved base-image lineage is *supposed* to deny images that do not. Those denials are the rule working. So the first step is triage, and it is manual. Each denied image goes into one of two buckets: - **True positive** — the image really does descend from something outside the approved lineage. The rule is right, the finding is real work, and the owning team has an action available to them. - **False positive** — the image complies in substance and the rule denied it anyway. In practice these cluster into a few shapes: the lineage label was renamed by a build template change, a later build stage produced the final image and dropped the label, or the base is an approved image mirrored under a different registry path so a string comparison failed. Suppose the split is 91 true and 27 false. ## Choosing the denominator on purpose Two rates are computable from that split and they answer different questions. | Rate | Arithmetic | What it answers | | --- | --- | --- | | False positives over denials | 27 / 118 = 23% | How much should a developer trust a red result from this rule? | | False positives over evaluations | 27 / 2400 = 1.1% | How much innocent work would this rule have stopped? | The first is a quality statement about the rule. The second is the ship/no-ship number, because it converts into something a lead can weigh: 27 wrongly denied builds across ninety days is roughly two a week, landing on whichever teams the triage named. Quote both, and say which one you are quoting — a bare percentage is ambiguous between them. ## Corrections that change the number **Deduplicate by service, not by build.** If one service rebuilds the same image nightly, ninety days of it contributes ninety rows that are all the same finding. Count distinct services or distinct image contents before quoting anything, or one broken pipeline will dominate the report. **Look at concentration.** 27 false positives spread across twenty teams is a rule problem — it is misreading a common pattern. 27 on a single service is a scoping problem — that service does something unusual and the rule may simply not apply to it. The same number supports two opposite decisions depending on how it is distributed. **Sample the allowed set.** The 2,282 images the rule permitted are where its misses live. Pull a small random sample, check by hand whether any of them descend from an unapproved base, and you get a rough sense of whether the rule is passing images it should catch. A rule that reads a field which is simply absent on many images can allow almost everything and still look immaculate in the denial report. ## Do not measure on the window you tuned against The single most common way this number becomes fiction: you run the shadow evaluation, you see 118 denials, you fix the rule so the 27 obvious false positives stop, and then you quote the new rate from the same corpus. That corpus is now a training set — the rule has been fitted to it, and its measured error there is optimistic by construction. Re-measure on data the rule has not seen: an earlier window, a held-out slice of services, or the next two weeks of builds. ## What you actually report The deliverable is not a percentage. It is: the rule would have denied 118 builds over ninety days, 91 of them correctly; the 27 wrong denials fall on these four teams and have these three causes; two of the causes are fixable in the rule and one is not. That is a statement someone can make a decision from, and it may still support a decision not to ship.
- Two denominators are available — denials or all evaluated images. When would you use each?False positives over denials tells a developer how much to trust a red result, so it argues about the rule's quality. False positives over all evaluations tells a lead how much legitimate work the rule would stop per week, so it argues about whether to ship. Quote both, and name which one a number refers to.
- You fixed the rule after seeing the first shadow result. Is the new rate still valid?No. Once the rule has been fitted to the denials that window produced, its measured error on that window is optimistic by construction. Re-measure on data the rule has not seen — an earlier window, a held-out set of services, or the next few weeks of builds — before quoting a rate to anyone.
- Triage of 118 denials is expensive. How do you cut the work without faking the number?Collapse duplicates first: distinct services and distinct image contents, not raw builds, which usually shrinks 118 rows to a few dozen cases. Then group by denial message, because identical messages tend to share a root cause, and triage one representative per group before checking whether the group is homogeneous.
saying these in an interview costs you the question
- Reports the raw denial count as the false-positive rate
- Assumes every denial is a genuine violation
- Ignores that the allowed set can hide misses
- Counts builds instead of distinct services
- Quotes a rate from the window the rule was tuned on