You know of three missed intrusions, each surfaced by a different channel — can you state a false-negative rate?
answer
- three observations, no denominator
- channels are not random sampling
- you can bound, you cannot rate
- known misses are a floor, not a total
- a designed sample buys a denominator
basics
~20 sNo. Three misses are observations of an unknown total, drawn by channels with different and unknown chances of surfacing anything. You can state a floor of three, and an upper bound on your detection rate — never a rate.
solid answer
~50 sA rate needs a denominator and a random sample, and you have neither. Each channel — an outside notification, a hunt finding, a back-timeline read — has its own unknown probability of surfacing any given intrusion, and those probabilities differ by intrusion type, so the three cannot be scaled up. What is defensible: **a floor** (at least three got through), and therefore that `detected / (detected + known misses)` is an *upper* bound on your true detection rate, since every unknown miss only pushes it down. Often more useful is the **pattern of channels**: if all three arrived from outside, your internal discovery machinery produced nothing all year. For a genuine rate you must buy a denominator by injecting chosen behaviours — a real number about a population you selected, not the intrusions you face.
go deeper
Be ready to say why counting three known misses does not give a percentage, and what a percentage would have needed that you do not have.
Explain selection bias in concrete terms here: each discovery channel has a different chance of surfacing a given intrusion, so the known misses are not a random draw.
Get the direction of the bound right and say it out loud — detected over detected-plus-known-misses is an upper limit on your detection rate, and every unknown miss lowers it.
Be prepared to refuse a number in front of leadership while still leaving the meeting with something usable, and to argue for the spend that buys a denominator you control.
## What a rate would require A false-negative rate is `misses / (misses + detections)`, and computing it honestly needs two things: the full set of intrusions in the period, and — if you intend to estimate rather than enumerate — a sample of that set drawn in a way whose selection probability you know. Three known misses give you neither. They are three observations of a population whose size is unknown, produced by samplers whose sensitivity is unknown and unequal. ## Why the three channels do not form a sample Suppose the three arrived as: one from a sharing-community notification naming your data, one from a hunt, and one from walking back a later incident. Each of those has a different, intrusion-dependent probability of firing at all: - outside notification is near-certain for an intrusion whose data gets published and near-zero for one whose objective was quiet persistence; - a hunt finding depends on retention, on which sources you collect, and on which hypothesis someone wrote down that quarter; - a back-timeline read is conditional on your having caught a later intrusion at all. So the probability of an intrusion appearing in your known-miss list is a function of what kind of intrusion it was. That is textbook selection bias, and it is not correctable, because correcting it would require knowing the selection probabilities — which requires knowing the population you are trying to measure. ## What you can legitimately say **A floor.** At least three intrusions were not detected in the period. This is a lower bound on misses, never an estimate of them. **A bound in the right direction.** If you detected `D` intrusions and know of `K` misses, then `D / (D + K)` is an *upper* bound on your true detection rate: the true miss count `M` satisfies `M >= K`, and increasing the denominator only lowers the ratio. Saying `our detection rate is at most 93%` is honest and useful; saying `our detection rate is 93%` is a fabrication. Interviewers listen specifically for this direction, because getting it backwards — presenting the figure as a floor, as if unknown misses could improve it — is the common error. **A statement about channels.** Every miss should carry the channel that revealed it. Three misses, all from outside notification, is a finding about your own machinery: internal discovery produced zero results, so the number of misses you can ever report is capped by other people's diligence. That is a defensible, actionable sentence, and it does not pretend to be a rate. **A statement about time.** For each known miss you can date the activity and the discovery, which gives an interval per case. Those intervals are also selection-biased upwards, so report them as individual cases rather than as an average that implies a population. ## Capture-recapture, and why it is intuition rather than arithmetic The classic estimator for an unseen population — `N ≈ n1 × n2 / overlap`, where two samplers found `n1` and `n2` items with `overlap` in common — is tempting here and must be handled carefully. It assumes a closed population, independent samplers, and equal catchability across members. All three fail: intrusions arrive and end continuously, the channels are correlated (a loud intrusion is more likely to be found by *every* channel), and catchability varies enormously by intrusion type. What survives is the direction of the intuition, and it is worth stating in an interview because it is uncomfortable and correct: **if two discovery channels keep finding almost entirely different intrusions, that is evidence the unfound population is large.** High overlap would suggest you are close to exhausting the set; disjoint results suggest you are sampling a fraction. Present it as a qualitative argument, never as a number on a slide. ## Buying a denominator The only way to get a real rate is to construct the population yourself: inject a chosen set of adversary behaviours into the estate under authorisation and count, for each one, whether telemetry was produced, whether something matched, and whether a case was worked. Now the denominator is exactly the set you injected, so the rate is arithmetic rather than inference. The cost is that this measures coverage of the behaviours *you chose*. It is a genuine, defensible, repeatable number about a population that is not the population of real intrusions, and it must always be reported with that caption. A team that presents `we detected 74% of executed behaviours` as `we detect 74% of attacks` has swapped one honest number for a dishonest one. ## Reporting it upwards The presentation that survives scrutiny has three parts: the count with provenance (`three known misses; all reported to us by third parties`), the bound with its direction (`our detection rate is at most X`), and a named proxy with its caption (coverage of executed behaviours, or of collection against the estate). Plus one sentence saying explicitly that a false-negative rate does not exist and why. That sentence is the hardest part and the most valuable. Inventing a rate feels cooperative in the meeting and is catastrophic the moment a fourth miss appears and the number moves without anything changing. Refusing to state a number you cannot support, while offering the two you can, is the behaviour the question is testing for.
- You detected 40 intrusions and later learned of 3 misses. What does 40/43 mean?It is an upper bound on your true detection rate, not the rate itself. The three are a floor on misses; any miss nobody has revealed sits in the denominator unaccounted for and can only push the true figure lower. Reporting 93% as the achieved rate inverts the direction of the only inference the data supports.
- All three known misses were reported by outsiders. What is the finding?That the outside is your only functioning discovery channel. Internal machinery — hunting, emulation, review of closed cases — produced nothing, so your miss count is bounded by other people's diligence rather than by your own visibility. It is a measurement of the programme, and it argues for funding a channel that does not require the intrusion to become somebody else's problem first.
- The board wants one number for detection effectiveness. What do you give them?A count with provenance, a bound stated in the correct direction, and a named coverage proxy with an explicit caption about what population it covers — plus one sentence saying a false-negative rate does not exist and why. Inventing a number costs more credibility later, when a fourth miss appears and the figure moves although nothing about the estate changed.
Three shipwrecks found by three different search methods tell you the sea floor is not empty. They tell you nothing about how many wrecks are down there, because each method only looks where it can see.
saying these in an interview costs you the question
- Divides known misses by total alerts to produce a rate
- Presents detected over detected-plus-known-misses as the achieved rate
- Applies capture-recapture as if the channels were independent
- Reports coverage of executed behaviours as coverage of real attacks
- Invents a percentage because leadership asked for one number