Your SOC's busiest detection rule fired 40,000 times last quarter - what does that count tell you about its value?
answer
- a big number with unknown meaning
- cost is not the same as value
- count outcomes, not matches
- right rule, authorised actor
- yield is cases per firing
basics
~20 sA firing count measures how often a pattern occurs in the estate, not how often it means an intrusion. Value comes from yield: the share of firings that became real, escalated cases. The busiest rule can yield zero.
solid answer
~50 sVolume and yield are different measurements and only yield speaks to value. A firing is one match against records; a case is what an analyst reached a verdict on; yield is cases per firing. The highest-volume rule on a Windows estate is very often a credential-access rule watching handle opens against `lsass.exe`, because backup agents, crash handlers and the EDR itself legitimately read that process. Forty thousand firings can therefore be forty thousand *benign true positives*: the behaviour really happened, the rule was right about the behaviour, and the actor was authorised. So the number I want printed next to the 40,000 is how many firings became cases and how many escalated. Low yield is a finding about the rule, not yet a verdict on it - a rule can be low-yield because the technique is genuinely rare rather than because the logic is bad.
go deeper
Be ready to separate how often a rule fired from how often it was right, and to name the three closing verdicts: true positive, false positive, benign true positive. Know that a busy rule can produce none of the first.
An interviewer expects you to say what a firing counts against - matched records, alerts raised, cases opened - and to explain why legitimate software such as backup agents and crash handlers reads lsass.exe memory constantly.
Show that you use volume as a cost measure and never as a ranking of worth, and that you refuse to convert low yield straight into a verdict without asking how often the underlying behaviour occurs at all.
Own the framing when this reaches leadership. A chart of the noisiest rules looks like a chart of your strongest controls, and correcting that impression before someone budgets against it is your job, not the analyst's.
## The two numbers people confuse Every detection rule in a SIEM or EDR produces a stream of **firings**: one firing is one match of the rule's logic against the records it reads. A firing usually becomes an **alert** in a queue. An analyst works the alert and opens or attaches it to a **case**, and the case eventually gets a **disposition** - a recorded verdict. *Volume* counts firings. *Yield* counts outcomes per firing: cases opened per hundred firings, or escalations per hundred firings, depending on which unit your SOC has agreed to use. They answer completely different questions: | Measurement | What it actually tells you | |---|---| | Firings per quarter | How often the pattern occurs in this estate, and roughly what the rule costs to work | | Cases or escalations per firing | How often the pattern, when it occurs, meant something a defender had to act on | Volume is a **cost** measurement. Yield is a **value** measurement. A chart that ranks rules by volume ranks them by expense, and it will put the rule that has never once been right at the top. ## The three verdicts, and the one people forget - **True positive** - the behaviour occurred and it was hostile. The rule earned its place on this firing. - **False positive** - the rule was wrong about the records. The thing it claims happened did not happen: it matched a field it misread, a parser artefact, a name collision. - **Benign true positive** - the behaviour occurred exactly as described, and the actor was authorised. The rule was *right* and the answer is still 'nothing to do'. That third verdict is the whole reason a 40,000-firing rule can have zero value in the queue and still be technically correct 40,000 times. Recording benign true positives as false positives destroys the measurement, because 'the logic is wrong' and 'the logic is right about work we authorise' call for entirely different responses. ## The worked case Reads of `lsass.exe` process memory are the classic example. The Local Security Authority Subsystem holds credential material in memory, so a technique catalogued as OS credential dumping from LSASS memory (`T1003.001`) is high on any credential-access detection list, and endpoint telemetry can see it - Sysmon Event ID 10 records one process opening a handle to another and carries the granted-access mask. The problem is that legitimate software opens that handle constantly: backup and imaging agents, crash and error reporting, the antimalware engine, and frequently the EDR product that is also generating the alert. On a fleet of tens of thousands of endpoints that is a five-figure quarterly firing count with a handful of cases behind it, and it is usually the *last* rule anyone volunteers to touch, because the technique it names is genuinely serious. ## What the 40,000 does tell you It tells you cost. Multiply the firings that reached a human by the minutes each one consumed and you have the analyst-hours this single rule spent. It also tells you the shape of the queue - if a dozen rules produce most of the alerts, the SOC's daily experience *is* those dozen rules, whatever the other 888 were written for. ## What low yield does not license It does not by itself prove the rule is bad, and it does not prove the rule is broken. Yield depends on how often the underlying behaviour occurs, and rare-technique rules are supposed to be quiet. It also does not prove the estate is clean: a rule's own output can never tell you what it failed to see. The count is an input to a judgement about the detection set, not the judgement itself.
- What separates a false positive from a benign true positive on that rule?A false positive means the rule was wrong about the records - it claimed a handle open against `lsass.exe` that did not happen, or misread the process. A benign true positive means the read genuinely happened and the reader was a backup agent or the EDR itself. The first says fix the logic; the second says the logic is right and the activity is authorised. Recording them under one label makes both untreatable.
- The 40,000 firings produced three escalations, all authorised backup activity. Is the rule worthless?No - that measurement establishes its cost is terrible and says nothing about whether the technique matters. The rule may be the only content watching that technique, and credential access is not something a SOC wants unwatched. What the number licenses is putting the rule at the top of the list to be examined; the decision about what to do with it is a separate one that has to price what you would stop seeing.
- Which single extra column would you add to a volume chart of the top rules?Cases opened per rule, alongside escalations. Volume plus outcomes is the smallest pair that turns a cost chart into something you can reason about, because it lets you see the rules that are loud and productive, loud and empty, and quiet but almost always right - three completely different situations that a volume ranking flattens into one.
A smoke detector that shrieks every time someone makes toast is not the best detector in the house because it is the loudest. It is the most expensive one, and you still do not know whether it would notice a fire.
saying these in an interview costs you the question
- Treats the busiest rule as the most valuable rule
- Assumes every read of lsass.exe memory is an attacker dumping credentials
- Closes authorised backup activity as a false positive
- Concludes zero cases means the rule must be broken
- Reads a quiet quarter as proof the estate was clean