skip to content

How do you compute per-rule true-positive yield across 900 SIEM detection rules?

level: middleimportance: should knowfreq 46%

answer

  1. two systems, one stable id
  2. the unit decides the answer
  3. age and scope are denominators
  4. who gets credit for the case
  5. closing codes cap the quality

basics

~20 s

Join the SIEM's own rule-execution metadata to case dispositions - firings, cases opened and verdicts per rule - and normalise per rule-day live, since a rule deployed last month is not comparable to one live for four years.

solid answer

~50 s

The data lives in two places: the platform's rule-execution metadata (rule id, firings, last-fired, enabled state, deployment date) and the case system (which rule opened the case, and the verdict recorded). I join on a stable rule id and fix the unit first - firings and cases are not interchangeable, because one intrusion can consume five hundred firings of a single rule inside one case. Then I normalise per rule-day live and per asset in scope, so a six-week-old rule is not compared with a four-year-old one. The number is only as good as the disposition vocabulary: if benign true positive is not a distinct closing code from false positive, yield cannot distinguish a wrong rule from a right rule that keeps describing authorised work. Corroborating rules that fire inside somebody else's case also score zero unless you credit contribution as well as origination.

code

text · 7 lines
text
rule_id  name                                    days_live  firings  cases  escalated  last_fired
DR-0042  LSASS handle open (T1003.001)                1621    40118     61          3  2026-08-27
DR-0117  Security audit log cleared (event ID 1102)    980      212     97         11  2026-08-19
DR-0508  kubectl exec into a production pod            412     1830   1802          4  2026-08-27
DR-0733  Federation trust added to the tenant          640        3      3          2  2026-05-02
DR-0777  New OAuth grant to an unverified app           41        0      0          0  never
...      (895 further rules)

go deeper

for a junior

Know that this number is not a built-in report. It is a join between what the detection platform recorded about its rules and what the case system recorded about the verdicts analysts reached.

for a middle

Be ready to define the unit explicitly - firings, alerts or cases - and to say why they differ by orders of magnitude when one incident consumes hundreds of firings from a single rule.

for a senior

Demonstrate the normalisations that stop the table lying: per rule-day live, scope size, enabled state through the window, and separating cases a rule originated from cases it merely corroborated.

for a principal

The measurement you standardise becomes the language your content review argues in. Insist on a controlled disposition vocabulary first, because no amount of analysis rescues yield computed from free-text closures.

## Two datasets, one join Per-rule yield is not a metric your platform ships; it is a join you build. **Side one - rule-execution metadata.** Every detection platform records, per rule: a rule id, an enabled/disabled state, a schedule or streaming mode, when it was deployed or last edited, how many times it executed, how many times it matched, and when it last matched. This is the side that tells you a rule exists and whether it produces anything. **Side two - case dispositions.** The case management system records, per case: which alert opened it, which rule that alert came from, who worked it, how long it took, and the closing verdict. This is the side that tells you whether the firings meant anything. The join key must be a **stable rule id**, not the rule's name. Names get edited; a portfolio that joins on display name silently splits one rule's history into three the moment somebody fixes a typo. ## Fix the unit before you compute anything The single largest error in this measurement is mixing units. Three candidates, and they differ by orders of magnitude: - **firings** - every match. One noisy rule can match half a million times on a fleet. - **alerts reaching a human** - firings that survived aggregation and grouping. - **cases** - what an analyst actually opened a verdict on. One case can swallow hundreds of firings from the same rule. Yield expressed as *cases per firing* and yield expressed as *escalations per case* are both defensible; a table that quietly uses one for some rules and the other for others is not. Write the definition down and hold it across the whole portfolio, because the entire point of this exercise is comparing rules to each other. ## Normalise, or the ranking is an artefact of age A rule live for four years and a rule deployed six weeks ago cannot be compared on raw counts. The minimum normalisation is **per rule-day live** - firings and cases divided by the days the rule was enabled during the window. Two further denominators matter in a real estate: the **number of assets or accounts in the rule's scope** (a rule scoped to twelve domain controllers will never out-fire one scoped to forty thousand workstations), and **whether the rule was enabled for the whole window** at all. Rules disabled mid-window for an incident or a migration produce a truncated numerator against a full denominator, which makes healthy rules look dead. ## Attribution decides who gets the credit Most case systems attribute a case to the rule whose alert *opened* it. That choice has consequences: - A rule that fires ten minutes later and confirms the same activity records zero cases forever, even though it is doing real work. If your platform supports it, count **contributing** firings on a case separately from **originating** ones, and report both. - Rules that feed a risk-scoring or aggregation layer instead of raising their own alerts will show a firing count and no cases at all. They need their own bucket; scoring them as zero-yield detections is simply wrong. - Deduplication and suppression upstream of the queue mean some firings never reach a case by design. If suppression happens in the pipeline rather than the rule, your denominator includes matches that no human could ever have converted. ## The disposition vocabulary is the ceiling on quality Yield inherits every weakness of the closing verdicts it counts. Free-text closures are uncountable. A vocabulary with only *true positive* and *false positive* forces analysts to record authorised, correctly-detected activity as a false positive, which makes a rule that is working perfectly look defective. At minimum the closing codes need to separate: hostile activity confirmed; authorised activity correctly identified (benign true positive); the rule was wrong about the records; and insufficient evidence to decide. Whether those recorded verdicts are themselves *correct* is a different exercise from this one, and it is worth saying so out loud rather than pretending the join produces truth. ## What you produce One row per rule: id, name, days live, firings, firings per rule-day, cases originated, cases contributed to, escalations, verdict mix, last fired. Read as a distribution rather than a league table, that single table is what a content review actually argues over.

  • One rule fired 500 times inside a single confirmed intrusion. How does that land in the table?
    It depends entirely on the unit. Counted as firings it is 500 hits with one case, which reads as terrible yield; counted as cases it is one case, one escalation, perfect yield. Both are true statements about the same events, which is why the definition has to be fixed portfolio-wide before anyone reads the numbers. I would report cases as the primary unit and keep firings as the cost column beside it.
  • Why join on rule id rather than rule name?
    Names are edited - a typo fix, a retitle to match a technique identifier, a team prefix added. Joining on name splits one rule's history across several rows and makes each fragment look newer and quieter than the rule really is. A stable id issued at creation and never reused is what keeps a rule's firings, cases and deployment date attached to each other across the window.
  • What do you do with rules that never raise their own alerts because they feed a risk score?
    Bucket them separately and measure them differently. They will always show firings with zero cases, so including them in a yield ranking guarantees they sit at the bottom looking worthless. Their contribution is to the score that raised somebody else's case, so the honest measurement is how often their signal was present on cases that escalated - not cases they originated.

saying these in an interview costs you the question

  • Mixes firings and cases in the same yield column
  • Ranks a six-week-old rule against a four-year-old one on raw counts
  • Joins the two datasets on the rule's display name
  • Scores corroborating rules as zero-yield because they never open cases
  • Treats recorded dispositions as unquestionably correct

context