skip to content

Measuring Detection & Response

How good is the SOC: detection and containment intervals, dwell time inferred after the fact, and coverage claims a heatmap cannot support. Interviewers use it to separate operators from believers.

on this pageshow

explore

questions

29

Your SOC's busiest detection rule fired 40,000 times last quarter - what does that count tell you about its value?

level: juniorimportance: must knowfreq 58%

answer

  1. a big number with unknown meaning
  2. cost is not the same as value
  3. count outcomes, not matches
  4. right rule, authorised actor
  5. yield is cases per firing

basics

~20 s

A firing count measures how often a pattern occurs in the estate, not how often it means an intrusion. Value comes from yield: the share of firings that became real, escalated cases. The busiest rule can yield zero.

solid answer

~50 s

Volume and yield are different measurements and only yield speaks to value. A firing is one match against records; a case is what an analyst reached a verdict on; yield is cases per firing. The highest-volume rule on a Windows estate is very often a credential-access rule watching handle opens against `lsass.exe`, because backup agents, crash handlers and the EDR itself legitimately read that process. Forty thousand firings can therefore be forty thousand *benign true positives*: the behaviour really happened, the rule was right about the behaviour, and the actor was authorised. So the number I want printed next to the 40,000 is how many firings became cases and how many escalated. Low yield is a finding about the rule, not yet a verdict on it - a rule can be low-yield because the technique is genuinely rare rather than because the logic is bad.

go deeper

for a junior

Be ready to separate how often a rule fired from how often it was right, and to name the three closing verdicts: true positive, false positive, benign true positive. Know that a busy rule can produce none of the first.

for a middle

An interviewer expects you to say what a firing counts against - matched records, alerts raised, cases opened - and to explain why legitimate software such as backup agents and crash handlers reads lsass.exe memory constantly.

for a senior

Show that you use volume as a cost measure and never as a ranking of worth, and that you refuse to convert low yield straight into a verdict without asking how often the underlying behaviour occurs at all.

for a principal

Own the framing when this reaches leadership. A chart of the noisiest rules looks like a chart of your strongest controls, and correcting that impression before someone budgets against it is your job, not the analyst's.

## The two numbers people confuse Every detection rule in a SIEM or EDR produces a stream of **firings**: one firing is one match of the rule's logic against the records it reads. A firing usually becomes an **alert** in a queue. An analyst works the alert and opens or attaches it to a **case**, and the case eventually gets a **disposition** - a recorded verdict. *Volume* counts firings. *Yield* counts outcomes per firing: cases opened per hundred firings, or escalations per hundred firings, depending on which unit your SOC has agreed to use. They answer completely different questions: | Measurement | What it actually tells you | |---|---| | Firings per quarter | How often the pattern occurs in this estate, and roughly what the rule costs to work | | Cases or escalations per firing | How often the pattern, when it occurs, meant something a defender had to act on | Volume is a **cost** measurement. Yield is a **value** measurement. A chart that ranks rules by volume ranks them by expense, and it will put the rule that has never once been right at the top. ## The three verdicts, and the one people forget - **True positive** - the behaviour occurred and it was hostile. The rule earned its place on this firing. - **False positive** - the rule was wrong about the records. The thing it claims happened did not happen: it matched a field it misread, a parser artefact, a name collision. - **Benign true positive** - the behaviour occurred exactly as described, and the actor was authorised. The rule was *right* and the answer is still 'nothing to do'. That third verdict is the whole reason a 40,000-firing rule can have zero value in the queue and still be technically correct 40,000 times. Recording benign true positives as false positives destroys the measurement, because 'the logic is wrong' and 'the logic is right about work we authorise' call for entirely different responses. ## The worked case Reads of `lsass.exe` process memory are the classic example. The Local Security Authority Subsystem holds credential material in memory, so a technique catalogued as OS credential dumping from LSASS memory (`T1003.001`) is high on any credential-access detection list, and endpoint telemetry can see it - Sysmon Event ID 10 records one process opening a handle to another and carries the granted-access mask. The problem is that legitimate software opens that handle constantly: backup and imaging agents, crash and error reporting, the antimalware engine, and frequently the EDR product that is also generating the alert. On a fleet of tens of thousands of endpoints that is a five-figure quarterly firing count with a handful of cases behind it, and it is usually the *last* rule anyone volunteers to touch, because the technique it names is genuinely serious. ## What the 40,000 does tell you It tells you cost. Multiply the firings that reached a human by the minutes each one consumed and you have the analyst-hours this single rule spent. It also tells you the shape of the queue - if a dozen rules produce most of the alerts, the SOC's daily experience *is* those dozen rules, whatever the other 888 were written for. ## What low yield does not license It does not by itself prove the rule is bad, and it does not prove the rule is broken. Yield depends on how often the underlying behaviour occurs, and rare-technique rules are supposed to be quiet. It also does not prove the estate is clean: a rule's own output can never tell you what it failed to see. The count is an input to a judgement about the detection set, not the judgement itself.

  • What separates a false positive from a benign true positive on that rule?
    A false positive means the rule was wrong about the records - it claimed a handle open against `lsass.exe` that did not happen, or misread the process. A benign true positive means the read genuinely happened and the reader was a backup agent or the EDR itself. The first says fix the logic; the second says the logic is right and the activity is authorised. Recording them under one label makes both untreatable.
  • The 40,000 firings produced three escalations, all authorised backup activity. Is the rule worthless?
    No - that measurement establishes its cost is terrible and says nothing about whether the technique matters. The rule may be the only content watching that technique, and credential access is not something a SOC wants unwatched. What the number licenses is putting the rule at the top of the list to be examined; the decision about what to do with it is a separate one that has to price what you would stop seeing.
  • Which single extra column would you add to a volume chart of the top rules?
    Cases opened per rule, alongside escalations. Volume plus outcomes is the smallest pair that turns a cost chart into something you can reason about, because it lets you see the rules that are loud and productive, loud and empty, and quiet but almost always right - three completely different situations that a volume ranking flattens into one.

A smoke detector that shrieks every time someone makes toast is not the best detector in the house because it is the loudest. It is the most expensive one, and you still do not know whether it would notice a fire.

saying these in an interview costs you the question

  • Treats the busiest rule as the most valuable rule
  • Assumes every read of lsass.exe memory is an attacker dumping credentials
  • Closes authorised backup activity as a false positive
  • Concludes zero cases means the rule must be broken
  • Reads a quiet quarter as proof the estate was clean

context

open as a page

A technique cell on your ATT&CK coverage heatmap is green - what three different claims could that colour be making?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Green usually stands for one of three very different facts: the log source that would show the behaviour is collected, a detection rule exists over that data, or the behaviour was actually executed and the rule caught it.

open as a page

Why can a SOC measure its false-positive rate but not its false-negative rate?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Every alert that fires gets a verdict, so false positives are countable. An intrusion nobody detected leaves no case and no verdict, so a miss rate has neither a numerator nor a denominator you can read out of SOC data.

open as a page

In SOC outcome reporting, what events start and stop the MTTD and MTTC clocks?

level: juniorimportance: must knowfreq 72%

basics

~20 s

MTTD runs from the adversary's first malicious action to the moment your organisation knows it is compromised. MTTC runs from that same moment to containment actually executed. Dwell time is the MTTD span measured on one intrusion.

open as a page

What does 'zero security incidents this quarter' actually prove about your estate?

level: juniorimportance: must knowfreq 52%

basics

~20 s

It proves only that nothing was detected, worked and declared. Zero is consistent with a genuinely quiet estate and with an intrusion nobody saw. The number on its own cannot tell you which one you had.

open as a page

Which date starts the dwell clock when the first malicious action predates every surviving log record?

level: middleimportance: must knowfreq 58%

basics

~20 s

Anchor on the earliest date the malicious code could have acted — a deployment or change record often survives when host telemetry has rotated. Retention gives a floor on dwell, never the start date, and every published span must carry the source that anchored it.

open as a page

An auto-close criterion closes any process alert whose parent image is your patch agent - how would an intruder abuse it?

level: middleimportance: must knowfreq 56%

basics

~20 s

Parent process is adversary-influenceable. Getting code launched by the deployment agent, spoofing the parent process id, or dropping a binary at the allowlisted path all satisfy the criterion, so the one case that mattered is closed unread.

open as a page

Your verdict audit finds a third of false-positive closures were authorised activity the rule caught correctly - what has that mislabel already broken?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Those closures assert the rule was wrong when it was right. A false-positive label justifies suppression, so the mislabel licenses tuning that carves out the exact behaviour the detection exists to catch, and every yield figure inherits it.

open as a page

Why does a SOC sample already-closed alerts and re-review their verdicts?

level: juniorimportance: should knowfreq 50%

basics

~20 s

A closed verdict is the one SOC output nothing downstream contradicts: a wrong dismissal produces no ticket and no complaint. Re-working a sample of closed cases blind is the only routine way an intrusion hidden inside a dismissal resurfaces.

open as a page

Auto-closing a detection's alerts versus deleting the rule - what changes about catching an intruder?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Auto-closing keeps the rule firing and writing records you can search later; only the human verdict is automated. Deleting the rule ends detection entirely, so an intruder using that technique leaves nothing behind to find.

open as a page

How do you compute per-rule true-positive yield across 900 SIEM detection rules?

level: middleimportance: should knowfreq 46%

basics

~20 s

Join the SIEM's own rule-execution metadata to case dispositions - firings, cases opened and verdicts per rule - and normalise per rule-day live, since a rule deployed last month is not comparable to one live for four years.

open as a page

What does a green T1218 cell overstate when only two of its signed-binary proxy sub-techniques were tested?

level: middleimportance: should knowfreq 48%

basics

~10 s

It overstates the unit of the claim. Sub-techniques use different binaries, command lines and parent processes, so catching rundll32 and regsvr32 says nothing about the ten siblings nobody executed.

open as a page

How do you select and blind thirty closed alerts for an independent re-review?

level: middleimportance: should knowfreq 38%

basics

~20 s

Draw a stratified random sample across rules, analysts, shifts and close speed, then hand the reviewer the raw evidence with the verdict, notes and analyst name stripped. An unblinded re-review measures agreement with a label already seen.

open as a page

What are the sources that reveal an intrusion your SOC never alerted on?

level: middleimportance: should knowfreq 52%

basics

~20 s

Four channels: an outside party tells you, an authorised emulation exercise runs a behaviour and nothing fires, a hunt turns up unalerted activity, or a later investigation's back-timeline reaches earlier activity. Each finds a different class of miss.

open as a page

What does a purple-team test where your detection fired prove about a quiet quarter?

level: middleimportance: should knowfreq 42%

basics

~20 s

It proves the chain from telemetry to rule to analyst worked for the one technique executed, on the hosts in scope, at that hour. It is a positive control for the pipeline, not evidence no adversary was present.

open as a page

400 of your SOC's 900 detection rules have never fired - what does that entitle you to conclude?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A never-fired rule has at least four causes - never really enabled, no telemetry reaching it, the behaviour genuinely absent here, or logic that cannot match - so portfolio silence proves nothing about how much you are missing.

open as a page

Your Kubernetes estate emits no telemetry for a technique your Windows fleet detects - how should that coverage cell read?

level: seniorimportance: should knowfreq 40%

basics

~20 s

As a distinct no-visibility state, not a failed detection. Red says a rule is missing and detection engineering can fix it; no telemetry says the raw material is absent, a platform team owns it, and it costs money.

open as a page

Median time-to-close falls each week as your alert backlog grows - improvement or clearing?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Falling close time is good news only if you can point at the change that caused it. Compare within one rule and severity, not across the mix, then settle it by blind re-reviewing a sample weighted to the fastest closures.

open as a page

While scoping an intrusion you find an alert closed as a false positive nine months ago — what do you do?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Treat it as evidence first: the intrusion starts nine months earlier, so re-scope that window before retention expires. Then correct the verdict — a rule that fired and was misjudged is a triage failure, not a coverage gap.

open as a page

Auto-close closed 4,000 alerts last quarter and a confirmed intrusion started in that window - how do you find the swallowed alert?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Query the closure store with what the investigation now gives you - hosts, accounts, times, artefacts - not by re-reading 4,000 cases. Whether you can answer at all depends on the closure retaining the original event and the values that matched.

open as a page

How do you present a quarter with no intrusion as evidence a detection programme worked?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Report what the programme observably did: near-misses contained early, alerts correctly closed as benign, techniques executed and detected, and which sources were reporting. State the 'we were blind' reading in the same deck.

open as a page

Your MTTD improved 40% quarter-over-quarter after the clock's start event was redefined — what do you present to the board?

level: principalimportance: should knowfreq 36%

basics

~20 s

Split the change into its definitional and performance parts before presenting anything: recompute both quarters under both definitions, show how much of the 40% is the new clock and how much is faster detection, and break the trend line where the definition changed.

open as a page

You know of three missed intrusions, each surfaced by a different channel — can you state a false-negative rate?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

No. Three misses are observations of an unknown total, drawn by channels with different and unknown chances of surfacing anything. You can state a floor of three, and an upper bound on your detection rate — never a rate.

open as a page

Your quarterly dwell figure is the median over cases closed that quarter — what does that number miss?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Aggregating over cases closed in a quarter is censored: long investigations close in later quarters, so short cases dominate and the median is biased low. It also describes only intrusions you detected, and it changes retroactively when a case is reopened.

open as a page

In-house detections yield ten times more cases per rule than your managed vendor pack - how do you act on that?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat the gap as a confounded comparison before treating it as a vendor failure: in-house rules exist because someone already saw the problem here, and a pack written for every customer carries content for platforms you do not run.

open as a page

Before your ATT&CK coverage heatmap goes into a customer questionnaire under your signature, what do you change?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

Everything the internal version leaves implicit: what each colour asserts and when it was last shown true, which platforms it covers, and any cell you cannot support, downgraded to a dated, owned blind spot.

open as a page

How do you run a closed-verdict audit without it becoming analyst performance management?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Fix the unit of analysis before the first sample: findings attach to rules, playbooks and enrichment gaps, reporting is aggregate, nobody is named. Agree with managers and HR up front what the results may never feed.

open as a page

You are deleting the only detection for a named ATT&CK technique - who signs that off, and what must the record say?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

A named risk owner with the budget to fund the alternative signs it, not the SOC alone. The record names the technique and estate scope, the residual visibility, the options priced, an expiry date and the trigger to revisit.

open as a page

The CFO reads two quiet years as over-investment and wants your second analyst cut - what do you argue?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Argue about capability, not fear. Say what the second person actually delivers in hours and cases, present the options with their honest costs, and let the budget holder choose knowingly. Never claim a number of breaches prevented you cannot evidence.

open as a page