skip to content

Why does a SOC sample already-closed alerts and re-review their verdicts?

level: juniorimportance: should knowfreq 50%

answer

  1. nothing downstream disagrees with a close
  2. a wrong dismissal is silent
  3. the only feedback triage ever receives
  4. re-work it blind, then compare

basics

~20 s

A closed verdict is the one SOC output nothing downstream contradicts: a wrong dismissal produces no ticket and no complaint. Re-working a sample of closed cases blind is the only routine way an intrusion hidden inside a dismissal resurfaces.

solid answer

~50 s

Almost everything else in a SOC gets a second look. An escalation is reviewed by another analyst, a declared incident is reconstructed afterwards, a noisy rule shows up as queue volume. A close gets nothing: the case leaves the queue and the record says an analyst decided it was not malicious. If that decision was wrong - the credential-audit alert that really was an intruder, the odd sign-in that really was a stolen token - nothing argues back, because an adversary's whole objective is to generate no further complaint. So the SOC deliberately re-opens a sample: draw a set of closed cases, hide the original disposition and notes, have a different analyst work them from the raw evidence, then compare verdicts. The output is not a score for whoever closed them; it is a list of rules, playbooks, reference data and enrichment gaps that made the wrong close easy.

go deeper

for a junior

Be ready to say why a wrongly closed alert is harder to notice than almost any other mistake in a SOC, and to describe the basic loop: sample closed cases, hide the verdict, work them again, compare.

for a middle

Explain the mechanics that make the result trustworthy - blinding, independence, adjudicating disagreements, and the retention window that limits which cases can be re-worked at all.

for a senior

Show that you would turn findings into detection content and reference data rather than into individual feedback, and that you know a clean sample of thirty certifies nothing about a population of forty thousand.

for a principal

Own the framing you give leadership: this is quality control over detection content with a remediation loop and a named owner, and it stops producing usable data the moment its results are attached to people.

## The problem this practice exists to solve Every alert a SOC works ends in a disposition, and the overwhelming majority end in some form of "not an incident". That verdict is then treated as fact by everything downstream: the queue empties, the metrics count it, the tuning backlog reads it, and the case is never opened again. What makes that dangerous is the **asymmetry of the two error types**. If an analyst wrongly escalates a harmless case, the error surfaces almost immediately - a second analyst, an incident lead or an engineer whose host is about to be isolated pushes back, and the mistake is corrected within hours. If an analyst wrongly closes a real intrusion, **nothing happens at all**. There is no failing test, no bounced ticket, no user complaint. The adversary is actively working to keep it that way. The error is silent by construction, and it stays silent until the intrusion becomes loud enough to be found some other way - often by a third party telling you. So a closed verdict is the only major SOC output with no feedback path. A verdict audit builds one on purpose. ## What the practice actually is The core loop is a **blind re-review** of a sample: 1. **Draw a sample** of closed cases from a period recent enough that the underlying telemetry still exists. 2. **Strip the answer.** The reviewer gets the original alert and the raw evidence, not the disposition, the case notes or the name of the analyst who closed it. 3. **Re-work the case independently** to a fresh verdict. 4. **Compare** the two verdicts, record disagreements *before* anyone discusses them, and adjudicate the disputes with a third, more senior reviewer. 5. **Convert findings into content changes** - a rule that needs context, an enrichment step that was missing, a playbook that told the analyst to close on a weak signal, a piece of authorised-activity reference data nobody had. The important design commitment is in step 5. A verdict audit is **content quality control**, not an examination of people. The moment its output is a per-analyst overturn table, analysts stop closing anything and the queue - and the data - collapse. ## What the audit can and cannot tell you It can tell you that a class of cases is being systematically mislabelled: for example that alerts on an administrator's credential-audit tooling are routinely recorded as "false positive" when the rule was in fact correct and the activity merely authorised. That is a finding with real teeth, because a false-positive label is the normal justification for suppressing a rule, and suppressing it carves out exactly the behaviour an intruder would reuse. It cannot certify the population. Thirty re-reviewed cases out of forty thousand closures bound the gross error rate very loosely, and a targeted intrusion is a rare event that a small random sample will usually miss entirely. "We sampled thirty and found nothing" means the process is not obviously broken; it does not mean the queue is clean. It is also bounded by **retention**. You can only re-review a case while the evidence behind it still exists - the process telemetry, the sign-in logs, the proxy or flow records the original analyst looked at. Once those age out, the case is un-auditable, which is why the audit runs on a cadence inside the hot retention window rather than as an annual exercise. ## How it differs from neighbouring measurements Counting how many alerts a rule produced tells you about the rule's output volume. Counting how many were escalated tells you what analysts did with it. Neither tells you whether the *verdicts* were right, because both are computed from the same dispositions whose correctness is in question. That circularity is the reason the audit has to go back to the raw evidence rather than to the ticket text. ## What good looks like A mature programme runs a small blind sample on a regular cadence, seeds it with a few cases whose true answer is already known (an in-scope red-team action, an alert that preceded a confirmed incident), reports findings at the level of rules and playbooks, tracks each finding to a named owner and a fix, and publishes what changed. Analysts cooperate with an audit that repairs their tooling; they defend themselves against one that grades them.

  • How is this different from just measuring how many alerts each detection rule produced?
    Volume tells you what the rule emitted; the audit tells you whether the verdicts on that output were right. The two are not substitutes, because rule metrics are computed from analyst dispositions - if the labels are wrong, a rule can look perfectly tuned precisely because every case was closed quickly and wrongly.
  • Where would you expect a verdict audit to find its worst cases?
    In the fastest closures; on high-volume rules where a standing assumption has grown up that this alert is always benign; in hours when the queue is deepest; and in cases whose host, account or domain later turned up in a different alert. Those are the strata worth over-sampling.
  • The audit re-reviewed thirty closed cases and found no errors. Has it proved the queue is clean?
    No. Thirty cases loosely bound the gross error rate and say nothing about a targeted intrusion the sample never touched. It shows the process is not obviously broken. Sample size, seeded known-answer cases and cadence are what make the result mean more over time.

A hospital re-reads a sample of scans that were already reported as normal. Nothing else in the workflow ever revisits a normal report, so the only way a missed finding comes to light is by deliberately looking again.

saying these in an interview costs you the question

  • Claiming no complaint after a close proves the alert was benign
  • Treating the audit as a way to rank individual analysts
  • Assuming the platform would re-alert if the first close was wrong
  • Sampling only the cases the analyst already flagged as uncertain
  • Re-reading the case notes instead of the raw evidence

context