skip to content

How do you select and blind thirty closed alerts for an independent re-review?

level: middleimportance: should knowfreq 38%

answer

  1. not the last thirty, not the easy ones
  2. stratify by rule, analyst, shift, speed
  3. over-sample the fastest closes
  4. strip the verdict, notes and name
  5. agreement measures reproducibility, not correctness

basics

~20 s

Draw a stratified random sample across rules, analysts, shifts and close speed, then hand the reviewer the raw evidence with the verdict, notes and analyst name stripped. An unblinded re-review measures agreement with a label already seen.

solid answer

~50 s

Start from the sample frame: closed cases inside the window where the underlying telemetry still exists, because a case whose evidence has aged out cannot be re-worked at all. Stratify - by detection rule or rule family, by analyst, by shift, and by seconds-to-close - and draw randomly inside each stratum, deliberately over-weighting the fastest closures and the highest-severity rules, since that is where mislabels concentrate. Seed a few cases whose true answer you already know, such as an in-scope red-team action or an alert that preceded a confirmed incident, so you have something to calibrate against. Then blind: the reviewer sees the alert and the raw telemetry, never the disposition, the case notes or who closed it. Reviewers work independently with no discussion, both verdicts are recorded before anyone talks, and disagreements go to a third adjudicator.

go deeper

for a junior

Know the two things that make a re-review meaningful: the cases are chosen rather than convenient, and the second analyst cannot see the original verdict before deciding.

for a middle

Explain the strata you would sample across and why each exists, how blinding is enforced in practice, and why disagreements are recorded before anyone discusses the case.

for a senior

Demonstrate that you read the agreement figure sceptically - base-rate inflation, shared assumptions, seeded known-answer cases - and that you turn the disagreements into named content fixes rather than a score.

for a principal

Own the cadence and scope: how often the audit runs, what it costs the queue, which strata are non-negotiable for high-severity content, and how the reporting stays at the level of rules rather than people.

## Why the sample design carries the whole result A blind re-review is only as good as the cases it pulls. Two failure modes ruin it before a reviewer looks at anything. The first is the **convenience sample**: taking the last thirty closures, or the ones an analyst volunteered, or the ones that happen to have long notes. Those are systematically the cases that were worked carefully, so the audit passes and tells you nothing. The second is the **naively uniform sample**. Closed cases are dominated in volume by a handful of noisy rules. Draw thirty uniformly and you will get twenty-five cases from two rules and nothing from the high-severity content you most need to trust. ## The sample frame Start with what is actually re-reviewable: closed cases inside the retention window for the telemetry behind them. If a case turned on process-creation records or sign-in logs that are now aged out, a reviewer cannot form an independent verdict; they can only read the ticket and agree with it. That constraint sets the audit cadence - it runs inside the hot window, not annually. ## Strata worth using - **Detection rule or rule family.** Guarantees coverage of high-severity, low-volume content that uniform sampling drowns. - **Analyst.** Not to score people, but so a finding is not an artefact of one person's caseload. - **Shift and hour.** Overnight and peak-queue hours behave differently from a quiet Tuesday morning. - **Seconds-to-close.** Over-weight the fastest bucket. Close speed is the cheapest available proxy for how much investigation happened, and it is where systematic mislabels concentrate. - **Disposition.** Include some cases closed as escalate-then-stand-down, not only outright dismissals. A fixed count like thirty is a working batch size a team can actually re-work each cycle, not a statistical guarantee. Treat the numbers it produces as directional and let the cadence, not the batch, build confidence. ## Seeding known answers Agreement between two reviewers measures reproducibility. It does not measure correctness, because both can hold the same wrong standing assumption - for instance that a particular administrative tool on a particular host is always authorised. To get any grip on correctness you need cases with an independent answer: an in-scope red-team action with an execution log to compare against, an alert that preceded a confirmed intrusion, a purple-team run. Seeding a handful of these into the sample gives the audit a small amount of ground truth alongside the agreement figure. ## Blinding mechanics The reviewer receives the original alert, the entity involved and access to the raw telemetry. They do **not** receive the disposition, the case notes, the escalation history or the closing analyst's identity. Two things follow: - Reviewers must work independently and record their verdict before any discussion. A five-minute chat before recording turns two independent judgments into one. - Disagreements go to a named adjudicator, usually a senior analyst or the detection engineer who owns the rule, and the adjudication is recorded as a separate outcome rather than overwriting either verdict. Blinding also runs the other way in the reporting: the write-up names rules, playbooks and gaps, not people. ## Reading the agreement number Suppose the two reviewers agree on twenty-seven of thirty. Three cautions: 1. **Agreement is not accuracy.** It says the process is reproducible, not that it is right. 2. **Raw percent agreement is inflated by the base rate.** When one disposition dominates the population - and in a SOC queue it heavily does - two reviewers who guessed "not malicious" every time would agree most of the time. This is why chance-corrected agreement statistics exist and why a raw percentage alone is a weak claim. 3. **The disagreements are the value.** Three disputed cases that all trace to the same missing piece of context is a better finding than the twenty-seven that matched. ## What comes out The deliverable is a short list of content findings with owners: this rule needs an authorised-activity reference before an analyst can close it correctly; this playbook licenses a close on a signal that does not support it; this enrichment step would have changed the verdict; this log source was missing at the moment of decision. Each gets a fix and a date. The agreement figure is the headline; the findings are the product.

  • Two reviewers agree on twenty-seven of thirty cases. What does ninety per cent agreement prove?
    That the process is reproducible, not that it is right - both reviewers can share the same wrong standing assumption. Raw agreement is also inflated when one disposition dominates the population, since agreeing by default is common. Correctness needs seeded cases with a known answer or an adjudicated reference set.
  • Why over-sample the fastest closures rather than drawing uniformly?
    Because seconds-to-close is the cheapest proxy available for how much investigation actually happened, and mislabels concentrate there. A uniform draw buries those cases under the volume of whichever rules are noisiest, so the audit spends its effort on the cases least likely to be wrong.
  • What limits how far back a verdict audit can sample?
    Telemetry retention. A reviewer can only form an independent verdict while the evidence behind the case still exists; past that point they can read the ticket and agree with it, which is not a re-review. That is why the audit runs on a cadence inside the hot retention window.

saying these in an interview costs you the question

  • Sampling the most recent thirty closures as a batch
  • Letting the reviewer read the original case notes
  • Reporting raw agreement as if it were accuracy
  • Letting reviewers discuss cases before recording verdicts
  • Sampling a period whose telemetry has already aged out

context