Your tuning evidence is 612 analyst verdicts — how do you check the verdicts themselves are right?
answer
- a verdict is a judgement, not truth
- look at the time-to-close distribution
- blind re-review a random sample
- the minority cluster is the signal
- false negatives are absent by construction
basics
~20 sA disposition records one analyst's judgement, not ground truth. Blind re-review a sample, look at the time-to-close distribution, re-check the recorded deciding value against the raw event, and explain the closures that do not fit the majority cluster before requesting any change.
solid answer
~50 sStart from what the data actually is: 612 rows saying an analyst concluded something, under time pressure, on a rule they had learned to distrust. Three checks are cheap and catch most contamination. First, the time-to-close distribution — a run of eighteen-second closes cannot have checked a process ancestry or a destination. Second, a blind re-review of a random sample by a second analyst, judged on the raw events rather than the first verdict. Third, cross-check the recorded deciding value against the underlying event: if the closure claims the executing process was the render daemon, was it? Then look hardest at the minority: the 6% of closures from hosts outside the render farm are not noise to be dropped, they are the part the majority explanation does not cover, and they are exactly where a genuine hijack would sit. Tune on contaminated evidence and you cut a blind spot in precisely that shape.
go deeper
Know that closing an alert records your judgement, not a proven fact, and that a rushed close can feed a wrong conclusion into a rule change later. Be honest in the record about how much you actually checked.
Explain why a set of closed alerts can only ever evidence false positives: a rule that misses produces no record, so silence and safety look identical. Be able to name a cheap reliability check such as the time-to-close distribution.
Show that you validate evidence before acting on it — blind re-sampling, field-level cross-checks, and re-opening the closures the majority pattern does not explain. An interviewer wants to hear you say the outliers are where an intrusion would hide.
Own the standard for what counts as usable tuning evidence across the SOC, and the culture question underneath it: re-review has to read as quality assurance of the data or the data degrades the moment people feel scored.
## What a disposition proves, and what it does not An alert closed `false positive` establishes one thing: **a human decided that this firing was not what the rule was looking for**. It does not establish that the activity was benign, that it was reviewed properly, or that a second analyst would agree. When six hundred of those judgements are aggregated into a request to change a detection, the request inherits every mistake in them — and unlike a rule, a bad verdict leaves no error message. This matters more here than in most places, because the evidence and the defect share a cause. The reason there are 612 closures is that the rule is noisy; the reason the closures may be sloppy is *also* that the rule is noisy. Habituated closing produces exactly the dataset you are about to tune on. ## Three checks that cost almost nothing **Time-to-close distribution.** Plot it. A verdict that took eighteen seconds did not include pulling the process tree or resolving a destination — it was pattern recognition on the alert title. That is not automatically wrong, but a rule whose closures are overwhelmingly sub-thirty-second closures is producing verdicts with very little checking behind them, and the packet should say so rather than hide it. **Blind re-review of a sample.** Take a random sample — thirty is usually enough to see whether there is a systematic problem — strip the original verdict, and have a second analyst work them from the raw events. Disagreement rate is the number you want. Run it as quality assurance of the *evidence*, announced in advance and reported in aggregate, not as an audit of named individuals; otherwise the next month's closures are written defensively and the data gets worse. **Field-level cross-check.** The closure claims a deciding value. Go back to the underlying event and confirm it held. If the analyst wrote that the executing process was the render daemon but the event shows a process running from a user-writable temp path, that closure is not evidence for tuning — it is an incident nobody opened. ## The minority cluster is the whole point The seductive move is to describe the 94% and delete the rest. Resist it. The closures that the majority explanation does not cover are the ones carrying information: hosts outside the render farm, closures with no deciding field recorded, closures whose reason text contradicts the structured field. Re-open a sample of those before the packet goes anywhere. A resource hijack on a build server looks, in a summary table, exactly like a rounding error. This is also the practical safety argument for the packet as a whole. If you request a narrowing on evidence that is 94% sound and 6% unexamined, the narrowing is shaped by the sound part and the hole it leaves is shaped by the unexamined part. ## The asymmetry you cannot fix from the queue Even perfectly reliable dispositions can only ever tell you about **alerts that fired**. Nothing in a closure set says what the rule missed. False negatives are not underrepresented in this data — they are absent from it by construction, because a detection's silence produces no record at all. So a feedback loop built on dispositions is a false-positive instrument only, and a packet that implies otherwise is overclaiming. Establishing that a rule still detects the behaviour requires executing that behaviour against it, not reading its output. ## What you say in the packet Be explicit about the quality of your own evidence: sample size re-reviewed, disagreement rate, how many closures carried no deciding field, and how many outliers you re-opened and what they turned out to be. A detection engineer who has been handed unreliable feedback before will look for exactly this, and a packet that volunteers it is the one that gets acted on.
- Can this closure data tell you anything about what the rule is missing?No. The queue contains only what fired, so misses leave no record at all — a rule that detects nothing and a quiet estate produce identical output. Any claim about false negatives has to come from somewhere else: executing the behaviour against the rule and checking whether it fires.
- How do you sample closures for re-review without it landing as an attack on the analysts?Frame and run it as quality assurance of the evidence, not of the person. Sample by rule rather than by analyst, announce it in advance, strip the original verdict so the reviewer works blind, and publish only aggregate disagreement rates. If people believe they are being scored, the next month's closures get written defensively.
- A re-reviewed closure turns out to have been a real detection. What happens next?It stops being a tuning question and becomes an incident: open a case from the original event time, and accept that the dwell clock started when the alert first fired, not when you found it. Then check the neighbouring closures on the same asset group before the packet goes out, because one wrong verdict of that kind rarely stands alone.
saying these in an interview costs you the question
- Treats a false-positive verdict as verified ground truth
- Drops outlier closures as statistical noise
- Claims the closure set measures the rule's false negatives
- Ignores that a twenty-second close checked nothing
- Re-reviews closures as an audit of named analysts