skip to content

Your verdict audit finds a third of false-positive closures were authorised activity the rule caught correctly - what has that mislabel already broken?

level: seniorimportance: must knowfreq 58%

answer

  1. the rule was right; the activity was authorised
  2. a false-positive label licenses a suppression
  3. yield numbers inherit the wrong label
  4. telemetry, rule, analyst, response
  5. the dismissed red-team case names the failing stage

basics

~20 s

Those closures assert the rule was wrong when it was right. A false-positive label justifies suppression, so the mislabel licenses tuning that carves out the exact behaviour the detection exists to catch, and every yield figure inherits it.

solid answer

~50 s

The rule fired correctly on activity that turned out to be authorised - a benign true positive, not a false positive. Recording it as a false positive claims the detection was wrong, and four things consume that claim. Tuning: the normal remedy for a false positive is a suppression or a narrowed rule, so you blind yourself to the precise behaviour an intruder would reuse from that same host or account. Metrics: any per-rule yield or precision figure is computed from dispositions, so the rule now looks worse than it is and becomes a retirement candidate. Coverage claims: a technique you report as detected may in fact be suppressed. And purple-team attribution: an in-scope red-team action closed as a false positive is recorded as a detection gap, when telemetry arrived and the rule fired and an analyst dismissed it - a response failure with an entirely different fix.

go deeper

for a junior

Recall that a rule firing on real, authorised activity is not the same as a rule being wrong, and that recording it as wrong is what leads someone to switch the detection off.

for a middle

Explain the chain from a label to a suppression to a blind spot, and name the metrics that are computed from dispositions and therefore inherit the error.

for a senior

Show you would find the whole population rather than fixing thirty tickets, audit every exclusion that grew from those cases, and attribute a dismissed exercise action to the response stage rather than to detection.

for a principal

Own the systemic fix: reference data for authorised administrative activity, a deconfliction path for exercises, and a case system whose disposition options do not force analysts into the wrong label in the first place.

## What was actually recorded An alert on an administrator's credential-audit tooling, or on an in-scope red-team action, is not a false positive. The detection did exactly what it was written to do: the behaviour occurred, the telemetry captured it, the rule matched. What made it not an incident was **authorisation**, which lives outside the detection entirely. Recording that outcome as "false positive" makes a claim about the rule that is simply untrue, and a verdict audit that finds a third of a rule's closures carry that label has found a content defect, not thirty individual mistakes. ## The four downstream consumers of a disposition **1. Tuning.** This is the expensive one. The standard, reasonable response to a false positive is to suppress or narrow the rule - exclude that binary, that account, that host, that path. Applied to a benign true positive, the exclusion removes a genuine detection. If the exclusion is written broadly ("suppress this tool on the admin jump host"), an intruder who reaches that host and runs the same tooling now generates nothing at all, and the exclusion is invisible in the alert stream by definition. Mislabels do not merely distort a number; they buy the adversary a blind spot and pay for it with a ticket nobody re-reads. **2. Rule yield and precision figures.** Every per-rule quality number a SOC reports is derived from analyst dispositions. If a third of one rule's closures are labelled false positive when the rule was right, the rule's apparent precision collapses and it moves to the top of the retirement or rewrite list. The measurement agrees with the error instead of exposing it, which is precisely why the labels themselves have to be audited against raw evidence rather than trusted as ground truth. **3. Coverage claims.** If a suppression grew out of the mislabel, a technique cell that leadership reads as covered may be covered only for cases nobody excluded. The coverage statement and the exclusion list have drifted apart, and nothing reconciles them automatically. **4. Purple-team and exercise attribution.** This is the sharpest case. An in-scope red-team action was executed, the telemetry arrived, the rule fired, an alert reached the queue - and the analyst closed it as a false positive. Reported naively, the exercise records a *detection miss*. It was not. The chain has four stages - telemetry collected, rule fired, analyst acted, response landed - and the failure here is at stage three. Sending that to a detection engineer produces a new rule duplicating one you already have, while the actual fixes are analyst context, a playbook that does not license a close on this signal, and a deconfliction record so the exercise is recognisable. ## What you do about it - **Re-label the sampled cases** with the correct disposition and record why, so the audit trail shows the correction rather than quietly overwriting history. - **Find the population, not just the sample.** Pivot from the sampled cases to their rule, entity and note pattern, and pull every closure that looks like them. A third of a sample being wrong means the same standing assumption is running across the queue. - **Review every suppression that traces back to one of those tickets.** This is the step that actually restores detection. Narrow each exclusion to the specific authorised context - this service account, from this source, on this schedule - rather than the tool or the host. - **Recompute the rule's yield with corrected labels** before making any decision about retiring or rewriting it. - **Return the red-team case to the exercise record with the right stage attribution**, so the remediation goes to the response side. ## Preventing the next one The underlying reason a good analyst mislabels here is that they had no way to establish authorisation quickly. The durable fixes are reference data and process: an inventory of authorised administrative tooling with the accounts and hosts entitled to run it; a change or maintenance record the analyst can query; a deconfliction channel for exercises; and a disposition option in the case system that actually expresses "the rule was right, the activity was authorised", because if the tool offers only true and false positive, analysts will keep choosing the wrong one. One caution in the other direction: an authorised-looking action is not self-certifying. The same command run by an intruder who has taken the administrator's session looks identical in the telemetry. The case is closed correctly only when the *authorisation* was corroborated from a source the adversary does not control - a change record, an approval, the administrator confirming out of band - not merely because the account is one that usually does this.

  • The mislabelled case was an in-scope red-team action. Which stage of the chain actually failed?
    Response. Telemetry arrived, the rule fired and an alert reached the queue; the analyst dismissed it. Reporting it as a detection gap sends work to a detection engineer who will duplicate a rule that already works. The real fixes are analyst context, a playbook that will not license that close, and a deconfliction record.
  • How do you find the rest of the population once the sample shows the pattern?
    Pivot from the sampled cases to their rule, entity and note text, and pull every closure that matches that shape over the retention window. Then check which suppressions or exceptions were created from those tickets, narrow each one to the specific authorised context, and recompute the rule's yield with corrected labels.
  • Why does an authorised administrator's tool run still deserve a case record rather than a suppression?
    Because the same command run by an intruder is indistinguishable in the telemetry; only the corroborated authorisation separates them. The case's lasting value is the reference record - who runs this, from where, on what schedule - which makes the next close fast and keeps any exclusion narrow enough to still catch the adversary.
  • The suppression that came out of the mislabel has been in place for a year. How would you assess the damage?
    Treat it as a blind spot with a known start date: identify what the exclusion covered, hunt over the retained telemetry for that behaviour in that scope, and check whether any incident in the period touched the excluded host or account. Then replace the exclusion with a narrower one and note that its silent period cannot be fully reconstructed beyond retention.

saying these in an interview costs you the question

  • Calling any alert on authorised activity a false positive
  • Suppressing the tool or host after one authorised run
  • Reporting a dismissed red-team action as a detection gap
  • Trusting per-rule precision computed from unaudited dispositions
  • Closing on the account being one that usually does this

context