skip to content

Alert Triage & Investigation

The interval between a detection landing in a queue and someone concluding it is adversarial: judging one alert, widening it by pivots, and closing it with a verdict someone else acts on.

on this pageshow

explore

questions

page 2 of 2

An unopened cloud detection is 94 days old and your control-plane logs retain 90. What now?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The item is permanently unanswerable: the telemetry needed for a verdict expired while it waited. Close it as undetermined, never as benign, and pivot to what survives — current resource state and any longer-retention source.

open as a page

Your out-of-hours provider closed a 4769 RC4 service-ticket burst unaided. What authority should it have had?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Not the authority to close it. A provider holding only identity telemetry cannot tell an authorised assessment from an intruder, so its contract should grant a short list of unaided actions and hold-and-hand-over for every context-dependent verdict.

open as a page

A Sysmon Event ID 1 record names a parent that had already exited — how do you rebuild that lineage?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Join upward on the parent process GUID, never on the parent process id, because Windows reuses ids and a dead parent's number can be occupied by something unrelated. If the parent's own creation record is missing, the grandparent is unknowable and you say so.

open as a page

Four hours into an unexplained SSH key case with no second host and no execution — how do you close it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Close it inconclusive rather than leaving it open. Record the population, accounts and per-source window you actually covered, the question still unanswered, the residual risk in plain words, an owner who accepts it, a dated re-check and a trigger that reopens the case.

open as a page

Scoping an unexplained authorized_keys entry, you find a config-management run deployed it — why not close as benign?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the run explains the mechanism, not the authorisation. The next question is which commit put the key in the configuration repository, who wrote it and who reviewed it. That answer also widens the scope to every host in the group.

open as a page

Your SOC closes 400 alerts a day but its case notes are unusable a year later - how do you fix that without halving throughput?

level: principalimportance: should knowfreq 36%

basics

~20 s

Tier note depth by disposition, capture from the tooling what analysts should not retype, mandate a short skeleton rather than a long template, and sample closed cases to test whether a second analyst reaches the same verdict from the note alone.

open as a page

An employee whose authorised transfer you flagged asks what your closed case says about them. What do you answer?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A detection fired on activity they performed, the activity was verified as authorised, and the case closed as a benign true positive. That record is a statement about the rule, not a finding against the person.

open as a page

Your tuning evidence is 612 analyst verdicts — how do you check the verdicts themselves are right?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A disposition records one analyst's judgement, not ground truth. Blind re-review a sample, look at the time-to-close distribution, re-check the recorded deciding value against the raw event, and explain the closures that do not fit the majority cluster before requesting any change.

open as a page

The CI runner pod behind your alert is already deleted — how do you keep pivoting?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Pivot on entities that outlive the workload: the identity it ran as, the image digest, the pipeline job id, the node. The pod's filesystem and memory are gone, and that is an evidence gap to state, not a clean result.

open as a page

Desktop engineering says the RMM chain on the alerted host 'is us' — is that enough to close the case?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Not by itself. The owner's word establishes that such automation exists, not that this particular execution was theirs. Ask for something checkable — a job record with a run identifier and window, or the script that matches the decoded argument — then close.

open as a page

Your CMDB's owner field is wrong for a third of production systems - how do you fix triage?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

You cannot fix a system you do not own. Price the triage cost in numbers the owning teams answer for, and stop relying on one declared field by deriving ownership from deploy history, on-call rotas and code owners.

open as a page

Leadership wants to retire the phish-report button because 97% of reports are benign - what do you argue?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Hit rate is the wrong measure. The button's value is the campaigns where a human report was the first signal because the gateway had already delivered the mail, and how fast it arrived. Attack the queue's cost, not the sensor.

open as a page

An auditor asks what your 9,000 unopened medium-severity detections since March contained. What do you say?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Never assert an unworked backlog was benign. Characterise the population, work a random sample to full verdicts, report the estimated true-positive rate with its uncertainty, and state how many items can no longer be adjudicated at all.

open as a page

Would you run a tierless SOC at a 200-person firm with one in-house analyst and an overnight provider?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

With one in-house analyst there is no internal tier to escalate to, so the SOC is tierless by construction. The boundary worth writing is the organisational one with the provider, backed by an incident-response retainer for the depth you cannot staff.

open as a page

showing 31–44 of 44