Alert Triage & Investigation
The interval between a detection landing in a queue and someone concluding it is adversarial: judging one alert, widening it by pivots, and closing it with a verdict someone else acts on.
on this pageshowhide
explore
- First Verdict20 questions
- Malicious or Normal3 questions
- Enrichment and Asset Context4 questions
- Tier Model and Authority4 questions
- Unworked Detection Backlog4 questions
- A User Forwards a Message5 questions
- Widening the Picture12 questions
- Entity Pivots4 questions
- Process Ancestry4 questions
- Scope and Time Box4 questions
- Disposition and Feedback12 questions
- Benign True Positives4 questions
- Case Notes and Handover4 questions
- Verdicts as Tuning Evidence4 questions
questions
page 2 of 2An unopened cloud detection is 94 days old and your control-plane logs retain 90. What now?
basics
~20 sThe item is permanently unanswerable: the telemetry needed for a verdict expired while it waited. Close it as undetermined, never as benign, and pivot to what survives — current resource state and any longer-retention source.
Your out-of-hours provider closed a 4769 RC4 service-ticket burst unaided. What authority should it have had?
basics
~20 sNot the authority to close it. A provider holding only identity telemetry cannot tell an authorised assessment from an intruder, so its contract should grant a short list of unaided actions and hold-and-hand-over for every context-dependent verdict.
A Sysmon Event ID 1 record names a parent that had already exited — how do you rebuild that lineage?
basics
~20 sJoin upward on the parent process GUID, never on the parent process id, because Windows reuses ids and a dead parent's number can be occupied by something unrelated. If the parent's own creation record is missing, the grandparent is unknowable and you say so.
Four hours into an unexplained SSH key case with no second host and no execution — how do you close it?
basics
~20 sClose it inconclusive rather than leaving it open. Record the population, accounts and per-source window you actually covered, the question still unanswered, the residual risk in plain words, an owner who accepts it, a dated re-check and a trigger that reopens the case.
Scoping an unexplained authorized_keys entry, you find a config-management run deployed it — why not close as benign?
basics
~20 sBecause the run explains the mechanism, not the authorisation. The next question is which commit put the key in the configuration repository, who wrote it and who reviewed it. That answer also widens the scope to every host in the group.
Your SOC closes 400 alerts a day but its case notes are unusable a year later - how do you fix that without halving throughput?
basics
~20 sTier note depth by disposition, capture from the tooling what analysts should not retype, mandate a short skeleton rather than a long template, and sample closed cases to test whether a second analyst reaches the same verdict from the note alone.
An employee whose authorised transfer you flagged asks what your closed case says about them. What do you answer?
basics
~20 sA detection fired on activity they performed, the activity was verified as authorised, and the case closed as a benign true positive. That record is a statement about the rule, not a finding against the person.
Your tuning evidence is 612 analyst verdicts — how do you check the verdicts themselves are right?
basics
~20 sA disposition records one analyst's judgement, not ground truth. Blind re-review a sample, look at the time-to-close distribution, re-check the recorded deciding value against the raw event, and explain the closures that do not fit the majority cluster before requesting any change.
The CI runner pod behind your alert is already deleted — how do you keep pivoting?
basics
~20 sPivot on entities that outlive the workload: the identity it ran as, the image digest, the pipeline job id, the node. The pod's filesystem and memory are gone, and that is an evidence gap to state, not a clean result.
Desktop engineering says the RMM chain on the alerted host 'is us' — is that enough to close the case?
basics
~20 sNot by itself. The owner's word establishes that such automation exists, not that this particular execution was theirs. Ask for something checkable — a job record with a run identifier and window, or the script that matches the decoded argument — then close.
Your CMDB's owner field is wrong for a third of production systems - how do you fix triage?
basics
~20 sYou cannot fix a system you do not own. Price the triage cost in numbers the owning teams answer for, and stop relying on one declared field by deriving ownership from deploy history, on-call rotas and code owners.
Leadership wants to retire the phish-report button because 97% of reports are benign - what do you argue?
basics
~20 sHit rate is the wrong measure. The button's value is the campaigns where a human report was the first signal because the gateway had already delivered the mail, and how fast it arrived. Attack the queue's cost, not the sensor.
An auditor asks what your 9,000 unopened medium-severity detections since March contained. What do you say?
basics
~20 sNever assert an unworked backlog was benign. Characterise the population, work a random sample to full verdicts, report the estimated true-positive rate with its uncertainty, and state how many items can no longer be adjudicated at all.
Would you run a tierless SOC at a 200-person firm with one in-house analyst and an overnight provider?
basics
~20 sWith one in-house analyst there is no internal tier to escalate to, so the SOC is tierless by construction. The boundary worth writing is the organisational one with the provider, backed by an incident-response retainer for the depth you cannot staff.
showing 31–44 of 44