skip to content

Alert Triage & Investigation

The interval between a detection landing in a queue and someone concluding it is adversarial: judging one alert, widening it by pivots, and closing it with a verdict someone else acts on.

on this pageshow

explore

questions

page 1 of 2

What must a security alert's case note record beyond the verdict, so another analyst can act on it?

level: juniorimportance: must knowfreq 68%

answer

  1. someone else, without you
  2. facts and reasoning, not the verdict
  3. paste the query, not a summary
  4. one clock for every source
  5. zero results are findings too

basics

~20 s

A case note must let someone else reach the same verdict without you: the exact queries run with their time windows, what was found and what was ruled out, every timestamp in UTC, and the reasoning behind the disposition.

solid answer

~50 s

The test I apply is simple: can another analyst, or an auditor eleven months from now with nobody left to ask, reach my verdict from the note alone? That means recording observable facts and reasoning, not just a conclusion. Concretely: the entities involved (host, account, process, source address); the exact queries or tool actions I ran, pasted verbatim, with their data source and time window; what each returned, including the searches that returned nothing; what I ruled out and the evidence that ruled it out; every timestamp normalised to UTC, noting a source's local offset where it mattered; and the verdict with its reason - true positive, false positive, or benign true positive with the authorised activity named. `Checked, looks fine` is not a case note. It records that a human touched the ticket and nothing else.

go deeper

for a junior

Be ready to list what goes in a note: entities, the exact queries with time windows, results including empty ones, what was ruled out, UTC timestamps, and the verdict with its reason.

for a middle

Explain why each element is there - what a later reader cannot do without it, and why a paraphrased query or a local timestamp quietly destroys the note's value.

for a senior

Show the judgment about depth: which cases justify a full narrative, what you preserve versus what you only reference, and how you write so a stranger in another region can act at 22:00.

for a principal

Own the standard itself - what the organisation requires a closed case to carry, how long it is kept, and how that is reconciled with the minutes an analyst actually has per alert.

## What a case note is for A case note is not paperwork attached to an alert. It is the only durable output of triage. The alert console will age out, the search you ran will not be re-runnable at the same retention a year from now, and you will not remember this alert next week - let alone in eleven months when an auditor pulls a sample of closed cases. Everything triage produced that outlives the shift lives in the note. It has exactly two readers, and they want the same thing for different reasons: - **The next analyst.** They pick up your case cold, possibly in another region at 22:00, possibly because the same host alerted again. They need to know what you checked so they do not repeat it, and what you did not check so they know where the gap is. - **The auditor, regulator or investigator, much later.** They are asking why this alert was closed the way it was. They cannot ask you. They can only read. Both reduce to one acceptance test: **can someone else reach the same verdict from the note alone?** If not, the note is incomplete no matter how long it is. ## The contents **Identity and scope of the alert.** Which detection fired, on which entity, at what time, and what the rule was actually looking for. Rules change; a year later `suspicious archive creation` may mean something different, so record the rule name or identifier and the field values that triggered it. **The entities.** Hostname *and* an address or asset identifier, account name in a resolvable form (a SID or object id survives a rename), process name plus full command line, parent process, and any external address or domain. Names drift; identifiers do not. **The exact queries and actions, verbatim.** Paste the query text, not a paraphrase. `I searched the proxy logs` is unverifiable and unrepeatable; the query string, with its index or data source and its explicit time window, is both. The same goes for console actions: which tool, which view, which filter. **What each query returned - including nothing.** Zero-result searches are findings and belong in the note. So do their limits: which hosts were in scope, whether every host in scope actually reports, how far back the data goes. **What was ruled out, and by what evidence.** The conclusion `no lateral movement` is worthless on its own. The query that failed to find it, over which sources and window, is what a reader can evaluate. **Timestamps, in UTC.** A SOC handing cases between regions cannot reason about a timeline where one line is in local time and the next is not. Normalise every timestamp to UTC in the case timeline; where a source recorded local time, say so and say what offset you applied. Note whether a timestamp is event time or ingest time when the two differ enough to matter - the gap between the two is often what makes an ordering claim wrong. **The reasoning and the verdict.** State the inference, not just its result: `4624 logon type 3 to FS-07 came from the backup service account at its scheduled 02:00 window and matches the previous 30 days, so this is authorised activity`. Then the disposition - true positive, false positive, or benign true positive (the rule fired correctly on activity that turned out to be authorised) - with the authorising owner or change named if you have it. **What you preserved.** If you exported logs, saved a process tree or pulled a file, say what you took, when, and where it now lives, so the next person can find it rather than re-collecting it. ## The failure modes - **The conclusion-only note.** `Checked, looks fine.` It cannot be checked, corrected or defended. - **The chat-thread note.** The reasoning lives in a channel, the case says `see chat`. Channels are ordered by message, not by event; they are not retained on the case's schedule; and the participants leave the company. - **Local timestamps.** Two sources, two offsets, one confident ordering that is wrong by an hour. - **Paraphrased queries.** The reader cannot tell whether your search would have found the thing you say it did not find. - **Recording only the supporting evidence.** A note that lists what confirmed the verdict and omits what did not is an argument, not a record. ## What the note is not It is not a formal evidence-custody record, and it is not a postmortem. It is the working record of an investigation, written for a stranger, in the plainest possible terms, with the reasoning exposed so that a stranger can disagree with it.

  • Where should the record live - the case, or the chat thread where the work actually happened?
    The case. Chat is a working surface, not a record: it is ordered by message rather than by event, it is retained on a different schedule, and its participants leave. Link the thread if you like, but the evidence, the queries and the reasoning are copied into the case record so it stands alone.
  • Why record which detection rule fired and what it was looking for?
    Because rules change. A year later, reading only the alert name, nobody can tell whether the rule that fired then is the rule that exists now, or which field values tripped it. Recording the rule identifier, its version if you have one, and the matched values makes the original firing reconstructible.
  • Two sources recorded the same event in different local times. What does the note record?
    Normalise the case timeline to UTC and record, per source, the original timezone and the offset you applied, plus whether the timestamp is event time or ingest time. Otherwise a later reader cannot verify the ordering you built your conclusion on, and an ordering claim is usually the load-bearing part of the note.

saying these in an interview costs you the question

  • Writes only a verdict: checked, looks fine
  • Summarises the query instead of pasting it
  • Records local times with no timezone or offset
  • Leaves the reasoning in a chat thread
  • Omits searches that returned nothing
  • Assumes the next reader can just ask them

context

open as a page

When closing a security alert, what separates a false positive from a benign true positive?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A false positive means the detection was wrong: the behaviour it claimed to see did not occur. A benign true positive means the detection was right and the behaviour was authorised. One indicts the rule; the other clears it.

open as a page

What must a false-positive closure record carry for the detection engineer who owns the rule?

level: juniorimportance: must knowfreq 60%

basics

~20 s

The rule and the version that fired, the exact field and value that made the activity benign, the asset and its asset group, and the analyst's reason. A verdict label on its own is not evidence anyone can act on.

open as a page

What does a CMDB asset criticality and owner lookup tell you about whether an alert is malicious?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Nothing. Criticality and owner tell you what is at stake and who to ask, so they change urgency, routing and the response you may take. Only evidence about the behaviour itself moves a malicious-or-not verdict.

open as a page

A certutil download command alerts on a packaging workstation and turns out to be authorised: false positive or benign true positive?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A benign true positive. The behaviour really happened and the rule matched exactly what it was written to match; it was simply authorised. A false positive is a rule firing on activity that never matched its intent at all.

open as a page

A reported message passes SPF, DKIM and DMARC - what does that prove during triage?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Only that the sending domain authenticated and the signed parts were not altered in transit. It says nothing about intent. An attacker who registers a lookalike domain publishes his own records and passes all three cleanly.

open as a page

In a tiered SOC, which alerts may a tier-1 analyst close unaided and which must go to tier-2?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Tier-1 closes only what a written playbook fully covers: expected activity and known false-positive patterns. Anything the playbook does not cover, anything suggesting the adversary succeeded rather than merely tried, and anything on a critical asset goes to tier-2.

open as a page

Which fields in a Kubernetes audit record of a Secret read are worth pivoting on?

level: juniorimportance: must knowfreq 68%

basics

~20 s

The entity fields: the service-account username, the object's namespace and name, the source IP, and the user agent. Each becomes a search across other sources inside a time window. The audit ID only links stages of that one request.

open as a page

An alert fires on powershell.exe with an encoded command line — what does its parent chain add to the verdict?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The parent chain says who asked for it. Encoded PowerShell under a document application means a user opened something that executed code; the same command under a management agent's service means scheduled automation. The lineage, not the child, carries the verdict.

open as a page

One unexplained entry in a Linux server's ~/.ssh/authorized_keys: what defines the scope of your investigation?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Scope has three axes: how many hosts carry the same artefact, which accounts it grants access to, and over what time window you can see. One host with no accounts and no time bound is a finding, not a scope.

open as a page

How do you read an approved emergency change record as evidence when triaging a 04:00 database bulk export?

level: middleimportance: must knowfreq 62%

basics

~20 s

A change record proves only that a ticket exists with a stated scope, requester and approver. Join it to the activity on four axes - time, scope, identity and approval path - before it counts as corroboration, and never as authorisation.

open as a page

300 cloud detections arrive daily and you can work 40. How do you choose which 40?

level: middleimportance: must knowfreq 72%

basics

~20 s

Work order decides which 260 you never examine, not how many you finish. Order by what a late look costs: evidence about to expire, high blast-radius accounts and identities, and stages of an intrusion you cannot undo.

open as a page

At shift end an unexplained 6 GB password-protected archive is still growing on a file server - what must your handover note carry?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Hand over the hypothesis, not just the facts: what you think is happening and how sure you are, the evidence both ways, what you deliberately did not do and why, and the threshold at which the incoming analyst escalates without waiting for you.

open as a page

A prevalence lookup shows 412 of 6,000 hosts ran the alerted certutil download command in 30 days: what can you conclude?

level: seniorimportance: must knowfreq 61%

basics

~20 s

Only that the behaviour is common in that population. Prevalence measures a base rate, not authorisation. It can lower your suspicion cheaply, but only if those 412 hosts are comparable to the alerting one, and it never proves anything is malicious.

open as a page

A reported phishing message is confirmed malicious - how do you scope who else in the tenant got it?

level: seniorimportance: must knowfreq 63%

basics

~20 s

Search the mail trail on the campaign's stable elements, not the exact subject, over a window that opens before the report. Then split recipients into those who merely received it and the few with evidence of acting.

open as a page

A build alert gives you a token id, an image digest and a source address — which pivot runs first?

level: seniorimportance: must knowfreq 50%

basics

~20 s

The one whose result would change your next decision, weighted by how selective it is. In a build-infrastructure intrusion that is normally the image digest, then the token id; a shared or NAT'd source address answers least and usually runs last.

open as a page

What is the difference between a detection closed as a false positive and one nobody ever opened?

level: juniorimportance: should knowfreq 58%

basics

~20 s

A closed false positive is a verdict somebody reached. An unopened detection is the absence of a verdict: it tells you about the SOC's capacity, not about the estate, and it can never be counted as benign.

open as a page

In an alert case note, why record what you ruled out and the exact query that ruled it out?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because a conclusion is not evidence. A zero-result search only means something if the note records what was searched, over what window and which sources - otherwise a later reader cannot tell a clean host from a blind spot.

open as a page

A sanctioned 40 GB partner transfer fired an exfiltration rule. Why is closing it 'false positive' wrong?

level: middleimportance: should knowfreq 58%

basics

~20 s

The rule was not wrong. It reported the transfer exactly as it happened; only a signed contract makes the transfer acceptable. A false-positive label records a working detection as defective, and defective detections get narrowed or retired.

open as a page

How do you turn 612 closures on one misfiring detection rule into evidence its author can act on?

level: middleimportance: should knowfreq 48%

basics

~20 s

Aggregate the closures into a claim: which field decided them, how often, and on which asset groups. Add a handful of exemplar records, state what the rule did catch that was real, and hand it to a named owner. Supply the evidence, not the rewrite.

open as a page

A SIEM alert arrives labelled 'Critical' by the detection rule's vendor: what does that severity actually describe?

level: middleimportance: should knowfreq 54%

basics

~20 s

It describes the rule's content: how bad that behaviour would be in general, judged by an author who has never seen your estate. This alert's severity also depends on the host, the account and how ordinary the behaviour is here.

open as a page

Thirty-one staff report your own marketing newsletter as phishing - what verdict do you record?

level: middleimportance: should knowfreq 47%

basics

~20 s

A benign true positive, not a false positive: the reporters correctly spotted phishing-shaped mail and the campaign really is ours. Confirm it from the mail trail rather than from marketing's word, close the thirty-one as one case, and answer every reporter.

open as a page

What must a tier-1 escalation to tier-2 carry besides the alert identifier?

level: middleimportance: should knowfreq 58%

basics

~20 s

A stated hypothesis, every lookup already run and what each returned including the negative results, what was deliberately not checked, any action already taken on the estate, and the exact identifiers and time window. Tier-2 inherits an investigation, not a ticket.

open as a page

Two alerts a week apart share the CI service account ci-runner — what does that shared name prove?

level: middleimportance: should knowfreq 55%

basics

~20 s

Only that both events authenticated as the same identity. If every pipeline uses that account, the link is nearly worthless. It becomes evidence when a second, more selective entity such as an image digest is shared too.

open as a page

The same encoded PowerShell runs under WINWORD.EXE on one host and under an RMM agent on 4,000 — how does ancestry decide each verdict?

level: middleimportance: should knowfreq 62%

basics

~20 s

Compare the roots, not the leaves. One chain is rooted at an interactive user's document application; the other at the service control manager starting a management agent as SYSTEM in session 0. Identical children, different causes, opposite verdicts.

open as a page

An authorized_keys entry's mtime is 40 days old but auditd retains 14 days — how do you set your lookback window?

level: middleimportance: should knowfreq 52%

basics

~20 s

Set the window per source, not once for the case. auditd answers only its 14 days; reach further with config-management history, archived sshd authentication logs, bastion logs and host backups, and record whatever stays unreachable as a stated gap.

open as a page

The same authorised partner transfer alerts every month. What do you record instead of an allowlist entry?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Record the authorised activity, not a suppression: who owns it, the exact combination permitted, the evidence that authorises it, and an expiry or review date. The alert keeps firing and keeps being reviewed - it just resolves in minutes.

open as a page

Analysts closed the same nightly cryptomining alert by hand for fourteen months — why did nobody report it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because reporting cost more than closing and produced no visible result. Ninety seconds a night is cheaper than finding an owner two org units away, and analysts are measured on time-to-close, not on rules improved. The real damage is the trained reflex.

open as a page

How do you decide whether to call the DBA at 04:20 to confirm a bulk export you cannot explain?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Ask which hypothesis the call tests. Under a stolen-credential hypothesis, reaching the real person is fast and safe. If that person is the possible subject, the call warns them, and the decision belongs to the incident lead with legal and HR.

open as a page

showing 1–30 of 44