Auto-close closed 4,000 alerts last quarter and a confirmed intrusion started in that window - how do you find the swallowed alert?
answer
- pivot from what you now know
- replay needs values, not references
- cases outlive the events they point at
- no record has four different meanings
- the timeline probably moves earlier
basics
~20 sQuery the closure store with what the investigation now gives you - hosts, accounts, times, artefacts - not by re-reading 4,000 cases. Whether you can answer at all depends on the closure retaining the original event and the values that matched.
solid answer
~50 sWork backwards from the confirmed intrusion. You now know the hosts, accounts, time range and artefacts, so scope the closure store by those instead of reviewing cases in bulk. For each hit, you need the record to replay the decision: which criterion and version closed it, which field values matched, and a pointer to the retained raw event rather than a case summary. Two things usually break the exercise: the criterion matched against a mutable list, so replaying it today reproduces today's answer rather than the one made then; and the raw events aged out before the investigation started. Also separate the failure stages - telemetry missing, rule never fired, criterion closed it, or nobody acted - because finding no closure proves nothing unless you can show the telemetry existed. The deliverable is a narrowed or retired criterion plus a fixed trail.
code
json · 11 lines{
"case_id": "SOC-2026-114882",
"rule": { "id": "WIN-SCHTASK-CREATE", "version": 7 },
"alert_time": "2026-03-04T02:41:09Z",
"closed_time": "2026-03-04T02:41:11Z",
"closed_by": "auto-close",
"criterion": { "id": "AC-19", "version": 3 },
"matched": { "field": "ParentImage", "allowlist_ref": "ALLOW-PATCHAGENT" },
"raw_event_ref": "sysmon:1:e91f...",
"raw_event_retained_until": "2026-04-03"
}go deeper
Know that an automated closure should leave a record you can search, and that finding an auto-closed alert on a compromised host means the rule fired and a criterion matched - not that anyone looked at it.
Explain the fields a closure record needs to be replayable: rule and criterion identity and version, the field values matched at the time, a pointer to a retained original event, and both timestamps.
Demonstrate the method: pivot from the confirmed intrusion's hosts, accounts and window; distinguish missing telemetry from a missing rule from an automated closure; and expect the exercise to move the intrusion start time earlier.
Own the standard. Decide what evidence an automated closure must carry before automation is allowed at volume, and how retention for closed cases is aligned with the investigations that will need them.
## What is actually being asked This is the question that decides whether your automation is auditable. Any SOC can close alerts automatically; only some can answer, months later, *did the automation close the one that mattered, and why*. The interviewer wants a method, an honest list of what the trail must have kept, and the discipline not to confuse absence of a record with absence of an event. ## Do not read 4,000 cases The investigation has already produced pivots: a set of hosts, one or two accounts, a time window with a defensible start, file hashes or names, destination addresses, a scheduled task or service name. Query the closure store with those. A well-run exercise is a handful of targeted searches, not a bulk review, and it should also cover the period *before* the currently believed start - the whole point of the exercise is that your believed start may be later than the real one because an early alert was closed. ## What a closure record has to contain To reconstruct a decision you need all of: - **The rule identity and version** that fired, so you know what was being detected. - **The criterion identity and version** that closed it. "Closed by automation" is not an answer. - **The values that matched at closure time**, not a reference to a list. If the record says the criterion matched allowlist entry `ALLOW-PATCHAGENT`, and that entry has been edited twice since, you cannot say what it contained then. - **A pointer to the retained original event**, not a rendered summary. Summaries drop the fields you did not know you would need. - **Timestamps for both the alert and the closure**, so you can measure how long the estate had an unread signal. ## The two failures that end the exercise **The mutable lookup.** Closure criteria are usually written against tables - asset owners, exempted accounts, approved paths - that change constantly. Replaying the criterion against today's table tells you what would happen today. If you cannot recover the table as of the closure timestamp (a versioned list, a change log, or the matched values copied into the record), the honest answer is "the decision is not reconstructable", and that is itself a finding. **Retention asymmetry.** Case records are often kept for years while raw telemetry is kept for thirty or ninety days. An intrusion confirmed months later meets a closure record that references an event which no longer exists. The fix is to copy the evidence into the closure, not the reference, for any criterion closing at volume. ## Separate the stages before you conclude A missing closure record has at least four explanations, and they lead to different remediation: 1. **The telemetry never arrived** - the sensor was not installed, the log source had stopped shipping, or the event type was not collected. Blindness, not automation. 2. **The telemetry arrived but no rule matched** - a detection gap. 3. **A rule fired and the criterion closed it** - what you are looking for. 4. **A rule fired, a human closed it** - a triage failure, not an automation one. Proving (1) apart from (2) requires evidence that the records existed for that host in that window; a sensor-health or ingest-volume record is what carries that. Without it you cannot claim the automation is innocent. ## What the closure proves, and what it does not If you find the alert: the rule fired and the criterion matched. It does **not** follow that the criterion was unreasonable - a well-built criterion can be satisfied by an adversary who shaped activity to fit, which is a different lesson from a sloppy predicate. Read the matched values and decide which it was. If you do not find it, you have proved nothing about the automation until you have shown the alert was ever generated. ## The output The exercise is worth running only if it changes something. Expect three deliverables: a verdict on the specific criterion (narrowed, corroborated, scoped, or retired); a fix to the trail itself so the next reconstruction is a query rather than an archaeology project; and a corrected intrusion start time, because a swallowed alert usually moves the timeline earlier - which changes scope, notification obligations and what else you have to go back and check.
- You find no closure record for the compromised host in that window. What do you conclude?Nothing yet. Absence has four readings: telemetry never arrived, telemetry arrived but no rule matched, a human closed it, or the case store does not cover that period. Prove the records existed for that host in that window - ingest volume or sensor-health evidence - before you clear the automation. Otherwise you are treating a blind spot as an all-clear.
- The criterion matched an allowlist entry that has since been edited. How do you salvage the reconstruction?Recover the entry as of the closure timestamp from the list's version history or change log, and replay against that. If no history exists, say so plainly: the decision is not reconstructable, which is a finding in its own right. Then fix it forward by copying the matched values into each closure record instead of a reference.
- What changes for the wider investigation if you do find a swallowed alert two weeks before the believed start?The intrusion timeline moves earlier, and everything scoped from it moves with it: which accounts and hosts were in play, which backups predate compromise, what data was reachable in the extra fortnight, and any notification clock that runs from the earliest known access. Re-derive scope from the new start rather than patching the old narrative.
saying these in an interview costs you the question
- Proposes reviewing all 4,000 closures by hand
- Treats no closure record as proof the automation was fine
- Replays the criterion against today's allowlist
- Relies on a case summary after the raw event aged out
- Concludes the criterion was wrong without reading the matched values