skip to content

Four hours into an unexplained SSH key case with no second host and no execution — how do you close it?

level: seniorimportance: should knowfreq 40%

answer

  1. a time box is a budget, not a verdict
  2. decide the exit before you start widening
  3. inconclusive is not false positive
  4. residual risk needs words and an owner
  5. a re-check date and a reopen trigger

basics

~20 s

Close it inconclusive rather than leaving it open. Record the population, accounts and per-source window you actually covered, the question still unanswered, the residual risk in plain words, an owner who accepts it, a dated re-check and a trigger that reopens the case.

solid answer

~50 s

A time box is a decision about the value of the next hour against the rest of the queue — it is not a verdict that the activity was benign, and I say so explicitly. Ideally the exit conditions were written before the widening started, so the four-hour mark is a pre-committed decision rather than fatigue. At the close I record: what was covered (hosts searched, accounts checked, window per source), the specific question left unanswered, and the residual risk in words a system owner can act on — "a key we cannot attribute grants shell as this service account". Then a label of **inconclusive**, not false positive; an owner for the residual risk; a re-check date; and a reopen trigger, such as any further write to that path or an authentication accepted for that fingerprint. A closed case with a date beats an open case nobody returns to.

go deeper

for a junior

Know that 'inconclusive' is a legitimate outcome with its own label, and that closing means writing down what you covered and what you could not — not deciding it was probably fine.

for a middle

Be able to distinguish false positive, benign true positive and inconclusive, and to list what a close record must carry: scope covered, unanswered question, residual risk, owner, re-check date.

for a senior

Show that you set exit conditions before widening and that your close is a resource decision you can defend, with a reopen trigger that the estate raises rather than a person remembering.

for a principal

Own the policy question: how much analyst time an ambiguous artefact is worth against the rest of the queue, who is allowed to accept residual risk, and how the SOC keeps inconclusive closes from becoming a way to make work disappear.

## Why a time box exists at all Analyst hours are the scarce resource in a SOC, and every hour spent widening one ambiguous artefact is an hour the rest of the queue does not get — and those alerts are also real. An investigation with no stopping rule does not end in a verdict; it ends when the analyst runs out of energy or the shift ends, which is the worst of both outcomes because nothing is recorded and nobody owns what is left. So the discipline is: decide *before you widen* what would make you escalate and what would make you stop. Written down, with a rough budget. That pre-commitment does two things — it stops the open-ended dig, and it makes the eventual close defensible to someone reading the case later. ## "Inconclusive" is a real outcome, and it has its own label Three case labels are routinely confused, and they are not interchangeable: - **False positive** — the detection was wrong about the activity being suspicious. - **Benign true positive** — the activity genuinely happened and was authorised. - **Inconclusive** — the evidence available did not decide either way. Stamping an inconclusive case as a false positive is the common and damaging move. It feeds a tuning signal that the rule misfires, and it tells the next analyst — and the next one after a re-alert — that this was already looked at and was nothing. Say what you actually mean: you could not determine it, within the scope and window you covered. ## What the close record has to contain Write it for the person who picks this up in six weeks, not for a metric: 1. **Scope actually covered.** Be specific: which hosts ("the twelve hosts in configuration group `app-prod`"), which accounts, which sources and the window each one could answer. Vagueness here is what makes a later re-alert start from zero. 2. **The unanswered question, stated as a question.** "When was this key added, and by whom?" — not "no evidence of compromise found", which reads as a finding of absence and is not one. 3. **The residual risk, in plain business language.** "An unattributed public key grants shell access as the application service account on twelve production hosts." That sentence is the only part a non-analyst will read. 4. **What was done and by whom.** Any remediation decision belongs to whoever owns that call; the case records it rather than assuming it. 5. **An owner and a re-check date.** Residual risk needs someone who accepts it, and a date is a commitment where "we'll keep an eye on it" is not. 6. **A reopen trigger.** Concretely: any further write to that `authorized_keys` path, the same public key appearing on another host, or an authentication accepted for that key's fingerprint. This converts a manual intention into something the estate will tell you about. ## The two failure modes it prevents **The open case nobody returns to.** It has no owner, no date and no trigger. It sits in "in progress" distorting the queue, and the ambiguity that started it never gets resolved — it just stops being visible. **The reassuring close.** "No evidence of compromise" without stating what was searched and over what period is a sentence that sounds like a conclusion and contains none. If someone quotes it back to you in three months when the same key turns up on another host, you want the record to say exactly which hosts and which window you covered, so the gap is visible rather than implied. ## Handing it over A time-boxed close is also a handoff. The next reader inherits a hypothesis, not just a status: what you believed, what you checked, what you could not check and why. Written that way, a re-alert costs twenty minutes to reassess instead of another four hours to redo. ## The sentence to be able to say "I spent four hours on it, covered twelve hosts, three accounts and every source that reaches back past the key's timestamp, and I still cannot say when it was added or by whom. Here is the residual risk, here is who accepts it, here is the date I look again, and here is what will page us before then." That is a complete answer, and it is a defensible one.

  • Why not simply leave the case open until something else turns up?
    An open case with no owner and no date is not monitoring — nobody is assigned to it, nothing will prompt a revisit, and it quietly distorts the queue. A close with a stated residual risk, an owner, a dated re-check and a reopen trigger is an actual commitment that something will bring it back.
  • What would reopen this particular case?
    Any further write to that authorized_keys path, the same public key appearing on another host, an authentication accepted for that key's fingerprint, or a related commit surfacing in the configuration repository. Each is something the estate can tell you about rather than something a person has to remember.
  • Who accepts the residual risk you have written down?
    The analyst states it; acceptance belongs to whoever owns the affected system, usually with the SOC lead in the loop. Record who accepted it and when. An unowned residual risk is functionally the same as not having recorded it at all.

saying these in an interview costs you the question

  • Closes the case as a false positive because nothing was proven
  • Leaves it open indefinitely with no owner or date
  • Writes 'no evidence of compromise' without saying what was searched
  • Treats running out of hours as evidence the activity was benign
  • Decides the stopping point only once fatigue sets in

context