skip to content

The platform owner won't rebuild 291 servers you can't prove are clean — how do you declare eradication complete?

level: principalimportance: should knowfreq 39%

answer

  1. you will not get proof
  2. criteria, not a measurement
  3. no host defaults to clean
  4. residual risk needs a named owner
  5. quiet is not the same as clean

basics

~20 s

Stop claiming proof and state criteria instead: what was swept, what could not be, what compensates for the gap, and which named business owner accepts the remainder. Then run a time-boxed re-entry watch with specific tripwires and conditions that reopen the incident.

solid answer

~50 s

Eradication complete is a decision under uncertainty, not a measurement, so write it as criteria someone signs. Every host in scope ends in exactly one column — rebuilt, cleaned with the mechanism enumerated, or explicitly accepted as unverified — with none silently defaulting to clean. The entry vector is closed and the closure verified independently. The access they held is invalidated. For the accepted column you propose compensating controls: restricting what those hosts can reach, forcing the missing telemetry on, re-sweeping on a schedule, and a funded rebuild in the next maintenance window. The residual risk is stated in plain language and accepted by the owner who controls that estate and its budget, not by the SOC, with a review date. Then a 30-day watch with named tripwires, each owned and tested, and explicit conditions that reopen this incident rather than open a new one. Thirty quiet days is bounded confidence, never proof.

go deeper

for a junior

Know that an incident is not closed because alerts stopped, and that hosts nobody checked are recorded as unchecked rather than counted as clean.

for a middle

Explain why the silence of a detection carries no information about the estate, and describe what a re-entry watch would actually consist of on a Linux fleet.

for a senior

Design the closure: the per-host accounting, verified closure of the entry vector, compensating controls for the unverified hosts, and tripwires that are owned, tested and tied to reopen criteria.

for a principal

Own the risk conversation. Put the uncertainty in business language, place acceptance with the person who controls the budget, attach a review date, and resist letting bounded confidence be reported upward as proof.

## Why this is a decision and not a measurement There is no observation that returns 'the adversary is gone'. Detections produce evidence when they fire; their silence is produced equally by an empty estate, a dormant intruder, a technique nobody modelled, and a sensor that stopped reporting. So the honest form of 'eradication is complete' is never a proof — it is a documented decision, made against criteria, with the remaining uncertainty named and owned by somebody who can act on it. That reframing is also what gets you out of the deadlock with the platform owner. They are refusing to rebuild 291 servers on a maybe, which is a reasonable position to hold about an unbounded demand. What they cannot refuse is being told, precisely, what is unknown and what happens next. ## Criterion one: every host lands in exactly one column The deliverable is an accounting where nothing defaults. Each host in scope is either **rebuilt** from verified sources; **cleaned**, with the mechanism fully enumerated and the privilege reached bounded and evidenced; or **accepted as unverified**, meaning nobody claims it is clean and the decision to leave it in service is recorded as such. The silent fourth category — 'no report, therefore fine' — is the one that turns an incident into a recurrence. If 187 hosts produced no sweep data because their agent had not checked in, those 187 belong in the accepted column with their names attached, not in the swept total. ## Criterion two: the way in is closed and the closure verified Eradication that removes footholds without closing the entry vector produces a clean estate that is re-entered through the same route. 'Verified' means someone other than the person who applied the change confirmed it independently — the exposed service no longer reachable, the vulnerable version actually replaced, the exposed key no longer accepted anywhere. ## Criterion three: compensating controls for the accepted column Accepted risk is only defensible when it is reduced. What you propose for the un-rebuilt hosts is concrete and mostly cheap: constrain what they can reach and what can reach them, so a surviving foothold has a much smaller blast radius and a much noisier path out; force on the telemetry that was missing, which converts the same hosts from unknowable to monitorable within days; re-run the sweep on a schedule rather than once; and put a funded rebuild into the next maintenance window so 'accepted' has an expiry date rather than becoming permanent. ## Criterion four: the acceptance is signed by the right person The SOC does not accept this risk, because the SOC neither owns the systems nor controls the budget that would remove the risk. The accepting party is the executive who owns that estate — and the value of the signature is not blame, it is that a named person now holds a decision with a review date, which is what makes the funded rebuild happen later. Your job is to state the residual risk in language they can weigh without a security background: something like 'on 187 servers we have no record of what ran during the eleven days the intruder held access; we did not find persistence there, and we could not have found it if it were present.' ## The watch: tripwires, not vigilance 'We'll keep an eye out' is not a control. A 30-day watch is a set of specific detections, each with an owner and each tested before you rely on it, aimed at the behaviour and access this actor demonstrated: authentication or connections involving the infrastructure seen in the incident; new keys appearing in `authorized_keys` anywhere in the fleet; new or modified systemd units and package hooks on hosts that should be static; an agent on the accepted hosts going silent again; use of the credentials that were exposed. Each tripwire gets a defined response, and the response is escalation to the incident lead rather than a ticket in a queue. Define in advance what reopens this incident instead of opening a new one — typically any signal tied to the same access, credentials, or infrastructure — because that distinction decides whether the second event is treated as a continuation with all the context, or as a fresh low-severity alert triaged by someone with none of it. ## What thirty quiet days does and does not tell you At the end of the window you know that no tripwire fired. Given that the tripwires were tested and aimed at behaviour this actor actually used, that is real information — it is evidence against the hypothesis of an active foothold using the known techniques. It is not proof of absence: a patient adversary can simply wait out a watch they can see coming, and a tripwire only detects what it was built to detect. Say that out loud when you close the incident, because the alternative is an organisation that believes it has certainty it never bought, and defunds the rebuild on the strength of it.

  • Who signs the residual risk, and why not the SOC?
    The executive who owns that estate and its budget. The SOC produces facts — what was swept, what was not, what we could and could not have detected — and a recommendation. Accepting risk means choosing to run systems in a state you cannot verify, which is a business choice funded by whoever could remove it. A SOC that signs it has quietly absorbed a decision it cannot act on, and the rebuild never gets funded.
  • Thirty quiet days pass. Has the adversary gone?
    It means no tripwire fired. That is evidence against an active foothold using the techniques we modelled, and it is worth something because the tripwires were tested and aimed at this actor's demonstrated behaviour. It is not proof of absence: a patient intruder goes dormant, and any watch only sees what it was built to see. I would report bounded confidence with the basis stated, never 'they are gone'.
  • What would make you reopen this incident rather than open a new one?
    Any signal tied to the same access, credentials or infrastructure — a login from the same source, use of a credential exposed in this intrusion, the same persistence mechanism appearing on a new host. Reopening keeps the timeline, scope and stakeholder context intact; opening a fresh ticket hands the follow-up to an analyst who has none of it and will likely close it as a low-severity oddity.
  • The owner refuses even the compensating controls. What then?
    I document the refusal with the specific risk it leaves and escalate it as an accepted risk to whoever sits above both of us, with a review date. That is the end of my authority and it should be — but an unrecorded refusal is the worst outcome, because the organisation then carries a risk nobody chose. I would also make the smallest independent improvement available to me, such as monitoring those hosts from the network side.

saying these in an interview costs you the question

  • Declares eradication done because nothing has alerted since
  • Lets the security team accept the business's residual risk
  • Claims a clean sweep proves the adversary is gone
  • Leaves a watch with no owner, no tests and no reopen criteria
  • Counts hosts that produced no data among the swept total

context