On a 300-server Linux fleet with confirmed persistence on nine hosts, which do you rebuild rather than clean?
answer
- privilege reached decides it
- can you account for every action
- the host grading its own homework
- what does the rebuild pull from
- rebuild does not revoke anything
basics
~20 sRebuild wherever the adversary reached root, where the package manager or kernel could have been touched, or where telemetry cannot account for what they did. Clean in place only where the mechanism is fully enumerated and the host's activity is fully observed — and check what the rebuild pulls from first.
solid answer
~50 sTier the fleet by evidence and privilege, not by which hosts alerted. Rebuild any host where the intruder had root or kernel-level access, where the package database, boot chain or the agent you would trust could have been altered, or where you have no telemetry covering the dwell window — on those, a clean result is the host grading its own homework. Cleaning in place is defensible only where the mechanism is fully enumerated, the privilege reached was bounded, and the records let you account for what ran. Before ordering rebuilds, verify what the rebuild source is: if the persistence was committed into the configuration-management repository, the base image or a package mirror, a rebuild reinstalls it faithfully at scale. And be clear about what a rebuild does not do — it does not invalidate the credentials, keys or tokens taken off the host, and it does not close the entry vector.
go deeper
Know the basic split: a host where an intruder had root cannot be argued clean from its own output, so it is rebuilt, while a bounded, fully-observed compromise can sometimes be cleaned.
Explain why root access undermines the host's own evidence — agent, package database, file metadata and local logs all sit within the adversary's reach — and what independent evidence would strengthen a clean verdict.
Produce the tiering rule under real constraints, including the hosts with no evidence either way, and name the trap of rebuilding from a configuration repository or base image the adversary could have modified.
Own the negotiation: turn an unbounded rebuild demand into tiers with owners, costs and stated residual risk, and be honest that a third of the fleet lacking telemetry is what made this expensive.
## Why 'rebuild everything' and 'clean everything' are both wrong answers A fleet-wide rebuild is the safest technical position and it is often unaffordable in the window you have; cleaning in place is affordable and, on a host where the adversary had root, close to unfalsifiable. The interviewer is looking for a tiering rule you can defend to the platform owner who has to execute it, not a slogan. ## The rule: rebuild is forced by what you cannot account for Three conditions each force a rebuild independently. **1. Privilege reached.** If the intruder held root — or you cannot prove they did not — then everything on that host that would attest to its own cleanliness is within their reach: the configuration agent, the package database, file metadata, the local logs, the boot chain. A clean sweep result from that host is the host describing itself. Cleaning is only defensible when the access was bounded to an unprivileged service account and you can show it never escalated. **2. Accountability of actions.** Can you enumerate what ran during the dwell window? If the host had process telemetry, shell history you can trust, and audit records covering execution, you can reason about what was done and confirm that the persistence you removed was the whole of it. If a third of the fleet has no agent and audit was never configured for execution, then for those hosts your list of what the adversary did is a list of what you happened to see — which is not the same list. **3. Cost of being wrong versus cost of the rebuild.** On a configuration-managed, stateless fleet, rebuilding is a pipeline run rather than an artisan effort, and it is frequently cheaper than the analyst-days it would take to argue a host clean. The picture inverts for a stateful host — a database primary, a machine holding data nobody else has — where the rebuild is a project, and where the correct move may be to accept a narrower cleaning with heavier monitoring. ## The trap: what does the rebuild pull from? This is the point that separates a good answer from a rehearsed one. A rebuild is only as trustworthy as its sources: the configuration-management repository and its catalogue, the base image, the internal package mirror, and the credentials the build process injects. An intruder who committed a change into the configuration repository has achieved persistence that survives an unlimited number of rebuilds, and a fleet-wide rebuild becomes a fleet-wide redeployment of their access performed by you at speed. So before the first host is rebuilt: review repository history over the intrusion window, check who could commit and whether their credentials were exposed, verify the base image and mirror against known-good state, and rotate the build pipeline's own secrets. ## What a rebuild does not fix Say this explicitly, because it is where over-confident answers fail: - **Credentials.** Keys, tokens and passwords read off the host stay valid after it is rebuilt. Invalidating them is separate work and it is not optional. - **The entry vector.** If the exposed service or stolen key that admitted them is unchanged, the rebuilt host is a fresh target with the same door. - **Their footholds elsewhere.** A rebuild is a statement about one machine, and eradication is a statement about an estate. - **Your evidence.** The rebuild destroys the host's state. Whatever you will need later — to scope the rest of the fleet, or to defend the claim afterwards — has to be collected before the machine goes. ## Handling the hosts with no evidence either way The hard tier is the one that neither alerted nor produced a usable sweep result: no agent, no drift report, no audit records. They are not clean; they are unknown. Sort them by exposure — reachable from the compromised segment, sharing the credential that was stolen, holding the data that matters — and rebuild the reachable and stateless ones because it is cheap. The remainder get compensating controls and a named watch rather than an assumption. ## Making the argument to the platform owner The platform owner is right that rebuilding 300 servers on a maybe is not free, and the useful reply is not louder insistence. It is a tiered list with a reason per tier: these nine are confirmed and non-negotiable; these are stateless and config-managed, so rebuilding is one pipeline run and no argument is needed; these are stateful, so we propose cleaning plus specific monitoring; these have no evidence and here is what we need — telemetry, a maintenance window, or an explicit acceptance that they stay as they are. That turns an unbounded demand into a set of decisions with owners, which is the form in which it actually gets done.
- The platform owner says a fleet rebuild means a week of downtime. What is your counter?That the demand is tiered, not fleet-wide. The nine confirmed hosts are non-negotiable. Stateless, configuration-managed hosts rebuild through the existing pipeline rather than by hand, so the cost is far lower than the estimate implies. Stateful hosts get cleaning plus named monitoring. And I would put the alternative cost in the same sentence: an un-provable host is how the same intruder comes back, at which point we do this again with a longer outage.
- What can a rebuild not remove?Credentials, keys and tokens taken off the host, which remain valid regardless of the disk being wiped; the initial access vector, which will admit them to the new host just as it did the old one; footholds on other machines; and any persistence sitting in what the rebuild pulls from — a poisoned catalogue, base image or package mirror gets faithfully reinstalled. It also removes your evidence, so collect what you need first.
- When is cleaning in place genuinely the safer choice?When the rebuild source is not yet trusted — rebuilding from a repository the adversary may have committed to is worse than leaving the host alone under watch. Also when the host carries state or evidence you cannot afford to lose, and when the access was demonstrably unprivileged and fully observed, so the mechanism you removed can be shown to be the whole of it.
saying these in an interview costs you the question
- Rebuilds without checking whether the rebuild source was touched
- Cleans a root-compromised host because the file was deleted
- Decides by alert count rather than by privilege and evidence
- Assumes a rebuild invalidates the credentials taken from the host
- Counts hosts with no telemetry as clean because nothing was found