skip to content

After a confirmed intrusion, why is the most recent good backup usually the wrong restore point?

level: juniorimportance: must knowfreq 70%

answer

  1. chosen against a timeline, not a clock
  2. when they got in, not when it broke
  3. a later copy contains their work
  4. dwell time versus retention window
  5. the initial-access estimate only moves earlier

basics

~20 s

Because a restore point after an intrusion is chosen against the intrusion timeline, not against data freshness. Any copy written after the intruder got in can contain their accounts, implants and configuration changes, so restoring it restores them.

solid answer

~50 s

A backup is picked for two different reasons in the two worlds. Outside an incident you pick the newest one, because the goal is to lose as little data as possible. After an intrusion the goal is to come back without the intruder, so the copy has to predate their initial access, not merely predate the damage you noticed. Everything written after they got in is a faithful snapshot of a machine they controlled: their scheduled task, their service account, their edited configuration, all of it backed up and verified along with the real data. So the first input to the decision is the investigation's earliest confirmed adversary activity, and you take margin below it, because that estimate only ever moves earlier as the investigation continues. Restoring old state also does not undo what they learned or took — credentials they harvested still work, and exfiltrated data is still gone.

go deeper

for a junior

Be ready to say plainly that the restore point is dated against when the intruder got in, not against how much data you would lose, and that a copy taken after that date contains their work too.

for a middle

Explain why the initial-access estimate moves earlier as an investigation proceeds, and why that means taking margin below it rather than restoring exactly at the current best guess.

for a senior

Show how you close the gap between an old restore point and today: carry validated data forward, rebuild executables and configuration from known-good sources, and keep credential resets and entry-path closure as separate workstreams.

for a principal

Own the framing that the acceptable data loss here is set by an adversary's dwell time rather than by any agreed objective, and that the resulting loss is a business risk decision someone accountable must take, not a responder's unilateral call.

## Two different questions that look like one When a disk fails or someone drops a table, "which backup do I restore?" has an obvious answer: the newest one that restores. The whole discipline around backups — frequency, retention, verification — exists to make that newest copy as recent and as reliable as possible. After a confirmed intrusion the question changes shape completely. You are no longer trying to lose as little data as possible; you are trying to come back to a state the intruder is not standing in. Those two goals point in opposite directions along the same axis of time, and recovery goes wrong when a responder answers the second question with the first question's habit. ## What a post-compromise backup actually contains A backup is a faithful copy of the source at the moment it ran. If the source had a scheduled task the intruder created, a service they installed, a local account they added, a modified startup script, or a web shell dropped in a content directory, the backup contains all of that — and it contains it as ordinary, healthy-looking data. The job succeeded. The integrity check passed. The restore test worked. None of those facts say anything about whether the source was clean; they say the copy matches the source. Restoring such a copy is not recovery, it is redeployment of the intrusion. This is why the restore point is dated against the **intrusion timeline** rather than against a recovery point objective. The relevant date is the earliest confirmed adversary activity that the investigation can evidence — the first suspicious authentication, the first execution of a tool they brought, the earliest artefact with their fingerprints on it. ## Take margin, because the date moves The estimate of initial access is provisional and it moves in one direction. Investigations start from the loudest, latest activity and work backwards; as more telemetry is examined, earlier activity keeps appearing. It is common for a first estimate of "about a week" to become "six weeks" after host timelines and identity logs are worked through. It is rare for the estimate to move later. So a restore point chosen exactly at the current best estimate of initial access is a restore point that will probably turn out to be inside the intrusion. Practitioners pick a copy comfortably before it, and they revisit the choice if the timeline moves again before the restore is executed. ## The gap between the restore point and today Choosing an older restore point costs whatever real work happened after it. This is the trade that makes the decision hard and it is a business decision as much as a technical one, but note what it is *not*: it is not an RPO discussion. An RPO is a design-time promise about acceptable data loss from a failure. Here the acceptable loss is dictated by an adversary's dwell time, which nobody agreed to in advance and which can easily exceed the retention window the RPO produced. The practical way out of the gap is to separate **data** from **executables and configuration**. Data can often be carried forward from the later, untrusted copy after validation, or reconstructed from independent sources that the intruder did not control. Binaries, system state, scheduled jobs and configuration should come from known-good sources rather than from the later copy, because that is exactly where their changes live. ## What restoring does not do Getting the direction of the claim right matters here: - Restoring to a point before initial access rolls back their **changes**. It does not roll back their **knowledge**. Credentials, key material and network layout they learned are still theirs, which is why credential resets are a separate workstream from recovery. - Restoring does not undo exfiltration. Data that left is still out. - Restoring an image is not eradication. If the same weakness that let them in is restored too, and it usually is, the path back is restored with it. - A backup that restores cleanly proves the copy matches what was backed up. It proves nothing about the health of the original. ## When there is no clean copy Sometimes the earliest confirmed activity predates every retained copy. Then no restore point is clean, and the honest answer is that you rebuild rather than restore: systems come back from installation media and configuration you can vouch for, and data is migrated forward with validation. That outcome is common when the dwell time is longer than the retention window — which is a property of how the backup policy was designed, and a good thing to notice before an incident rather than during one.

  • What if the earliest confirmed adversary activity predates every backup you still retain?
    Then no retained copy is a clean restore point, and you rebuild instead of restoring: systems come back from installation media and configuration you can vouch for, and data is migrated forward from the untrusted copies after validation. Carry data; do not carry executables, system state or configuration. It is also the moment to notice that the retention window was shorter than the adversary's dwell time.
  • Does restoring to a point before initial access mean the intrusion is over?
    No. It removes their changes, not their knowledge. Credentials and key material they harvested still work, data they exfiltrated is still gone, and the weakness that let them in is restored along with everything else. Recovery runs alongside credential resets, closing the entry path, and watching for re-entry — it does not replace any of them.
  • The backup job reported success every night and the last restore test passed. What does that tell you about the restore point?
    That the copy faithfully matches the source and can be read back. Both facts are about the copy, not about the health of what was copied. A machine under an intruder's control produces successful, verifiable backups of an intruder's work, so job success and restore tests are orthogonal to whether the restore point is inside the intrusion.

Rewinding a security camera to before the burglar entered is not the same as rewinding to before you noticed the mess.

saying these in an interview costs you the question

  • Picks the newest verified backup because it passed integrity checks
  • Chooses the restore point from the system's RPO
  • Assumes the intrusion began when the alert fired
  • Treats a successful restore test as proof the source was clean
  • Thinks restoring old state cancels the exfiltration

context