After forty days of domain-admin access, do you restore the domain controllers or rebuild the forest?
answer
- the directory's contents, not the servers
- can you bound what they changed?
- three options, not two
- restoring also discards the legitimate delta
- issued certificates survive any directory restore
basics
~20 sThe decision is about the directory's contents, not the servers. Restore controllers from a pre-compromise system state when you can bound what changed; rebuild the forest when the privileged object graph can no longer be vouched for.
solid answer
~50 sRebuilding a domain controller is easy; re-establishing trust in what the directory *says* is the hard part. Someone with domain-level privilege for forty days could edit group memberships, permissions on privileged objects, delegation, group policy, certificate templates and account attributes — and edit the records that would show it. So I ask whether I can bound their changes from evidence they could not curate. If I can, in-place remediation is defensible: re-baseline every privileged object against a known-good reference. If I cannot, the options are a forest recovery — restore one writable controller per domain from a system state predating initial access, in isolation, then rebuild the rest by promotion — which discards forty days of legitimate change, or a clean-source rebuild of a new forest with objects migrated in. Either way the certificate authority is a separate question: certificates it issued survive any directory restore.
go deeper
Know that recovering identity means recovering what the directory says about privilege, not just reinstalling the servers, and that a domain-level intruder could change both.
Explain the three options and their costs: in-place re-baselining of privileged objects, restoring controllers from a pre-compromise system state, and building a new forest and migrating into it.
Show the evidence argument — what independent telemetry would let you bound their changes — and handle the certificate authority and the administrative tiering as parts of the recovery rather than afterthoughts.
Be ready to defend a months-long rebuild against a days-long restore to an executive, in terms of what claim the organisation will be able to make afterwards and to whom.
## What you are actually trying to recover The instinct is to think of this as a server-recovery problem: controllers were compromised, restore or rebuild the controllers. That framing misses the point. Domain controllers are replaceable in an afternoon. What you cannot easily replace is confidence in the *contents* of the directory — who is in which privileged group, who holds which permissions on which objects, which accounts exist, what delegation is configured, what group policy applies where, and which certificate templates can issue what. Someone holding domain-level privilege for forty days had write access to all of that, and to the controllers that record changes to it. So the real question is: can you still make a defensible statement about the privileged object graph? ## Three options, not two **Remediate in place.** Enumerate every privileged object and compare it against a known-good reference: membership of the administrative groups, permissions on those groups and on the objects that protect them, delegation and rights on the directory root and on the controllers themselves, group policy objects linked at sensitive scopes, service accounts with elevated rights, accounts with unusual attributes or unexplained creation dates, and certificate templates and their enrolment permissions. Fix every difference. This is by far the cheapest option and it is defensible only if you can bound what they touched from evidence they could not curate — which, given they controlled the controllers, is a strong claim to make. It is most often the right call when the privilege window was short, well-instrumented, and corroborated by telemetry the intruder did not own. **Forest recovery.** Restore a single writable controller per domain from a system state backup predating initial access, in an isolated network; clean up the metadata of every controller you did not restore; then rebuild the remaining controllers by promoting fresh installations rather than by restoring more backups. This returns the directory to a state you can date. Its cost is that it also discards every legitimate change since that date — new employees, new machine accounts, group policy work, application service accounts — and much of that must be reconstructed by hand. It also depends on having a system state backup old enough and reachable enough to be trustworthy, which is exactly the question the backup platform's own trustworthiness decides. **Clean-source rebuild.** Stand up a new forest from installation media, build the tiered administration model properly this time, and migrate accounts, workstations and application trust into it with new credentials. Nothing is inherited, so nothing needs to be vouched for. The cost is measured in months and in every application that has a hard dependency on the old identity domain. Organisations choose it when the old directory's history is genuinely unknowable, or when the pre-incident administrative model was so flat that restoring it would just restore the conditions for the next intrusion. ## The tiered model is part of the recovery, not a follow-up A recurring finding in these incidents is that the estate had no working separation between administering the directory and administering ordinary workstations — the same accounts, or accounts whose credentials landed on the same machines. That is usually *how* forty days of domain privilege happened. Restoring the directory without changing that returns you to the pre-incident state, which is the state that failed. The governing idea is the clean source principle: a system's security depends on the security of everything that controls it, so a rebuilt, trusted directory must not be administered from, or built by, machines you have not also rebuilt. That constrains the recovery order — administrative workstations and the management plane come back before anything they will manage. ## The certificate authority is a separate decision A point that catches people out: restoring or rebuilding the directory does nothing to certificates an enterprise certificate authority already issued. If the CA was in reach, an intruder may hold certificates that authenticate as privileged accounts, or may have altered a template so they can obtain more. Those certificates remain valid until they expire or are revoked, and they do not care that the directory was rolled back. So the CA gets its own assessment: was its key material reachable, what was issued during the window, does the CA come with you or get rebuilt, and what has to be revoked or re-issued. ## How to argue the choice in an interview A strong answer does four things. It states that the object is the directory's contents, not the servers. It names the evidence that would let you bound the changes, and admits that the intruder's privilege undermines that evidence. It prices the options honestly — in-place remediation is hours to days but rests on a claim you may not be able to defend; forest recovery is days and costs the legitimate delta; clean-source rebuild is months and costs migration. And it treats the administrative model and the certificate authority as parts of the decision rather than as separate projects, because both determine whether the recovered directory is actually a clean one.
- What would push you towards in-place remediation rather than any kind of rebuild?Evidence they could not curate. If independent telemetry — endpoint records shipped off-host, identity logs in a separate tenant, network records — lets me bound what the privileged accounts actually touched, and the window is short, then re-baselining every privileged object against a known-good reference is defensible and enormously cheaper. The weaker that corroboration, the less in-place remediation is a claim I can defend later.
- Why does a rebuilt forest not automatically solve a compromised certificate authority?Because certificates already issued are self-contained credentials that remain valid until expiry or revocation, independent of the directory. If the CA key was reachable, or a template was altered so the intruder could enrol for privileged identities, those certificates keep working against services that trust the CA. The CA needs its own decision: assess what was issued, revoke, and rebuild rather than carry it over if its key material was in reach.
- What does the clean source principle mean for how you build the recovered directory?That nothing trusted may be built or administered from something less trusted. A recovered directory installed from an existing imaging server, or administered from an unrebuilt admin workstation, inherits whatever those systems carry. So administrative workstations and the build and management plane are rebuilt from known-good media first, and only then do they touch the new directory.
saying these in an interview costs you the question
- Treats it as a server rebuild rather than a directory-contents problem
- Assumes reimaging the domain controllers makes the directory trustworthy
- Forgets that a forest restore discards legitimate changes since that date
- Ignores certificates already issued by the enterprise CA
- Restores the same flat administrative model that enabled the intrusion