Which hosts do you exempt from EDR auto-isolation, and what risk are you accepting by exempting them?
answer
- cost of the action versus cost of the alert
- tier-0 identity and the tooling itself
- swap destructive for reversible
- an exemption is a slower promise, not none
- the list is a map of where you are slow
basics
~20 sExempt hosts where the automated cut-off hurts more than the intrusion it might stop: domain controllers, the certificate authority, trading-floor systems, jump hosts, the security tooling itself. The cost is human-speed containment on your most valuable hosts.
solid answer
~50 sThe criterion is not importance, it is asymmetry: exempt a host when the blast radius of the action exceeds the blast radius of the alert being right. In practice that is the tier-0 identity estate (domain controllers, the certificate authority, the identity provider's own infrastructure), safety-critical and revenue-critical systems such as a trading floor or plant control, shared jump hosts, and the response platform's own servers - isolate that last one and you have contained your ability to contain. An exemption is a promise of human-speed response, not of no response, so exempt hosts should still get a non-destructive automated action (kill the process, block the hash, force re-authentication, snapshot memory) plus an immediate page to a named person. Derive the list from an authoritative role source rather than hand-typed hostnames, and treat it as sensitive: it maps exactly where your response is slow.
go deeper
Know that automated isolation is not applied uniformly, and be able to name a couple of obvious exclusions - a domain controller, a plant control system - and say why cutting them off is worse than the alert being right.
Explain the criterion rather than reciting a list: exempt where the action's blast radius exceeds the alert's, and describe the non-destructive actions that replace isolation on those hosts.
Show that you have run this: derived membership from an authoritative role source, alerted on tier-0 assets missing an exemption, and staffed the human path the exemption implicitly promises.
Own the second-order risk. The exemption list concentrates your slowest response on your most valuable hosts and is a target list if it leaks, so argue for compensating telemetry depth and a funded review cycle.
### The criterion The wrong question is "which hosts are important?" - by that test you exempt everything. The right question is a comparison of two blast radii: > Exempt the host when the damage done by the **action**, if the verdict is wrong, exceeds the damage done by the **intrusion** in the minutes it takes a human to decide. That framing does two useful things. It makes exemption a statement about the *action*, not about the host's value, and it makes the exemption conditional on how fast a human really can decide - which forces you to be honest about 03:00 staffing. ### The classes that usually qualify - **The tier-0 identity estate.** Domain controllers, the certificate authority, the identity provider's own infrastructure, and the PAM system. Isolating a domain controller can take authentication down for everything that depends on it, and if the verdict was wrong you have caused the outage the adversary wanted. - **Safety-critical and revenue-critical systems.** Plant and process control, medical devices, a trading floor. Some of these cannot be network-isolated without physical consequences, and some carry a regulatory objection to an unattended machine cutting them off. - **Shared jump hosts and bastions.** Isolating one cuts off every responder using it, including the responders working the case. - **The security tooling itself.** The SIEM, the response platform's own servers, the log collectors. Contain those and you have contained your ability to contain, and blinded the investigation at the same moment. - **Single points of failure with no redundancy.** A build agent in a pool of forty is a fine candidate for automation; the one host running an unclustered legacy service is not. ### What you are accepting Three things, and a good answer names all three. 1. **Your highest-value hosts are exactly the ones a machine will not defend at machine speed.** That is a coherent choice only if you actually staff the human path - a page that reaches a named person, a pre-agreed decision, and out-of-hours coverage. 2. **The exemption list is a map.** It says, precisely, where response is slower. An adversary who learns it learns where to land and where to stage. That makes the list sensitive material, tightly access-controlled, and it argues for compensating depth on exempt hosts: richer telemetry, lower alert thresholds, hunts targeted there. 3. **Exemptions rot.** Host roles drift, machines are rebuilt and renamed, a service moves. A list of hand-typed hostnames silently stops covering the thing it was written for. ### Making it hold up - **Derive membership, do not type it.** Bind the exemption to an authoritative source of role - a tier-0 security group, a CMDB or cloud tag, an asset-inventory class - so a rebuilt domain controller is still exempt on the day it comes back. - **Alert on the gaps in both directions.** A tier-0 asset that is *not* covered by an exemption is a finding; so is an exempt host that no longer has the role that earned the exemption. - **Substitute rather than subtract.** Exempt hosts should not fall out of automation entirely. Replace the destructive action with reversible ones and with a page. - **Review on a cadence, with the owning team present.** Every exemption should have an owner, a reason and a date. An exemption granted once because a team complained, and never revisited, is how the list becomes most of the estate. - **Test it.** The only way to know a host is genuinely exempt is to see the automation decline to act on it - either in a controlled test or in a review of real actions taken over the last quarter. ### The trap The most common failure is treating the exemption list as a safety feature that needs no maintenance and no compensating control, so it becomes both stale and, once it leaks, a target list. Exemption is not an absence of decision; it is a decision to spend human time instead of machine speed, and it has to be funded like one.
- Does an exempt host get no automated response at all?It should not. Keep the automation and change the verb: kill the offending process, block the hash, force re-authentication, capture a memory snapshot, and page a named person with the evidence already collected. Those are reversible or additive actions, so a wrong verdict costs a wasted page rather than an outage, and you keep the speed where it is safe to have it.
- Someone proposes publishing the exemption list internally so teams know where they stand. What is the risk?It is a directory of where your response is human-speed. An adversary who obtains it knows which hosts to land on and stage from, and knows their tooling will not be cut off automatically there. Publish the policy - the classes and the criterion - and keep the membership access-controlled, with compensating telemetry depth on those hosts.
- How do you stop the exemption list going stale?Bind it to an authoritative role source - a tier-0 group, a CMDB class, a cloud tag - rather than hand-typed hostnames, so a rebuilt or renamed critical host keeps its exemption. Then alert in both directions: a tier-0 asset with no exemption, and an exempt host that has lost the role. Give every exemption an owner, a reason and a review date.
saying these in an interview costs you the question
- Exempts any host whose owning team complains loudest
- Treats exempt as meaning no monitoring and no action
- Keeps hardcoded hostnames that rot after a rebuild
- Insists nothing should ever be exempt from automation
- Grants exemptions permanently with no owner or review
- Forgets the security tooling itself can be auto-isolated