Auto-isolation cut off the CFO's laptop mid-board-meeting and the CIO now wants it disabled estate-wide. What do you argue?
answer
- not on-or-off, but scope
- bring the exchange rate, not the excuse
- the business owns availability risk
- benign true positive, not a bug to fix
- publish who reverses, and how fast
basics
~20 sRefuse the on-or-off framing. Make the exchange rate visible - what automation contained last quarter and how fast, against the cost of this one action - then propose scope: softer actions on executive and production hosts, a named reversal owner.
solid answer
~50 sThe decision belongs to the business, because the business owns the availability risk; your job is to make the trade legible rather than to defend the tool. Bring numbers you should already have: how many automated actions fired last quarter, how many were true positives, the median minutes from verdict to containment against your human median including out-of-hours, and the honest cost of this one outage. Then replace the binary with a graded proposal - keep automated isolation on the commodity fleet, downgrade to non-destructive actions on executive and production hosts, set the acting threshold above the alerting threshold, name who can reverse an action and in how many minutes, and notify the person whose machine went dark at the moment it happens so a cut-off is explained rather than mysterious. Be equally honest about what switching it off costs: containment falls back to human speed, including at 03:00.
go deeper
Understand that an automated containment has a business cost as well as a security benefit, and that a rule can be right about the behaviour while the behaviour was perfectly authorised.
Be able to explain the graded alternatives to switching automation off: scoping by host class, separating the acting threshold from the alerting threshold, and substituting non-destructive actions.
Show you can run the aftermath: establish the verdict, reverse on a stated criterion rather than on pressure, restore quickly, and adjust the action rather than deleting the detection.
Own the negotiation. Present the exchange rate honestly, put the decision with the business that carries the availability risk, and get a named owner and a reversal promise on the record before the next one happens.
### Read the room correctly The request is not really "prove the tool works". A senior person has just watched a machine remove a senior executive from a board meeting with no human involved, and they are asking who is in charge. The two losing responses are equally common: defending the rule as technically correct and asking for patience, or agreeing to switch everything off to end the conversation. The first tells the business that security does not price its own actions; the second discards a control on a sample size of one. ### Whose decision it is The availability risk belongs to the business, so the decision to expose a class of hosts to machine-speed containment belongs to the business too. That is not an abdication - it changes your job from advocate to honest broker. You supply the exchange rate; they choose the trade; the choice is recorded with a named owner. Recording the owner matters most on the day something like this happens, because there is then a person who chose the policy, rather than only a machine that executed it. ### The numbers to arrive with If you cannot produce these on the day, that is itself the finding. - Automated actions taken last quarter, split by host class. - How many were true positives, how many false positives, and how many were **benign true positives** - the rule was right about the behaviour and the behaviour was authorised, which is what happened here when an administrator's maintenance script matched. - Median and worst-case minutes from verdict to containment under automation, against the human median, measured separately for business hours and for nights. - The cost of this one action: minutes of executive time, meeting impact, the reversal effort. - How many of those automated containments a human would have got to at all before the damage was done. The argument is an exchange rate, not a virtue: *we buy N minutes off containment on M real intrusions and pay with K disruptive actions a quarter, one of which was very expensive.* Reasonable people can decide that trade is bad for one class of host and good for another. That is the outcome you want. ### The graded proposal - **Scope by host class, not estate-wide.** Executive and VIP laptops, single-point-of-failure production hosts and the tier-0 estate move to non-destructive actions - kill the process, block the hash, force re-authentication, capture memory - plus an immediate page. The commodity fleet keeps isolation. - **Separate the alerting threshold from the acting threshold.** A verdict good enough to raise an alert is not automatically good enough to take a machine off the network, and having two numbers makes that an explicit decision rather than an accident. - **Publish a reversal promise.** Who can reverse an automated action, on what criterion, within how many minutes, at 03:00 as well as at 11:00. A reversal made because an executive is angry is not a criterion; "the process tree matches the administrator's own account of what he ran, and he confirms it" is. - **Tell the human at the moment of the action.** A message on the device and to the person's manager, naming the action, the reason and who to call, converts a mysterious dead laptop into a known, bounded event. Most of the political damage of this incident is the surprise, not the minutes. - **Review every automated action monthly**, with the owning teams in the room, and report the disruption rate as a first-class metric next to the containment rate. ### Two things not to say Do not promise it will never happen again - the same rule matching an authorised administrator's living-off-the-land maintenance script is a benign true positive, and benign true positives are a permanent property of behavioural detection, not a bug to be finally fixed. And do not blame the machine. "The system did it" is the answer that guarantees the capability is removed; "the policy was approved by a named owner, the action was taken under it, and here is the log with the rule, the confidence and the time" is the answer that keeps the conversation about scope. ### The uncomfortable discovery to raise anyway Once this incident is over, look at who is now on the exemption list. Every painful automated action generates pressure to add hosts, and an exemption list grown by political pressure gradually becomes a list of exactly where your response is human-speed - which is useful to more people than you. That belongs in the same conversation, while the organisation is still paying attention.
- The CFO asks who authorised a machine to cut off her laptop. What is your answer?Name the human. The policy that allows automated isolation on that host class was approved by a named owner on a date; the action was executed under it by automation; and the action log records the rule, the confidence and the time. That answer keeps the conversation on whether the policy is right for her host class. Saying "the system did it" invites the capability to be removed outright.
- What do you actually do in the first hour after that isolation?Establish whether the verdict was right about the behaviour before deciding anything - here it was, and the behaviour was an authorised administrator's script, which makes it a benign true positive. Get the administrator's own account, match it against the process tree, then reverse on that stated criterion rather than on pressure. Restore, confirm the host is working, and keep the rule alerting while you strip its automated action for that host class.
- The platform team responds by asking for a blanket exemption for all production hosts. Why is that the wrong shape?It is most of the estate and the most valuable part of it, so it removes the capability by another name while pretending to scope it. Scope by action instead: production hosts keep automated non-destructive responses and an immediate page, and lose only the network isolation. Exemptions should be granted per class with an owner and a review date, not as a single blanket.
saying these in an interview costs you the question
- Defends the rule as correct and asks the business for patience
- Agrees to switch the capability off estate-wide to settle the argument
- Arrives with no data on what automation has actually contained
- Blames the tool instead of naming the policy owner
- Promises the rule will never misfire again
- Treats a benign true positive as a detection failure to be fixed