Across your organisation, incidents are detected in about three minutes but the median time to mitigate is around forty-five. As the engineering lead, what would you change so responders can stop user impact faster?
answer
- detection is already fast — look elsewhere
- permission is in the critical path
- measured times next to each lever
- an undrilled lever is a hypothesis
- some levers will fire unnecessarily
basics
~20 sAttack the two things that fill those forty-two minutes: decision latency and lever latency. Give each service a short menu of mitigations with measured times, pre-authorise on-call engineers to pull them without approval, and treat a slow rollback or an undrilled failover as a defect to fix.
solid answer
~60 sDetection is already solved, so the forty-two minutes are spent elsewhere, and they split into two buckets I can measure. **Decision latency** — the responder knows what would help but is waiting for permission, for the service owner, or for enough confidence to act. **Lever latency** — the decision is made and the mitigation itself is slow: a rollback that needs a full pipeline run, a failover nobody has exercised, a kill switch that requires a deploy. So: publish a short per-service mitigation menu with *measured* times next to each lever, and put it where the page takes you. Pre-authorise the on-call engineer to pull anything on that menu without approval. Set targets on lever latency itself and treat a violation as a bug. Drill the levers, because one that has never been exercised is not a control. The tradeoff I own explicitly: pre-authorisation means some levers fire unnecessarily. That is the price of speed, and reversible self-inflicted disruption is cheaper than tens of minutes of downtime — provided I measure the rate and the levers really are reversible.
go deeper
Understand that the time between noticing and stopping impact is a measurable number, and that waiting for someone's permission or for a slow deploy is often what fills it.
Be able to split the gap into decision latency and lever latency, and explain why a rollback requiring a full pipeline run is not really a mitigation.
Argue for pre-authorised, drilled levers with measured times, and be specific about what makes a lever real: no deploy required for a kill switch, genuine headroom behind a failover, shedding that already exists as a control.
Own the tradeoff explicitly — you are buying a lower time-to-mitigate with a measured rate of unnecessary, reversible pulls — set which levers are pre-authorised and which are not, and make time-to-mitigate a reported metric distinct from time-to-resolution.
## Decompose the forty-two minutes before changing anything The number to improve is time from detection to mitigation, and it is almost never one thing. Break it, per incident, into rough segments: acknowledgement, orientation (what is broken, what changed), decision (which lever), authorisation (who says yes), execution (the lever runs), and verification. Sample twenty recent incidents and the distribution is usually lopsided — and the two segments that dominate in most organisations are **authorisation** and **execution**, not orientation. That matters, because the intuitive fixes (more dashboards, more alerts, better tooling for diagnosis) target the segment that was already fast. ## Decision and authorisation latency Responders hesitate for identifiable, fixable reasons. **They do not know what levers exist.** The fix is a short, explicit menu per service — roll back, kill switch X, fail over to region B, shed class C traffic, scale group D — living where the page actually takes them, not in a wiki nobody can find at 3am. Five levers with one line each beats a thirty-page document. **They do not know what each lever costs or how long it takes.** Put *measured* numbers next to each entry: rollback 6 minutes; regional failover 90 seconds, requires the other region to be under 50% utilisation; kill switch immediate. Numbers turn an anxious judgement call into a comparison. **They are not allowed to pull it.** This is the big one. If the on-call engineer must find the service owner, or get a manager's approval, or file a change request, you have inserted a human availability problem into the critical path — and at 3am that is measured in tens of minutes. **Pre-authorisation** is the change: the on-call engineer may pull any lever on the menu, at any time, without asking, and will be supported afterwards even if it turns out to have been unnecessary. **They are afraid of being wrong.** Pre-authorisation only works if the organisation genuinely does not punish a good-faith unnecessary rollback. If the first time someone reverts a healthy release they are asked to explain themselves in a meeting, every subsequent responder will wait for certainty, and you are back to forty-five minutes. ## Lever latency is an engineering problem, not a culture problem The other half is that the levers themselves are slow, and this is where a lead has the most leverage because it is buildable. Set explicit targets — for example, every user-facing service must have at least one mitigation that takes effect in under five minutes — and treat a service that cannot meet it as carrying a defect, tracked like any other. Concretely that means: a revert must not require a full build-and-test cycle; a kill switch must take effect without a deployment; failover must have somewhere to fail over *to*, with real headroom (N+1 means the estate still serves after losing a unit; running two regions at 60% each does not qualify); and load shedding must be a control that exists rather than a thing someone would have to write during the incident. And every lever must be **drilled**. A failover path that has never been exercised is not a mitigation, it is a hypothesis — and incidents are where hypotheses go to be disproved expensively. Regular exercises turn the menu's numbers from estimates into measurements, which is also how the menu stays honest as the system changes. ## The tradeoff to state out loud This is the part that distinguishes a lead's answer. Pre-authorising fast levers **will** cause mitigations to fire when they were not needed: releases reverted that were innocent, traffic moved for a blip, features disabled for a false alarm. You are deliberately buying a lower time-to-mitigate with a small rate of self-inflicted, reversible disruption. That trade is only sound under two conditions, and both are yours to guarantee. The levers must be genuinely reversible and low-blast-radius — pre-authorising something destructive or data-lossy is a different decision and should stay behind a human check. And you must **measure the false-pull rate**, so the tradeoff stays visible: a few percent is the system working as designed, while a large fraction means the alerting is too noisy and you are treating an alerting problem with a mitigation lever. ## Measure the right thing afterwards Make **time to mitigate** a tracked, reported number per incident, separate from time to resolution. Teams optimise what is measured, and an organisation that only tracks time-to-resolution implicitly tells responders that understanding the failure is the goal. It is not, during the incident — stopping impact is, and the pursuit of the cause can then take exactly as long as it needs to.
- Which levers would you deliberately keep behind human approval rather than pre-authorising?Anything not cheaply reversible or with a very wide blast radius: actions that lose data, promote a replica in a way that risks split-brain, drop a customer's traffic entirely, or touch a regulated system. Pre-authorisation trades a small rate of unnecessary pulls for speed, and that trade only holds when an unnecessary pull is recoverable.
- How do you keep the mitigation menu's numbers from going stale?By measuring them during drills rather than estimating them. Regular game-day style exercises produce real timings for each lever and expose the ones that quietly broke since the last exercise. Timings taken from real incidents feed the same table. A menu whose numbers are guesses is worse than none, because responders make comparisons with them.
- Your false-pull rate after pre-authorisation reaches thirty percent. What does that tell you?That the problem has moved upstream. At that rate responders are reacting to signals that do not reliably indicate real user impact, so the fix is in alerting quality — symptom-based, user-facing conditions — not in tightening authorisation again. Re-adding approval would restore the forty-five-minute median without addressing why they are being woken for non-incidents.
saying these in an interview costs you the question
- Responds by adding more alerts to detect faster
- Keeps approval in the critical path for reversible levers
- Assumes an untested failover will work when needed
- Measures only time to resolution, not time to mitigate
- Denies that pre-authorisation causes any unnecessary pulls