skip to content

An incident is active and users are affected. You have several generic levers available: roll back the last release, fail over to another region or replica, flip a feature kill switch, shed load, or scale up. How do you choose between them in the first few minutes?

level: seniorimportance: must knowfreq 65%

answer

  1. what changed in the last hour
  2. fastest, narrowest, most reversible
  3. the failover target needs headroom
  4. scaling up is the slowest lever
  5. one change at a time

basics

~20 s

Match the lever to the most likely change vector — code, config, a dependency, traffic, or capacity — then prefer whichever is fastest to take effect, smallest in blast radius, and easiest to undo. Apply one lever at a time so the user-facing signal tells you which one worked.

solid answer

~60 s

Two questions drive the choice. First, **what changed?** A recent release points at rollback; one bad feature points at its kill switch; a single region or replica degrading points at failover; demand exceeding capacity points at scaling up or shedding load. Second, among the levers that fit, **which is fastest, narrowest and most reversible?** A kill switch takes effect in seconds and touches one feature. A rollback takes as long as your deploy pipeline. Scaling up is the slowest and only helps if saturation is genuinely the problem — new instances still have to boot, warm caches and fill connection pools. The constraint people forget is headroom. If two regions each run at 60%, failing all traffic to one puts roughly 120% of capacity on it and you have converted a partial outage into a total one. Failover is only a mitigation if the target has real N+1 headroom. I apply one lever at a time, announce each with a timestamp, and watch the user-facing SLI between them — otherwise I will never know which action recovered the service.

go deeper

for a junior

Know the levers exist and roughly what each does: revert the release, move traffic elsewhere, turn a feature off, drop some traffic deliberately, add capacity. Know that undoing the most recent change is usually the first thing to try.

for a middle

Explain the selection logic: identify the change vector, then prefer the fastest, narrowest and most reversible lever that fits it. Be able to say why a kill switch beats a rollback when one feature is implicated.

for a senior

Demonstrate the constraints that only show up in production — failover needs real headroom, scaling up can worsen a shared-dependency bottleneck, shedding load is a legitimate deliberate choice — and insist on one lever at a time with timestamped announcements.

for a principal

Own the fact that lever choice is bounded by what was built beforehand. Decide which levers each service must have, what latency each is allowed to have, how much headroom failover requires, and who is authorised to pull them without asking.

## Step one: name the change vector Incidents are not random. Something changed, and the class of change narrows the lever set immediately. The five that cover most incidents: - **Code** — a release went out. Lever: rollback. - **Configuration or a flag** — a setting, a limit, a routing rule, a feature enablement. Lever: revert that specific change, or flip its kill switch. - **A dependency** — a database, a downstream service, a third party, a zone or region. Levers: failover, or degrade the path that uses it. - **Traffic** — a spike, a bot, a retry storm, a hot tenant. Levers: shed load, throttle the source. - **Capacity or resource exhaustion** — disk filling, connection pool saturated, a leak. Levers: scale up, restart, free the resource. You are not proving the vector, you are ranking hypotheses in about ninety seconds. "What changed in the last hour?" answers this faster than any dashboard, which is why a change log covering deploys, config pushes and flag flips is the highest-value artifact in an incident. ## Step two: rank the fitting levers by speed, blast radius and reversibility Among the levers consistent with your leading hypothesis, prefer the one that is **fastest to take effect**, **narrowest in what it touches**, and **easiest to undo if you are wrong**. Rough profiles: **Feature kill switch** — usually the best lever when it applies. Effect in seconds, scoped to one feature, trivially reversible, and no deployment involved. It is strictly better than a rollback when a single new feature is implicated, because it removes the suspect behaviour while leaving every unrelated change from the same release in place. **Rollback** — the default when a release is implicated and no narrower lever exists. Its latency is your pipeline's latency: if reverting means a full build-and-test cycle, it is not a mitigation, it is a plan. Blast radius is the whole release, including good changes riding along with the bad one. **Failover** — powerful for a localised dependency or infrastructure failure, and the lever with the sharpest hidden constraint. Two constraints matter. **Headroom**: two regions at 60% utilisation each means the survivor sees roughly 120% of its capacity, so the failover causes the total outage the partial one had not. Real failover capacity means N+1 — the estate keeps serving after losing one unit — and expensive estates plan N+2 so they survive a failure during maintenance. **Statefulness**: promoting a replica may cost you writes, or risk split-brain if the old primary is not truly gone. **Shed load** — the right lever when demand genuinely exceeds capacity and you cannot add capacity in the time available. You are choosing to fail some requests deliberately so the rest succeed, which beats the alternative where overload makes everything fail. It is fast, and it is the lever candidates most often forget exists. **Scale up** — intuitive and usually the wrong first move. It only helps if saturation is the actual cause, it is the slowest lever because instances must boot, warm caches and establish connections, and if the bottleneck is a shared database or a downstream service, adding replicas makes the situation *worse* by adding load to the thing that is already drowning. ## Step three: one at a time, and say it out loud When three levers are plausible the temptation is to pull all three. Resist it, for a reason that is operational rather than procedural: if you roll back, scale up and shed load simultaneously and recovery follows, you have learned nothing about which one worked, so you cannot decide what to keep, what to revert, or what to write in the postmortem. Change one thing, watch the user-facing signal for long enough that the metric's own aggregation window has actually turned over, then decide. Announce each action in the incident channel with a timestamp before taking it. This prevents the genuinely dangerous failure mode where two responders mitigate in opposite directions — one scaling up while another sheds load — and it builds the timeline for free. ## When nothing fits Sometimes no hypothesis is strong after two minutes. Two moves are still available. **Widen the search cheaply**: check the change log, check whether one region, shard or tenant differs from the others, check whether throughput moved before the errors did. And **apply the broadest safe mitigation anyway** — shedding a fraction of traffic or disabling expensive non-essential features buys the system slack and buys you time, without requiring you to be right about the cause. A partial, reversible mitigation applied on a weak hypothesis is almost always better than a perfect one applied ten minutes later.

  • Why is scaling up so often the wrong first lever?
    Because it is the slowest to take effect and only helps if saturation is genuinely the cause. New instances must boot, warm caches and fill connection pools before they carry load. Worse, if the real bottleneck is a shared database or a downstream service, extra replicas add load to the thing already failing and deepen the incident.
  • You have a kill switch for the suspect feature and a rollback available. Which do you use?
    The kill switch, almost always. It takes effect in seconds rather than a pipeline cycle, touches only the suspect feature instead of every change in the release, and is trivially reversible if the hypothesis is wrong. The rollback stays as the next step if disabling the feature does not recover the SLI.
  • Your two regions each run at about 60% of capacity. What does that mean for regional failover as a mitigation?
    It means failover is not currently a mitigation. Moving all traffic to one region puts roughly 120% of its capacity on it, so the survivor saturates and you turn a partial outage into a total one. Real failover requires N+1 headroom — enough spare that losing a unit still leaves the estate serving — which has to be provisioned and tested before the incident, not discovered during it.
  • Why not pull three plausible levers at once to recover faster?
    Because you lose attribution. If recovery follows three simultaneous changes you cannot tell which one worked, so you cannot decide what to keep, what to revert, or what the postmortem should say. Simultaneous levers can also fight each other — one responder scaling up while another sheds load. Change one thing, watch the user-facing signal past one aggregation window, then decide.

saying these in an interview costs you the question

  • Reaches for scaling up regardless of the symptom
  • Fails over without checking the target's headroom
  • Pulls every available lever simultaneously
  • Never considers shedding load as an option
  • Waits for certainty about the cause before acting

context