skip to content

Rollback is the usual first mitigation, but sometimes fixing forward is genuinely the less risky choice during a live incident. Give the concrete conditions under which you would fix forward, and how you would bound that decision.

level: seniorimportance: should knowfreq 55%

answer

  1. the old build must still be able to read the data
  2. previous build may carry the same bug
  3. compare times to user recovery
  4. cover users while you write the fix
  5. agree the cutoff before you start

basics

~20 s

Fix forward when rollback is impossible, ineffective, or slower than the fix: an irreversible migration has already run, the previous build carries the same defect, or reverting takes forty minutes while a one-line change takes four. Bound it with a hard deadline and a cruder fallback mitigation.

solid answer

~60 s

Rollback is the default because it restores a build that demonstrably worked. I fix forward only when one of three things is true. **Rollback is impossible or unsafe** — an irreversible schema migration has run and the old code cannot read the new data, or reverting would undo a security fix. **Rollback would not help** — the defect predates the release, or the release merely exposed a latent bug, so the previous build fails the same way. **Rollback is slower than the fix** — reverting a large fleet takes forty minutes while a one-line configuration change takes four. The bound matters as much as the condition. I set an explicit deadline ("if this is not verified in fifteen minutes we take the cruder lever"), I keep a coarse mitigation applied meanwhile — kill switch, shed load, degrade the feature — so users are covered while the fix is written, and I require a second pair of eyes on the change. Fix-forward means untested code written under pressure, and its real risk is a second incident on top of the first.

go deeper

for a junior

Know that rolling back is the normal first choice because the previous build is known to work, and that a hot fix written during an incident is untested code. If asked, name one exception: a migration that the old version cannot read.

for a middle

State the three conditions — rollback impossible, rollback ineffective, rollback slower than the fix — and explain why the comparison is between times to user recovery rather than times to deploy.

for a senior

Show the bounding discipline: an announced deadline owned by someone other than the author, a coarse mitigation applied so users are covered while the fix is written, a minimal reviewed change, and a rule for abandoning after a failed attempt.

for a principal

Own the engineering that keeps the choice available: expand-then-contract migrations, rollback latency treated as a service-level property with a target, and a norm that a hot fix is never the only thing protecting users during an incident.

## Why rollback is the default at all A rollback returns the system to a state that was **empirically serving traffic minutes ago**. That is a far stronger guarantee than anything a fix-forward can offer, because a hot fix is code that has never run in production, written by someone stressed, reviewed under time pressure, and frequently deployed with the usual checks abbreviated. The base rate of hot fixes that introduce a second problem is not small. So the burden of proof sits on fix-forward, and the interview question is really: can you state the conditions that discharge it, rather than defaulting to whichever feels braver? ## The three conditions that justify fixing forward **1. Rollback is impossible or unsafe.** The canonical case is an irreversible data migration that has already run. If the new release migrated the schema or started writing data in a new format, the previous build may be unable to read what is now in the database — reverting the binary converts a partial failure into a total one, potentially with data loss. Similar cases: the release contained a security fix you must not undo; the old build depends on an external contract that has since been retired; the rollback path itself has never been exercised and you have no confidence it works. **2. Rollback would not fix it.** Two shapes here. The defect is **older than the last release** — it has been sitting latent and something else, such as a traffic pattern or a data value, triggered it today; the previous build contains it too. Or the release merely **exposed** the latent bug: it changed a timeout, a query shape or a call ordering that made a long-standing weakness suddenly matter. Rolling back may hide the trigger, which is a valid mitigation, but if the trigger is external — a spike, a bad record, a dependency change — it will fail identically after the revert. The tell is a rollback that completes cleanly with the error rate unmoved. **3. Rollback is slower than the fix.** This is the arithmetic case, and it is the one worth stating with numbers because it is the one interviewers can push on. If reverting a large fleet with cold caches takes forty minutes, and the failure is a single misconfigured limit that a four-minute config push corrects, fixing forward is not the reckless option — it is the option that stops user impact thirty-six minutes sooner. Note that the comparison is between *times to user recovery*, not times to deploy: a rollback that finishes in six minutes but leaves caches cold for twenty is not a six-minute mitigation. ## How to bound the decision Deciding to fix forward is not the end of the decision, because the failure mode of fix-forward is *open-ended optimism*: "nearly there" repeated for ninety minutes while users stay broken. Four bounds keep it honest. **Set a deadline before you start.** "If the fix is not deployed and verified by 14:35, we roll back / we cut the feature / we shed traffic." State it in the channel so it is a commitment, not an intention, and make someone other than the person writing the code own the clock. **Cover users in the meantime.** Fix-forward should almost never be the *only* thing happening. Apply a coarse mitigation immediately — disable the feature, shed the affected traffic class, serve degraded results — so the fix is being written on a system that is no longer failing users. This converts a race into ordinary work and is the single most valuable habit in this whole answer. **Keep the change minimal and reviewed.** The hot fix should do one thing. Resist bundling the cleanup, the logging improvement and the refactor. A second person reads it before it ships, however trivial it looks, because the whole risk of this path is a defect written under pressure. **Do not skip the safety rails that are fast.** Abbreviating a twenty-minute test suite is a defensible incident decision. Skipping the linter, the type check or the smoke test to save ninety seconds is not — those are exactly the checks that catch the typo class of second incident. ## The decision is reversible until it isn't A useful framing: fixing forward is itself a lever you can abandon. If the fix fails verification, do not iterate a second and third time by reflex. Each failed attempt is evidence that the model of the failure is wrong, and the cheap action — rolling back, or applying the cruder mitigation you were holding — becomes more attractive with each one. The deadline exists precisely so that this reassessment happens by agreement rather than by exhaustion.

  • How can a team make rollback viable even when a schema migration has already run?
    By keeping migrations backward-compatible for at least one release: add columns rather than renaming, have the new code write both formats while the old code can still read, and only remove the old path a release later. That expand-then-contract discipline preserves rollback as an option, which is exactly when it is most valuable.
  • You fix forward, the fix does not work, and someone proposes a second attempt. What do you do?
    Treat the failed attempt as evidence the model of the failure is wrong, not as bad luck. Two consecutive failed hot fixes mean the hypothesis is unreliable, so I fall back to the coarse mitigation or the rollback we agreed as the cutoff. Iterating a third time is how a twenty-minute incident becomes a two-hour one.
  • During a live incident, would you deploy the hot fix straight to production or through a canary?
    Through a fast canary if one exists and it costs a minute or two, because the fix is untested code and a canary limits the second incident to a small population. If canary analysis needs a long bake to be meaningful, that time is not available — I would deploy narrowly, to one region or shard first, and watch the user-facing signal before completing the rollout.

saying these in an interview costs you the question

  • Says rollback is always correct with no exceptions
  • Chooses fix-forward because reverting feels like giving up
  • Starts a hot fix with no deadline or fallback agreed
  • Forgets that the old build may not read migrated data
  • Leaves users failing while the fix is written

context