You set release-engineering standards for an organisation running several dozen services. Which services would you require to maintain a tested, single-action rollback path, and where would you accept that fixing forward is the only realistic option?
answer
- capability, not the in-incident decision
- undrilled rollback is a claim
- measure time-to-rollback like deploy time
- irreversible effects cannot be recalled
- pipeline time is the fix-forward floor
basics
~20 sMandate a drilled, measured rollback for stateless services on the critical user path, where it is cheap. Accept fix-forward where effects are already irreversible outside the system - then make the pipeline fast, because its duration becomes the recovery floor.
solid answer
~50 sI would grade services by two axes: how expensive an unrecovered minute is, and how much state or external effect a release entangles. Stateless request-serving services on the critical user path get a hard requirement - rollback in a single action, exercised on a schedule, with time-to-rollback measured and reported like deploy time, because a path nobody has run is a claim rather than a control. Services whose releases move data, take payments, send messages or ship code to clients you cannot recall genuinely cannot roll back the effects; for those I require an explicit, written statement that they are fix-forward, and I invest instead in shortening pipeline time, since the pipeline's duration becomes the floor on recovery. The requirement is not free: keeping every release compatible with the previous one turns simple migrations into multi-step parallel changes over days. Spend that where blast radius justifies it, and be honest where it does not.
go deeper
Know that not every service can be rolled back, and that the ones that cannot depend on how fast a fix can reach production instead.
Be able to name what makes rollback expensive for a given service - stateful data, completed migrations, effects already sent outside the system - rather than treating rollback as a universal tooling feature.
Argue the per-service call with evidence: measure time-to-rollback, exercise the path outside incidents, and state the rollback horizon so nobody plans a response around a capability that does not exist.
Own the standard and its cost. Tier services by blast radius, price the parallel-change and dual-write burden you are imposing, set an expedited-pipeline target where fix-forward is honest, and verify with evidence from deployment history rather than attestation.
## The standard is about capability, not about the moment During an incident, whether to roll back or fix forward is a live judgment call made with partial information. A release-engineering standard cannot make that call in advance. What it can do is guarantee that when the call is made, at least one of the two routes is fast, tested and known. That is the framing to bring to this question: you are deciding where to *invest* in reversibility, not scripting the incident. ## What makes a rollback capability real Four properties separate a capability from a claim: 1. **Single action.** If rolling back means twelve manual steps and an approval, its duration is a floor under every bad-release incident. One command or one revert. 2. **Measured.** Time-to-rollback should be a tracked number next to deploy time and change failure rate. What is not measured drifts. 3. **Exercised.** A rollback path that has never been run is untested code. Stateless services can exercise it on a schedule - roll back and forward again in a low-traffic window; others need a periodic drill. 4. **Bounded and known.** Every service should be able to state how far back it can go and why, because that horizon moves every time a migration completes. A rollback button that exists in a tool and has never been pressed is not one of these things. ## Where I would mandate it - **Stateless request-serving services on the critical user path.** Rollback is cheap here and the value is highest: minutes of user-visible impact avoided per incident. - **Edge, routing and configuration layers.** These fail globally and instantly; the ability to revert one revision is the primary control. - **Teams shipping frequently.** A high change rate multiplies the value of a cheap undo. For these, the compatibility rule follows automatically: each release must run against the state the previous one produced, migrations follow parallel change, and configuration is versioned with the artifact. ## Where fix-forward is the honest answer - **Irreversible external effects.** Funds moved, emails and notifications sent, third-party webhooks fired, physical fulfilment triggered. You can restore the code; you cannot recall the effect. Here the investment belongs in limiting the blast radius and in idempotency, not in a rollback that only half-works. - **Stateful stores and anything mid-migration.** Once data has moved, the previous version may simply be unable to run. - **Clients you cannot recall.** Shipped mobile or desktop builds, embedded devices, and to a lesser extent cached browser bundles. The real lever becomes server-side compatibility with older clients - which means the *server's* rollback safety is what you invest in. - **Long-running stateful streams and jobs** where restarting on the old version means reprocessing or losing position. For these I want the fix-forward path explicitly acknowledged, because the failure mode is a team that plans its incident response around a rollback that does not exist. ## Making fix-forward viable If a service cannot roll back, its recovery time is bounded below by how long it takes to get a change to production: commit, build, verify, deploy. That number is the service's real recovery floor, and shortening it is the equivalent investment - an expedited path with a reduced but honest verification set, pre-approved for incident use, and drilled so nobody is discovering it under pressure. Alongside that, invest in blast-radius limits so a bad release affects a small share of traffic before anyone has to choose a route at all. ## The cost of the mandate Requiring universal rollback readiness is not free, and a principal engineer should price it out loud: - every schema change becomes a multi-step parallel change spread over days; - dual-write phases cost latency and write amplification; - releases carry compatibility code that must be cleaned up, and if it is not, it accumulates; - version compatibility has to be tested, which means running the previous release against current state as a routine check. That is real engineering time. Spend it where an unrecovered minute is expensive, and be willing to say plainly that a low-traffic internal service does not need it. ## How I would express the standard A short tiering, applied per service and reviewed when its criticality changes: tier 1 requires single-action rollback, a stated horizon, regular exercise, and a measured time-to-rollback; tier 2 requires a documented route with an owner; tier 3 declares fix-forward and commits to an expedited pipeline target instead. Then hold it with evidence rather than attestation - a rollback that appears in the deployment history is proof, a checkbox is not.
- How would you verify compliance with a standard like this without relying on teams' self-attestation?Look for evidence in systems you already have: deployment history showing an actual rollback within the review period, a recorded time-to-rollback, and a stated horizon in the service's operational metadata. A checkbox in a spreadsheet proves someone read the policy. A rollback that appears in the deploy log with a duration attached proves the path works, and it costs the team almost nothing if the capability is genuinely there.
- A team argues that rollback readiness slows them down too much to be worth it. How do you respond?Price both sides for their service. The cost is real: multi-step migrations, dual-write phases, compatibility code to clean up. The benefit is the number of user-visible minutes a cheap undo saves, which depends on their change failure rate and their traffic. If they are a low-traffic internal service, they may be right and I would tier them down. If they are on the payment path, the arithmetic answers itself.
- For a service that can only fix forward, what target would you set and why?An expedited commit-to-production time, because that is the floor on its recovery. I would define a reduced but honest verification set for incident changes, pre-approve the route so no approval is invented under pressure, and drill it. I would also pair it with blast-radius limits, so fewer users are affected during the window the fix takes - reducing exposure is the substitute for reducing time.
saying these in an interview costs you the question
- Every service should be able to roll back, no exceptions
- A rollback button in the deploy tool means we have rollback
- Rollback readiness is free once the tooling exists
- Fix-forward is just what teams say when they are undisciplined
- A documented rollback procedure is as good as a drilled one