skip to content

A rollback of a bad release finished two minutes ago and the server-side error-rate graph is falling. What do you check before you call the rollback successful?

level: middleimportance: should knowfreq 48%

answer

  1. falling errors is weak evidence
  2. check the version label on traffic
  3. watch request rate beside error rate
  4. baseline, not the incident peak
  5. slice by region, tenant, client version

basics

~20 s

Confirm every instance actually runs the old version, compare the user-facing SLI against the pre-release baseline rather than the incident peak, watch request rate alongside errors, and clear residual bad state such as poisoned caches and parked messages.

solid answer

~50 s

First, verify the rollback actually completed everywhere: partial rollbacks are routine - a stuck instance, a second region, a scaling group still launching from the old template, or an edge cache still serving the bad assets. Slice the SLI by version label rather than trusting the deploy tool's summary. Second, read the right signal: measure at the user-facing edge, not just server-side, and watch request rate next to error rate, because errors also fall when clients give up. Third, compare to the pre-release baseline, not the incident peak - anything looks better than the peak. Give it a window at least as long as it took the problem to show up in the first place, and confirm the budget burn has come back to normal rather than eyeballing sixty seconds of graph. Finally, re-check the specific symptom sliced by the dimension it appeared on, since a one-region or one-client-version failure disappears in the aggregate.

go deeper

for a junior

Know that a rollback is not verified by the deploy tool's success message, and that you should check which version is actually serving traffic and whether errors returned to normal levels.

for a middle

Explain why request rate must be read alongside error rate, why the comparison baseline is the pre-release period rather than the incident peak, and why the observation window should match the original detection latency.

for a senior

Show that you hunt partial rollbacks and residual state deliberately - version-labelled SLIs, poisoned caches, parked messages - and that you re-check the specific slice the incident lived on rather than the aggregate.

for a principal

Define what "recovered" means as a standard: which SLI, measured where, over what window, against which baseline - so declaring recovery is a check anyone on call can run identically rather than a judgment call.

## Why "the graph is falling" is not enough A falling error rate two minutes after a rollback is weak evidence. It is consistent with recovery, and it is equally consistent with a partial rollback, with clients giving up, with a metric computed over a window that has not caught up, and with the aggregate hiding a cohort that is still broken. Verification is a short checklist, and running it is the difference between an incident that ends and one that reopens twenty minutes later. ## 1. Did the rollback actually land everywhere? Partial rollbacks are the single most common reason a rollback "didn't work": - one instance or node failed to cycle and is still running the bad version; - a second region, cluster or cell was never targeted by the rollback command; - an autoscaling group is still launching instances from a template that points at the bad release; - an edge or CDN cache is still serving the bad client bundle, so users get the new front end against the old backend; - a controller reconciling against source control is quietly putting the bad version back because the revert has not landed. The check is mechanical: emit the running version as a label on your request metrics and confirm that traffic on the bad version has gone to zero, rather than trusting the deployment tool's "succeeded". ## 2. Are you reading a signal that represents users? Measure as close to the user as you can. Server-side error rate misses everything that failed before reaching your handler - a load balancer returning 5xx, TLS failures, timeouts where the client gave up first. If the release broke something upstream of your application, your own error rate may look fine throughout. And always read **request rate alongside error rate**. A falling error *count* with falling traffic often means clients stopped trying: retries exhausted, a mobile client backing off, an upstream circuit opening. Recovery looks like errors falling while traffic returns to its normal shape. Where you have them, success-rate ratios rather than raw counts protect you from this, but the traffic curve is the sanity check. ## 3. Compare against the baseline, not the peak During an incident everything is measured against the worst moment, and everything is an improvement on the worst moment. The question is not "is it better than five minutes ago" but "is it back to where it was before the release". Pull up the same window from the previous day or the pre-release period and compare shape, not just level. A useful framing here is the burn rate: at a burn rate of 1 you are consuming your error budget exactly as fast as the window allows, so recovery means the burn has come back to its normal baseline - not merely that it has come down from the incident. Watching burn return to baseline is more honest than watching a raw error graph flatten. ## 4. Give it a window with meaning How long you watch should be tied to how long the problem took to appear. If the regression only showed up after five minutes of traffic - a slow leak, a cache filling, a periodic job - then a two-minute observation proves nothing. Watch for at least the original detection latency, and preferably across one cycle of whatever periodic work the service does. ## 5. Look for residue The rollback restores your code, not the state the bad release created: - caches holding objects in the wrong format keep failing until invalidated; - messages already published in a bad shape sit in queues and will be redelivered; - rows written incorrectly during the window stay written; - alerts may still be firing on stale evaluation windows, and silencing them is not the same as fixing them. ## 6. Slice by the dimension the incident lived on Aggregates hide cohorts. If the failure was specific to one region, one tenant, one client version or one endpoint, the global success rate can look completely healthy while that cohort is still at zero. Re-run the check on the specific slice where the symptom appeared, and on the specific request path, before declaring it over. ## What "successful" should mean A defensible bar: every instance serving traffic is on the intended version; the user-facing SLI is at its pre-release baseline over a window at least as long as the original detection latency; request volume is normal; budget burn is back to baseline; the specific failing slice is healthy; and any residual bad state is either cleared or has a named owner and a follow-up item. Anything less and you are hoping, not verifying.

  • Error rate dropped to zero and so did traffic. What do you conclude?
    Nothing good. Zero errors with zero requests usually means clients stopped reaching you - retries exhausted, an upstream circuit opened, or a load balancer removing every backend from rotation. Recovery looks like errors falling *while* traffic returns to its normal curve. Ratio-based SLIs are less misleading than counts here, but the traffic shape is the check that catches it.
  • How long should you watch before declaring the rollback successful?
    At least as long as the regression originally took to become visible, and ideally across one cycle of whatever periodic work the service performs. If the bad release only degraded after five minutes of traffic warming a cache, or after an hourly job ran, a two-minute observation proves nothing. Tie the observation window to the detection latency you actually experienced rather than to impatience.
  • Why is the deployment tool reporting "rollback succeeded" not sufficient evidence?
    It reports that the orchestration completed, not that all traffic is served by the intended version. Stuck instances, an untargeted region, a scaling group launching from a stale template, or an edge cache still holding the bad bundle all survive a green deployment status. Confirming from the request metrics themselves - version as a label, bad-version traffic at zero - is evidence; the tool's summary is a claim.

saying these in an interview costs you the question

  • The error graph is falling, so we are recovered
  • Deployment tool says success, so every instance is rolled back
  • Compare against the incident peak to judge recovery
  • Server-side metrics are enough to prove users are fine
  • Once the alert clears, there is nothing left to clean up

context