skip to content

A store outage stayed harmless for an hour, then a routine fleet-wide host replacement began — what turned it into an estate-wide outage?

level: seniorimportance: should knowfreq 44%

answer

  1. the outage and the damage are two events
  2. one dependency on every start-up path
  3. failures correlated, not independent
  4. crash loops feed the outage
  5. halt the wave on first failed start

basics

~20 s

The replacement moved every workload into the population that must fetch, all at once. Because every start-up path runs through the same store, those start-ups failed together rather than independently — a correlated failure the fleet's capacity planning never assumed.

solid answer

~40 s

For an hour nothing restarted, so nothing had to fetch and the outage was invisible. The host replacement then pushed the whole estate onto the start-up path simultaneously. Every workload resolves credentials from **one** store, so the failures were correlated: not a few percent of instances failing independently, but every replaced instance failing for the same reason at the same time. Three things amplified it — a failed fetch exits, so the platform replaces the instance again and it crash-loops; nothing halted the wave when new instances failed to become ready; and when the store returned, every waiting instance retried at once. The control that would have bounded it is unglamorous: stop the wave at the first instance that fails to start.

code

pseudocode · 15 lines
pseudocode
maxConcurrent = 1 per 20 instances
failedStarts  = 0

for each instance in replacementPlan:
    if storeReachable() == false and instance.reason != CRASHED:
        halt("store unreachable; restart wave paused")

    replace(instance)
    waitUntil(instance.startedServing, timeout = startupBudget)

    if instance.startedServing == false:
        failedStarts = failedStarts + 1
        halt("new instance never served; wave stopped after " + failedStarts)

    waitForFreeSlot(maxConcurrent)

go deeper

for a junior

Understand that an outage only hurts when something has to restart. Restarting a service during a store outage is not a neutral act; it is what turns a quiet problem into a loud one.

for a middle

Explain correlation: one store on every start-up path means start-ups fail together, so the usual assumption that a few instances fail while others carry the load no longer holds.

for a senior

Name the amplifiers and the controls: crash loops, waves with no stopping rule, retry storms on recovery; against them, halt-on-first-failure, restart freezes, bounded retry with jitter, and real headroom.

for a principal

Own the rule that connects routine operations to dependency health — no restart wave proceeds while a shared start-up dependency is down — and make a store-blocked restart a rehearsal the estate actually runs.

## Why the first hour was free Nothing restarted. Every process was serving on credentials it had resolved before the window opened, so nothing had to reach the store and nothing noticed it was gone. That hour is not evidence of resilience; it is evidence that no start-up happened. The outage and the damage are two different events, and the second one had not started yet. ## What the replacement wave changed A fleet-wide host replacement takes every instance and makes a new one. Each new process runs the start-up path, and the start-up path runs through the store. In one operation the population that must fetch went from zero to everything. The important property is **correlation**, not volume. Capacity planning assumes instance failures are roughly independent: a few fail, the rest carry the load, replacements arrive. Here every failure has the same cause at the same moment, so there are no healthy replacements to arrive. A shared start-up dependency converts a fleet of many independent things into one thing with many copies. ## The amplifiers - **Crash looping.** A process that exits because the fetch failed is, to the platform, an instance that needs replacing. It replaces it, the fetch fails again, and the loop adds load to the very system that is down. - **A wave that does not look back.** Replacement proceeds on its own schedule. If nothing checks whether the *new* instances are actually serving, the operation happily retires four hundred healthy processes and creates four hundred that never start. - **Health checks read as "replace me".** An instance that never becomes ready looks identical to a broken instance, so the platform keeps doing the one thing that makes it worse. - **Retry storms on recovery.** When the store comes back, everything that has been failing retries at once. Without jitter and backoff the recovery load can be larger than the normal load and can knock the store over again, which is how a single outage becomes three. - **The recovery path needing the recovered thing.** If restoring the store itself requires a credential that lives in the store, the outage has no exit. That credential belongs somewhere reachable without it, held by named people. ## What actually bounds it 1. **Halt the wave on the first failed start-up.** One rule, cheap to implement, and it is the difference between losing one instance and losing the estate. Cap concurrent replacements too, so "the first failure" arrives while there is still a fleet. 2. **Freeze restarts during a store outage.** Deploy freeze, scaling held, host maintenance paused, and an explicit instruction not to restart things to see whether it helps. This is the highest-value action available to the on-call engineer and the one most often skipped. 3. **Retry, with bounds, before exiting.** A start-up path that retries with backoff and jitter for a bounded period turns most store unavailability into a slow start rather than a failed one, and spreads the recovery load instead of concentrating it. 4. **Keep headroom.** If losing a third of the fleet is a user-visible outage, the store's availability is already your availability. 5. **Rehearse the restart, not the outage.** Block the store, restart one instance in a real environment, and watch what it does. Almost no estate has done this, which is why almost every estate is guessing about the answer. ## Naming the mechanism in the post-mortem The honest sentence is not "the store went down and we went down". It is: *for an hour the outage cost nothing because nothing restarted; the host replacement then required every instance to fetch; every fetch failed for the same reason; the wave had no stopping rule, so it kept retiring healthy processes and creating ones that could not start; and the recovery was slowed further by every failed instance retrying at the same second.* Each clause points at a different control, and only one of them is about the store. ## The uncomfortable conclusion The outage was survivable and the operation was routine. The estate-wide failure came from running them at the same time with no rule that connected them. That rule — *do not restart what you do not have to while the store is unreachable, and stop the wave when new instances do not come up* — costs nothing and is missing almost everywhere.

  • Why does a shared start-up dependency break ordinary capacity planning?
    Capacity planning assumes failures are roughly independent, so a small fraction fails and replacements arrive from a healthy pool. A shared dependency makes every start-up fail for the same reason at the same moment, so there is no healthy pool and no replacement arrives. The fleet behaves as one unit with many copies rather than as many units.
  • The store comes back. What can still go wrong in the next five minutes?
    Everything that has been failing retries simultaneously, and that synchronised load can exceed normal load by a wide margin and take the store down again. Backoff with jitter in the start-up path, and staged resumption of the replacement wave, keep the recovery from becoming the second outage.
  • What single control would have kept the wave from costing the estate?
    Stopping the replacement at the first instance that failed to become ready, with a cap on how many are replaced concurrently so that the first failure arrives while most of the fleet is still serving. It is a few lines in the rollout logic and it converts a fleet-wide loss into the loss of one instance.

saying these in an interview costs you the question

  • Concludes the estate is resilient because the first hour was quiet.
  • Treats instance start-up failures as independent events.
  • Lets a replacement wave continue without checking new instances serve.
  • Restarts more instances during the outage hoping to clear it.
  • Keeps the store's own recovery credential inside the store.
  • Ignores the synchronised retry load when the store returns.