skip to content

Surviving Store Outage

What a restart means while the store is unreachable, and the deliberate choice between refusing to start and coming up on a cached value. Asked because every workload now restarts through one system.

on this pageshow

questions

4

A secret store is unreachable for twenty minutes; why do already-running services keep serving while every instance that restarts in that window fails?

level: middleimportance: must knowfreq 58%

answer

  1. the outage is not the event
  2. start-up path, not request path
  3. resolved once, then held
  4. only what must fetch is affected
  5. a restart forces the re-fetch

basics

~10 s

Most workloads resolve credentials once, on the start-up path, then hold them in memory. An unreachable store therefore breaks only what must fetch during the window, and a restart is what forces a fetch.

solid answer

~40 s

Reaching the store is normally on the **start-up path**, not the request path: a process authenticates once, resolves the names it needs, and uses the values it holds for every request afterwards. So an outage of the store does not touch traffic that is already being served — it affects only a process that has to ask again. A restart is the ordinary way a process joins that population: it loses what it held and must re-fetch before it can serve. That is why the same twenty-minute outage is invisible on a quiet fleet and expensive on one that is deploying, scaling out or replacing hosts. The blast radius is not "the estate" but "whatever fetches inside the window", and restarts decide how large that is.

go deeper

for a junior

Know that a service usually asks for its credentials once, when it starts, and keeps them while it runs. That single fact explains why an unreachable store can look like nothing is wrong.

for a middle

Explain the split: the store sits on the start-up path, so an outage affects only processes that must fetch inside the window. List what forces a fetch — restart, scale-out, host replacement, crash loop, refresh.

for a senior

During an incident, size the exposure by counting what must fetch, freeze deploys, scaling and casual restarts, and alert on instances that never become ready rather than on request errors.

for a principal

Treat "the fleet served fine last time" as untested, not as evidence. The start-up path is the one nobody exercises, so make a store-blocked restart a rehearsal the estate owes, not an accident.

## Where the store actually sits A workload that uses a stored credential normally reads it **once, on the start-up path**: it authenticates to the store, asks for the names it needs, receives the values, and keeps them in process memory for as long as it runs. Every request after that uses the copy already in memory. The store is therefore a dependency of **starting**, not a dependency of **serving** — and that one fact decides what an outage does to you. The exception is real and worth naming. A design in which the process calls the store on every request, or on a very short cycle, does put the store on the request path, and such a service stops the moment the store does. That design is uncommon, because it turns every request into a network call to a system that has no reason to be faster than your database. Most estates land on resolve-once-then-hold, and the outage story below is the one they get. ## Why the restart is the event If a value is resolved once and held, then during an outage a process only notices the store when it has to ask again. A **restart** is the ordinary way that happens: the process loses everything it held, and the first thing its replacement does is the fetch that now fails. The population at risk during the window is not the estate — it is *whatever has to fetch inside the window*, and restarts are what move an instance into that population. | What the instance is doing during the outage | Does it need the store? | Outcome | |---|---|---| | Serving, started hours ago | No — it holds the resolved value | Unaffected | | Being replaced by a deploy | Yes — the new process must fetch | Fails to start | | Added by a scale-out | Yes | Fails to start | | Moved because its host was replaced | Yes | Fails to start | | Restarted after a crash | Yes | Fails to start, and may crash-loop | This is why one twenty-minute outage produces two completely different post-mortems. On a quiet fleet nobody notices. On a fleet that is deploying, scaling out or rotating hosts, capacity drains for twenty minutes and does not come back on its own. ## What else forces a fetch Restarts are the main trigger, not the only one: - a **scale-out** adding instances that never held anything; - a host being replaced, which is a restart wearing different clothes; - a **crash loop**, where the failed fetch is itself the reason for the next restart; - a background **refresh** coming due, so a running process asks again; - an operator restarting something to "clear" an unrelated error — the most common way a harmless outage becomes a visible one; - a configuration reload that re-runs the same resolution code the start-up path runs. (What happens when a credential a process is *already holding* stops being accepted while work is in flight is a separate subject with its own failure modes. Here the point is only that a refresh is another reason a running process talks to the store.) ## Reading the blast radius during the incident For whoever is on call, the useful moves follow directly: 1. **Count what must fetch**, not what is deployed. Four hundred instances with nothing restarting means zero instances at risk. 2. **Stop anything that creates fetches.** Freeze the deploy, hold the scale-out, and — hardest — stop people restarting things to see whether it helps. 3. **Watch start-up failures, not request errors.** The signal for this outage is instances that never became ready, which most request dashboards do not show as an error at all. ## What a surviving fleet does and does not prove A fleet that served normally through an outage proves exactly one thing: nothing in it had to fetch. It does **not** prove that the estate survives a store outage, because the path that would have failed — the start-up fetch — was never exercised. The only honest test is to block the store deliberately and restart one instance, and what that instance then does is the decision this leaf is really about: whether it refuses to start, or comes up on something it already had.

  • Besides a deliberate deploy, what else moves a running instance into the population that must fetch?
    A scale-out adding instances that hold nothing yet; a host being replaced under the workload; a crash loop, where the failed fetch causes the next restart; a background refresh coming due; a configuration reload that re-runs the resolution code; and an operator restarting a service to clear an unrelated symptom. The last one turns far more store outages into user-visible ones than any of the others.
  • Your fleet served normally through a store outage last month. What did that prove?
    That nothing in the fleet had to fetch during the window — no deploy, no scale-out, no host churn. It proves nothing about survivability, because the start-up fetch was never exercised. The only way to learn what an instance does is to block the store and restart one on purpose, which is the rehearsal almost nobody runs.
  • What signal would have told you early that the outage was costing capacity?
    Instances that were created but never became ready, and a gap between desired and serving instance counts. Request-level error rates stay flat, because the instances that are serving are fine; the failure is entirely on the start-up path, so the useful alert is on start-up failures and on ready-count drift, not on response codes.

saying these in an interview costs you the question

  • Says a store outage immediately takes every service down.
  • Assumes each request resolves its credential from the store.
  • Treats a flat error rate during the outage as proof nothing depends on the store.
  • Sees no problem with continuing a deploy while the store is unreachable.
  • Restarts an instance as a first response to store errors.
  • Thinks a healthy fleet during one outage proves start-up survivability.
open as a page

The store is unreachable at 2am and a payments instance restarts beside a cached host copy — what does refusing to start buy over coming up on it?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Refusing to start buys a loud, early failure and a guarantee that nothing serves on material that may already have been withdrawn. It costs capacity that never returns while the outage lasts, on a fleet that keeps losing instances.

open as a page

A store outage stayed harmless for an hour, then a routine fleet-wide host replacement began — what turned it into an estate-wide outage?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The replacement moved every workload into the population that must fetch, all at once. Because every start-up path runs through the same store, those start-ups failed together rather than independently — a correlated failure the fleet's capacity planning never assumed.

open as a page

Every service in your estate fetches its credentials from one store at start-up — how do you decide which of them may keep that dependency on the start-up path?

level: principalimportance: should knowfreq 33%

basics

~10 s

No service can have a higher start-up availability than whatever sits on its start-up path. Tier the estate, let most services keep the dependency, and remove it only where a tier cannot wait.

open as a page