A secret store is unreachable for twenty minutes; why do already-running services keep serving while every instance that restarts in that window fails?
answer
- the outage is not the event
- start-up path, not request path
- resolved once, then held
- only what must fetch is affected
- a restart forces the re-fetch
basics
~10 sMost workloads resolve credentials once, on the start-up path, then hold them in memory. An unreachable store therefore breaks only what must fetch during the window, and a restart is what forces a fetch.
solid answer
~40 sReaching the store is normally on the **start-up path**, not the request path: a process authenticates once, resolves the names it needs, and uses the values it holds for every request afterwards. So an outage of the store does not touch traffic that is already being served — it affects only a process that has to ask again. A restart is the ordinary way a process joins that population: it loses what it held and must re-fetch before it can serve. That is why the same twenty-minute outage is invisible on a quiet fleet and expensive on one that is deploying, scaling out or replacing hosts. The blast radius is not "the estate" but "whatever fetches inside the window", and restarts decide how large that is.
go deeper
Know that a service usually asks for its credentials once, when it starts, and keeps them while it runs. That single fact explains why an unreachable store can look like nothing is wrong.
Explain the split: the store sits on the start-up path, so an outage affects only processes that must fetch inside the window. List what forces a fetch — restart, scale-out, host replacement, crash loop, refresh.
During an incident, size the exposure by counting what must fetch, freeze deploys, scaling and casual restarts, and alert on instances that never become ready rather than on request errors.
Treat "the fleet served fine last time" as untested, not as evidence. The start-up path is the one nobody exercises, so make a store-blocked restart a rehearsal the estate owes, not an accident.
## Where the store actually sits A workload that uses a stored credential normally reads it **once, on the start-up path**: it authenticates to the store, asks for the names it needs, receives the values, and keeps them in process memory for as long as it runs. Every request after that uses the copy already in memory. The store is therefore a dependency of **starting**, not a dependency of **serving** — and that one fact decides what an outage does to you. The exception is real and worth naming. A design in which the process calls the store on every request, or on a very short cycle, does put the store on the request path, and such a service stops the moment the store does. That design is uncommon, because it turns every request into a network call to a system that has no reason to be faster than your database. Most estates land on resolve-once-then-hold, and the outage story below is the one they get. ## Why the restart is the event If a value is resolved once and held, then during an outage a process only notices the store when it has to ask again. A **restart** is the ordinary way that happens: the process loses everything it held, and the first thing its replacement does is the fetch that now fails. The population at risk during the window is not the estate — it is *whatever has to fetch inside the window*, and restarts are what move an instance into that population. | What the instance is doing during the outage | Does it need the store? | Outcome | |---|---|---| | Serving, started hours ago | No — it holds the resolved value | Unaffected | | Being replaced by a deploy | Yes — the new process must fetch | Fails to start | | Added by a scale-out | Yes | Fails to start | | Moved because its host was replaced | Yes | Fails to start | | Restarted after a crash | Yes | Fails to start, and may crash-loop | This is why one twenty-minute outage produces two completely different post-mortems. On a quiet fleet nobody notices. On a fleet that is deploying, scaling out or rotating hosts, capacity drains for twenty minutes and does not come back on its own. ## What else forces a fetch Restarts are the main trigger, not the only one: - a **scale-out** adding instances that never held anything; - a host being replaced, which is a restart wearing different clothes; - a **crash loop**, where the failed fetch is itself the reason for the next restart; - a background **refresh** coming due, so a running process asks again; - an operator restarting something to "clear" an unrelated error — the most common way a harmless outage becomes a visible one; - a configuration reload that re-runs the same resolution code the start-up path runs. (What happens when a credential a process is *already holding* stops being accepted while work is in flight is a separate subject with its own failure modes. Here the point is only that a refresh is another reason a running process talks to the store.) ## Reading the blast radius during the incident For whoever is on call, the useful moves follow directly: 1. **Count what must fetch**, not what is deployed. Four hundred instances with nothing restarting means zero instances at risk. 2. **Stop anything that creates fetches.** Freeze the deploy, hold the scale-out, and — hardest — stop people restarting things to see whether it helps. 3. **Watch start-up failures, not request errors.** The signal for this outage is instances that never became ready, which most request dashboards do not show as an error at all. ## What a surviving fleet does and does not prove A fleet that served normally through an outage proves exactly one thing: nothing in it had to fetch. It does **not** prove that the estate survives a store outage, because the path that would have failed — the start-up fetch — was never exercised. The only honest test is to block the store deliberately and restart one instance, and what that instance then does is the decision this leaf is really about: whether it refuses to start, or comes up on something it already had.
- Besides a deliberate deploy, what else moves a running instance into the population that must fetch?A scale-out adding instances that hold nothing yet; a host being replaced under the workload; a crash loop, where the failed fetch causes the next restart; a background refresh coming due; a configuration reload that re-runs the resolution code; and an operator restarting a service to clear an unrelated symptom. The last one turns far more store outages into user-visible ones than any of the others.
- Your fleet served normally through a store outage last month. What did that prove?That nothing in the fleet had to fetch during the window — no deploy, no scale-out, no host churn. It proves nothing about survivability, because the start-up fetch was never exercised. The only way to learn what an instance does is to block the store and restart one on purpose, which is the rehearsal almost nobody runs.
- What signal would have told you early that the outage was costing capacity?Instances that were created but never became ready, and a gap between desired and serving instance counts. Request-level error rates stay flat, because the instances that are serving are fine; the failure is entirely on the start-up path, so the useful alert is on start-up failures and on ready-count drift, not on response codes.
saying these in an interview costs you the question
- Says a store outage immediately takes every service down.
- Assumes each request resolves its credential from the store.
- Treats a flat error rate during the outage as proof nothing depends on the store.
- Sees no problem with continuing a deploy while the store is unreachable.
- Restarts an instance as a first response to store errors.
- Thinks a healthy fleet during one outage proves start-up survivability.