Why does a secret store partition that steady-state services barely notice break a service scaling out and a fleet on a rolling restart?
answer
- steady state barely touches the store
- a start-up dependency, not per-request
- new processes are first fetches
- a rollout discards warm copies
- freeze the change rate first
basics
~20 sThe store is a start-up dependency, not a per-request one. Already-running processes hold what they fetched and never call it, so a degraded store is invisible until something new starts — and scaling out and rolling restarts are the two activities that manufacture new start-ups.
solid answer
~50 sA process that has already fetched what it needs stops calling the store, so during a partition the running estate looks completely healthy. Every *new* process is a first fetch, and both a scale-out and a rolling restart exist to create new processes. The restart is the worse of the two, because it deliberately throws away the warm copies held by processes that were working, converting a degraded store into a spreading outage one batch at a time. Both are usually automated and will keep going: the scaling automation adds instances because peak demand is rising, and the rollout marks an instance healthy on a signal that says nothing about whether it ever reached the store. The first operator move is to freeze the rate of change — pause the rollout, hold the instance count — and let what is already running carry the peak.
go deeper
Know that many services read their credentials once at start-up and then hold them, which is why a store problem can be invisible while everything is running and obvious the moment something restarts.
Explain the mechanism: exposure scales with the number of processes starting, not with traffic. Be able to say why a rolling restart is worse than a scale-out, since it removes processes that were already working.
Show the incident judgment: freeze the rate of change before anything else, resist restarting failing instances, and know which consumers cannot start at all against a read-only store because their start-up path writes.
Own the coupling itself. Decide whether deployments and scaling are allowed to run while the store is degraded, who can halt them, and what dependency signal a rollout must consult before it takes working capacity away.
## Why steady state hides the failure Most consumers touch a secret store far less often than people assume. A service fetches the credentials it needs when it starts, holds them for as long as it runs, and then serves requests without going near the store again until something expires. That is exactly what makes the store survivable in normal operation — and exactly what makes a degraded store invisible. So during a partition at morning peak, the dashboards of a hundred running services are green. Their request rate against the store is close to zero. Nothing they do exercises the broken path. The store is a **start-up dependency**. Its failures are latent, and the latency is the gap between the failure starting and the next process starting. ## The two activities that convert a latent failure into an outage Both of these exist for good reasons, and both do the same dangerous thing: they create processes that have never fetched anything. 1. **Scaling out.** A customer-facing service adds instances because demand is climbing. Every new instance is a first fetch against a store that cannot serve it. The new instances fail to come up, the existing ones absorb the demand they were supposed to shed, and the scaling automation — seeing the load still rising — asks for more. The estate's response to peak becomes a stream of stillborn processes. 2. **A rolling restart.** A fleet of edge collectors is being replaced batch by batch. This is worse than scaling out in one specific way: it does not merely fail to add capacity, it **destroys capacity that was working**. Each batch takes processes that were happily holding values fetched before the partition and replaces them with processes that must fetch now. The failure spreads at the pace of the rollout. The read-only case sharpens both. Where a consumer starts by reading a value that is already stored, a read-only store serves it and the new process comes up fine. Where a consumer starts by asking for a credential the store generates per consumer, that is a write, and it fails instantly — so two services with identical restart behaviour have completely different outcomes, decided by which kind of credential they were designed to use. ## Why the automation keeps going The part that turns this into a long incident is that neither activity stops on its own: - The rollout marks an instance healthy on a signal that usually reflects the process itself, not whether it obtained what it needs from an external dependency. Where the new instance retries its start-up fetch for a while before giving up, the rollout can even count it as progressing. - The earlier batches genuinely succeeded, because they started before the partition. The operator watching sees a rollout that is going fine, right up until the batch that does not. - The scaling automation reads demand, not dependency health, so a degraded store makes it ask for *more* instances rather than fewer. ## What an operator actually does The store's state is usually not yours to fix in the first ten minutes, but the estate's rate of change is: - **Pause the rollout.** Every batch not yet started is an outage not yet caused. - **Hold the instance count.** Stop the scaling automation from manufacturing first fetches; the capacity you have running is the capacity you have. - **Do not restart the instances showing errors.** This is the reflex that turns a partial incident into a total one: a restarted process loses everything it was holding and joins the set that cannot come up. - **Find out which consumers need a generated credential to start**, because those are the ones that cannot come up even against a read-only store, and they will need the write path back before they can be moved at all. - **Know your expiry deadline**, since consumers holding something that will expire during the outage cannot renew and will restart into the same wall. What a restarting consumer *should* do when the store will not answer — refuse to start, or come up on a value it cached earlier, and for how long that copy stays usable — is a design decision that belongs to the consumer, and it is a separate subject from the store's own topology. What belongs here is the timing: the store's degradation and the estate's change rate multiply, and you control the second one. ## The point to make in an interview Say the store is a start-up dependency and that the estate's exposure is proportional to how many processes are starting, not to how much traffic it is serving. Then name the two activities that raise that number, say that a rolling restart also removes warm capacity, and give the first operator move: freeze the change rate before doing anything else.
- Two services restart during the same read-only window and only one comes up. What separates them?What each one asks the store for at start-up. A service that reads a stored value gets it, because reads still work. A service designed around a credential the store generates per consumer is making a write request, and a read-only store rejects it — so its start-up fails while its neighbour's succeeds. The difference is a design choice made long before the outage.
- Why is the scaling automation's behaviour during this window actively harmful rather than merely useless?Because it reads demand, not dependency health. Peak load keeps climbing, the existing instances saturate, and the automation responds by asking for more instances — each one a fresh start-up fetch against a store that cannot serve it. That adds churn and load to the degraded store at the moment you most want the estate to hold still.
saying these in an interview costs you the question
- Assumes green steady-state dashboards prove the store is healthy
- Thinks a rolling restart is safe because it proceeds gradually
- Lets a rollout continue because the earlier batches succeeded
- Believes scaling out relieves pressure on a degraded store
- Restarts failing instances first instead of freezing the change rate
- Treats a failed start-up fetch as a defect in the new instance