Every service in your estate fetches its credentials from one store at start-up — how do you decide which of them may keep that dependency on the start-up path?
answer
- the path sets the floor
- tier by cost of a failed start
- bound before you re-architect
- off the path means resting copies
- standardise, then rehearse it
basics
~10 sNo service can have a higher start-up availability than whatever sits on its start-up path. Tier the estate, let most services keep the dependency, and remove it only where a tier cannot wait.
solid answer
~50 sFrame it as a stake, not a preference: the store's availability is a floor under the start-up availability of everything that must fetch to start. Most of the estate should keep the dependency and simply bound the exposure — bounded retry, restart freezes during an outage, a rollout that halts on the first failed start-up, and headroom. For the small tier that genuinely cannot wait, the structural move is to take the store off the start-up path: the delivery mechanism writes the value to the host ahead of time and refreshes it in the background, so a restart reads locally and an outage stops only refreshes. That buys availability and pays for it with a credential resting on every host. Splitting into several stores is the expensive third option and usually trades one rare correlated outage for many small ones.
go deeper
Grasp the shape of the constraint: if a service must reach something to start, it cannot start more reliably than that thing. Everything else here follows from it.
Distinguish bounding the exposure from removing the dependency. Retries, freezes and headroom change the odds; only moving the fetch off the start-up path changes the structure.
Argue the trade honestly for a specific tier: what an hour without new instances costs, against a credential resting on every host and the chance of serving on a withdrawn value.
Own the standard rather than the case: a short list of permitted start-up behaviours per tier, a recorded age bound, a restart-freeze rule, and a rehearsal the estate actually runs.
## What the dependency really is If a process must reach the store to start, then the store's availability is a **floor** under that process's start-up availability. You cannot buy the service a better number than the thing on its path — you can only remove the thing from the path. Saying that plainly is most of the work, because it turns a vague unease into an ordinary architectural trade that leads can decide. It is also worth saying what the single store buys, because the decision is not one-sided. One store is what makes a real inventory possible, what makes rotation a single operation rather than a hunt, and what makes an access trail exist at all. Correlated failure is the price of those, and the usual right answer is to **bound** the price, not to abandon what it bought. ## Tier before you engineer Sort services by what a failed start-up during an outage actually costs: - **Can wait.** Batch work, internal tooling, anything whose instances are not being replaced hour to hour. Keep the dependency; do nothing structural. - **Must not lose capacity.** Customer-facing services that scale, deploy and replace hosts continuously, where an hour without new instances is an outage. These are candidates for coming off the path. - **Must not serve on a doubtful value.** Services where running on a credential that may have been withdrawn is worse than being down. These should refuse to start, and the answer for them is headroom and restart discipline, not a local copy. The third tier is the one teams get wrong on their own, because the locally rational choice is always "come up". ## Three moves, in increasing cost | Move | What it changes | What it costs | |---|---|---| | Bound the exposure | Bounded retry, restart freeze, halt-on-first-failure rollouts, headroom | Almost nothing; no structural change | | Take the store off the start-up path | Value written to the host ahead of time and refreshed in the background; a restart reads locally | A credential resting on every host, and a value that may already have been withdrawn | | Split the dependency | Separate stores per region or per tier | Multiplied operational surface: inventory, rotation, trail and access rules in several places | Move one is where every estate should start, and it is usually enough. Move two is the real answer for the small tier that cannot wait: it converts the store from a start-up dependency into a refresh dependency, so an outage stops updates rather than stopping starts. Move three reduces correlated failure and reliably trades one rare estate-wide event for several smaller ones plus permanently higher operating cost; take it for isolation reasons, rarely for availability alone. ## What you standardise With many teams, the failure mode is not a bad decision but **no decision anyone can see**. Standardise so the estate's behaviour is knowable: 1. The **start-up behaviour per tier** — refuse, or a bounded local start — chosen from a short list rather than invented per service. 2. The **maximum age** a local copy may have at start-up, and the requirement that the fetch time is recorded with it. 3. That a local start is **logged and alerted**, so a responder knows the fleet is coming up on copies while they are mid-incident. 4. The **restart-freeze rule**: while a shared start-up dependency is unavailable, deploys, scaling actions and voluntary host maintenance stop, and rollouts halt at the first instance that fails to become ready. 5. A **rehearsal requirement**: each tier proves, on a schedule, that an instance restarted with the store blocked does what its tier says it should. The start-up path is the one nobody exercises, so it is the one that has to be tested deliberately. ## What not to do Do not let each team choose silently. An estate where half the services refuse to start and half come up on local copies, and nobody can say which is which, has the costs of both designs and the guarantees of neither — and the question "what comes back if the store is down for two hours?" has no answer until it happens. And do not let the recovery path depend on the thing being recovered. Whatever is needed to bring the store back must be reachable without the store, held by named people, and exercised — otherwise every other decision here is theoretical.
- What does taking the store off the start-up path actually mean in practice?The delivery mechanism fetches the value ahead of need and writes it where the host can serve it, refreshing in the background. A restarting process reads locally and never calls the store, so an outage stops refreshes rather than starts. The price is a credential resting on every host and a value that may have been withdrawn since the last refresh.
- Why is splitting into several stores usually the wrong first answer?It reduces correlated failure but multiplies everything the single store made possible: inventory, rotation, access rules and the audit trail now exist in several places and drift apart. It trades one rare estate-wide event for several smaller ones and a permanent increase in operating cost. Take it for isolation or blast-radius reasons, not for availability alone.
- What would you require every tier to prove, and how often?That an instance restarted with the store deliberately blocked behaves the way its tier says it should — refuses to start, or comes up on a copy inside the allowed age with an alert fired. Run it on a fixed schedule in a real environment. Without that rehearsal, the estate's start-up behaviour is a belief rather than a measured property.
saying these in an interview costs you the question
- Promises a service more start-up availability than the thing on its path.
- Treats sharding the store as the default fix for correlated failure.
- Leaves each team to choose its own start-up behaviour silently.
- Assumes a local copy is free of standing exposure.
- Keeps the store's recovery credential inside the store itself.
- Never exercises a restart with the store blocked.