The store is unreachable at 2am and a payments instance restarts beside a cached host copy — what does refusing to start buy over coming up on it?
answer
- two failures, opposite directions
- refusing ratchets capacity down
- the copy needs the downstream to accept it
- withdrawn beats rotated-with-overlap
- bound the age, alert on every cached start
basics
~20 sRefusing to start buys a loud, early failure and a guarantee that nothing serves on material that may already have been withdrawn. It costs capacity that never returns while the outage lasts, on a fleet that keeps losing instances.
solid answer
~50 sThe two behaviours fail in opposite directions. **Refusing to start** means the instance exits and takes no traffic: nothing runs on a credential that may have been withdrawn, the failure is early and in one place, and the instance count only ratchets downward for the length of the outage. **Coming up on the host copy** restores capacity, but the copy only works if the downstream system still accepts it — a rotation with an overlap window is survivable, a withdrawal is not, and in that case the instance starts, passes its health check, takes traffic and fails on first use, which is a worse shape than refusing. The copy also has to rest on the host to survive a restart at all. Accept it only up to a recorded age, and fail closed past that.
code
pseudocode · 24 lineson startup:
result = fetchFromStore(name) with bounded retry and backoff
if result.ok:
writeHostCopy(name, result.value, fetchedAt = now)
serve(result.value)
return
copy = readHostCopy(name)
if copy == MISSING:
log "no credential and no local copy"
exit(1) # fail closed
if startPolicy == REFUSE_ON_CACHED:
log "store unreachable; policy forbids a cached start"
exit(1) # fail closed
if now - copy.fetchedAt > maxAcceptedAge:
log "local copy too old", age = now - copy.fetchedAt
exit(1) # fail closed
alert "starting on a local copy", age = now - copy.fetchedAt
serve(copy.value) # may already have been withdrawngo deeper
Know the two outcomes by name: the process either refuses to start, or starts on a value it kept from last time. They fail in opposite directions and both are choices someone made in advance.
Explain why a cached start is not automatically a save: the value is only useful while the downstream system still accepts it, and a withdrawal turns a clean start-up failure into a failure after traffic arrives.
Design the middle: bounded retry, a recorded fetch time, a maximum accepted age, and an alert on every cached start. Be able to say what your service does today and when it last did it.
Decide it per tier rather than per team, and own the consequence: a fleet that refuses to start has staked its recovery on the store's, and a fleet that starts on copies has accepted a credential resting on every host.
## The two behaviours, stated precisely When the start-up fetch fails there are exactly two outcomes, and conflating them is the commonest error in this material. - **Failing closed** means the process **refuses to start**: it logs, exits non-zero, and the instance never takes traffic. - **Starting on cached material** means the process **comes up on a copy it already had**, written to the host by an earlier successful fetch, and serves on a value that may have been replaced or withdrawn while the store was unreachable. Both are legitimate. Neither is free, and which one a service should do is a property of that service, not a fact about secret management. ## What refusing to start costs During the outage the instance count is a one-way ratchet. Scale-outs, host replacements and crash loops all consume instances; nothing produces one, because nothing can start. A fleet with no headroom walks down toward a user-visible outage on a timetable set by someone else's system. The cost people only discover at 2am is that **you cannot change your mind**. Changing the start-up behaviour means shipping a change, and shipping it means starting new processes, which need the store. The decision was made when the code was written. What it buys is real: the failure is loud, early and in one place; nothing serves on a credential someone deliberately killed; and no process comes up with an empty value that some code path quietly treats as "no credential configured". ## What coming up on the copy costs - The copy only helps if the **downstream system still accepts it**. A cached start buys availability only inside the credential's remaining validity. - If the value was **rotated** during the window and both values are accepted for an overlap period, the instance works. If it was **withdrawn**, the instance starts, passes its start-up health check, takes traffic and fails on first use — worse than refusing, because it fails *after* accepting work. - If the value was withdrawn as part of an **incident response**, every cached start re-arms a credential somebody killed on purpose. - The copy must **rest on the host** to survive a restart at all. That is a standing exposure the fail-closed design simply does not have. - Unless the copy was written together with the time it was fetched, nobody can say how old it is — and its age is the whole question. | | Refuse to start | Start on the host copy | |---|---|---| | Capacity during the outage | Falls and stays down | Restored | | Serves on a withdrawn value | Never | Possible | | Where the failure shows | Start-up, immediately | First downstream call, after taking traffic | | Standing exposure | None added | A credential resting on every host | | Reversible at 2am | No | No | ## The bound that makes the middle option honest Most services should do neither purely. Build four things: 1. Write the copy with the **time it was fetched**, so its age is a fact rather than a guess. 2. Make the start-up fetch a **bounded retry with backoff** first, so a five-second blip costs nothing under either policy. 3. Define a **maximum age you will accept** at start-up; past it, fail closed. 4. **Log and alert on every cached start**, so "we came up on yesterday's copy" is something you know that night, not something you reconstruct later. ## Who chooses, and when The service owner chooses, at design time, and the choice is already encoded in the start-up path by the time the outage happens. An operator has no lever unless one was deliberately built — an explicit degraded-mode input the platform can set, which is itself a control somebody has to be authorised to use. The tiering is usually intuitive: a payments service whose owner would rather be down than run on a credential that may have been revoked; a read-mostly internal service whose owner would much rather it came up. ## The scenario, named At 2am the payments instance starts from a copy written by the last successful fetch, nine hours earlier. It survives if and only if the downstream account that credential belongs to still accepts it. If it is a long-lived static value nobody touched, it does. If it is a generated credential minted with a short validity, nine hours is almost certainly past its life, and the cached start buys nothing at all: the instance comes up, takes traffic and fails — the worst of both designs.
- Under what condition does the cached start actually work, and under what condition does it not?It works while the downstream system still accepts that value. A rotation with an overlap window leaves the old value accepted, so the instance serves. A withdrawal — the value revoked, the downstream account disabled — leaves it useless, and the instance discovers that only on its first real call, after it has already been given traffic.
- Why is a bounded retry at start-up worth building whichever policy you choose?Because most store unavailability is seconds, not hours. A short retry with backoff absorbs blips, partial failovers and restarts of the store itself, so neither policy is ever invoked for a five-second event. It also separates "could not reach the store right now" from "cannot reach the store at all", which are different decisions.
- The value was revoked deliberately during an incident. What does a cached start do then?It re-arms the revoked credential on every instance that restarts, which is the opposite of what the responder intended. This is the case that argues for a maximum accepted age and an alert on every cached start: the responder needs to know that restarts are handing the killed value back out, and needs a way to make the fleet stop doing it.
- Can the on-call engineer switch the behaviour during the outage?Only if a degraded mode was built with an input the platform can set without the store — an explicit, authorised control. Otherwise no: changing the start-up path means shipping a change, and the new processes that would carry it need the store to start. The decision was frozen when the code shipped.
A night porter who cannot reach the key board either lets nobody in until morning, or uses yesterday's spare — which works only if the locks were not changed overnight.
saying these in an interview costs you the question
- Treats failing closed as the safe default with no cost.
- Assumes a cached copy always works when the store is unavailable.
- Confuses rotation with withdrawal when judging whether the copy is usable.
- Accepts a cached copy without recording or bounding its age.
- Thinks the on-call engineer can flip the behaviour during the outage.
- Starts on a copy silently, with no alert that it happened.