skip to content

The store is unreachable at 2am and a payments instance restarts beside a cached host copy — what does refusing to start buy over coming up on it?

level: seniorimportance: must knowfreq 66%

answer

  1. two failures, opposite directions
  2. refusing ratchets capacity down
  3. the copy needs the downstream to accept it
  4. withdrawn beats rotated-with-overlap
  5. bound the age, alert on every cached start

basics

~20 s

Refusing to start buys a loud, early failure and a guarantee that nothing serves on material that may already have been withdrawn. It costs capacity that never returns while the outage lasts, on a fleet that keeps losing instances.

solid answer

~50 s

The two behaviours fail in opposite directions. **Refusing to start** means the instance exits and takes no traffic: nothing runs on a credential that may have been withdrawn, the failure is early and in one place, and the instance count only ratchets downward for the length of the outage. **Coming up on the host copy** restores capacity, but the copy only works if the downstream system still accepts it — a rotation with an overlap window is survivable, a withdrawal is not, and in that case the instance starts, passes its health check, takes traffic and fails on first use, which is a worse shape than refusing. The copy also has to rest on the host to survive a restart at all. Accept it only up to a recorded age, and fail closed past that.

code

pseudocode · 24 lines
pseudocode
on startup:
    result = fetchFromStore(name) with bounded retry and backoff

    if result.ok:
        writeHostCopy(name, result.value, fetchedAt = now)
        serve(result.value)
        return

    copy = readHostCopy(name)

    if copy == MISSING:
        log "no credential and no local copy"
        exit(1)                      # fail closed

    if startPolicy == REFUSE_ON_CACHED:
        log "store unreachable; policy forbids a cached start"
        exit(1)                      # fail closed

    if now - copy.fetchedAt > maxAcceptedAge:
        log "local copy too old", age = now - copy.fetchedAt
        exit(1)                      # fail closed

    alert "starting on a local copy", age = now - copy.fetchedAt
    serve(copy.value)                # may already have been withdrawn

go deeper

for a junior

Know the two outcomes by name: the process either refuses to start, or starts on a value it kept from last time. They fail in opposite directions and both are choices someone made in advance.

for a middle

Explain why a cached start is not automatically a save: the value is only useful while the downstream system still accepts it, and a withdrawal turns a clean start-up failure into a failure after traffic arrives.

for a senior

Design the middle: bounded retry, a recorded fetch time, a maximum accepted age, and an alert on every cached start. Be able to say what your service does today and when it last did it.

for a principal

Decide it per tier rather than per team, and own the consequence: a fleet that refuses to start has staked its recovery on the store's, and a fleet that starts on copies has accepted a credential resting on every host.

## The two behaviours, stated precisely When the start-up fetch fails there are exactly two outcomes, and conflating them is the commonest error in this material. - **Failing closed** means the process **refuses to start**: it logs, exits non-zero, and the instance never takes traffic. - **Starting on cached material** means the process **comes up on a copy it already had**, written to the host by an earlier successful fetch, and serves on a value that may have been replaced or withdrawn while the store was unreachable. Both are legitimate. Neither is free, and which one a service should do is a property of that service, not a fact about secret management. ## What refusing to start costs During the outage the instance count is a one-way ratchet. Scale-outs, host replacements and crash loops all consume instances; nothing produces one, because nothing can start. A fleet with no headroom walks down toward a user-visible outage on a timetable set by someone else's system. The cost people only discover at 2am is that **you cannot change your mind**. Changing the start-up behaviour means shipping a change, and shipping it means starting new processes, which need the store. The decision was made when the code was written. What it buys is real: the failure is loud, early and in one place; nothing serves on a credential someone deliberately killed; and no process comes up with an empty value that some code path quietly treats as "no credential configured". ## What coming up on the copy costs - The copy only helps if the **downstream system still accepts it**. A cached start buys availability only inside the credential's remaining validity. - If the value was **rotated** during the window and both values are accepted for an overlap period, the instance works. If it was **withdrawn**, the instance starts, passes its start-up health check, takes traffic and fails on first use — worse than refusing, because it fails *after* accepting work. - If the value was withdrawn as part of an **incident response**, every cached start re-arms a credential somebody killed on purpose. - The copy must **rest on the host** to survive a restart at all. That is a standing exposure the fail-closed design simply does not have. - Unless the copy was written together with the time it was fetched, nobody can say how old it is — and its age is the whole question. | | Refuse to start | Start on the host copy | |---|---|---| | Capacity during the outage | Falls and stays down | Restored | | Serves on a withdrawn value | Never | Possible | | Where the failure shows | Start-up, immediately | First downstream call, after taking traffic | | Standing exposure | None added | A credential resting on every host | | Reversible at 2am | No | No | ## The bound that makes the middle option honest Most services should do neither purely. Build four things: 1. Write the copy with the **time it was fetched**, so its age is a fact rather than a guess. 2. Make the start-up fetch a **bounded retry with backoff** first, so a five-second blip costs nothing under either policy. 3. Define a **maximum age you will accept** at start-up; past it, fail closed. 4. **Log and alert on every cached start**, so "we came up on yesterday's copy" is something you know that night, not something you reconstruct later. ## Who chooses, and when The service owner chooses, at design time, and the choice is already encoded in the start-up path by the time the outage happens. An operator has no lever unless one was deliberately built — an explicit degraded-mode input the platform can set, which is itself a control somebody has to be authorised to use. The tiering is usually intuitive: a payments service whose owner would rather be down than run on a credential that may have been revoked; a read-mostly internal service whose owner would much rather it came up. ## The scenario, named At 2am the payments instance starts from a copy written by the last successful fetch, nine hours earlier. It survives if and only if the downstream account that credential belongs to still accepts it. If it is a long-lived static value nobody touched, it does. If it is a generated credential minted with a short validity, nine hours is almost certainly past its life, and the cached start buys nothing at all: the instance comes up, takes traffic and fails — the worst of both designs.

  • Under what condition does the cached start actually work, and under what condition does it not?
    It works while the downstream system still accepts that value. A rotation with an overlap window leaves the old value accepted, so the instance serves. A withdrawal — the value revoked, the downstream account disabled — leaves it useless, and the instance discovers that only on its first real call, after it has already been given traffic.
  • Why is a bounded retry at start-up worth building whichever policy you choose?
    Because most store unavailability is seconds, not hours. A short retry with backoff absorbs blips, partial failovers and restarts of the store itself, so neither policy is ever invoked for a five-second event. It also separates "could not reach the store right now" from "cannot reach the store at all", which are different decisions.
  • The value was revoked deliberately during an incident. What does a cached start do then?
    It re-arms the revoked credential on every instance that restarts, which is the opposite of what the responder intended. This is the case that argues for a maximum accepted age and an alert on every cached start: the responder needs to know that restarts are handing the killed value back out, and needs a way to make the fleet stop doing it.
  • Can the on-call engineer switch the behaviour during the outage?
    Only if a degraded mode was built with an input the platform can set without the store — an explicit, authorised control. Otherwise no: changing the start-up path means shipping a change, and the new processes that would carry it need the store to start. The decision was frozen when the code shipped.

A night porter who cannot reach the key board either lets nobody in until morning, or uses yesterday's spare — which works only if the locks were not changed overnight.

saying these in an interview costs you the question

  • Treats failing closed as the safe default with no cost.
  • Assumes a cached copy always works when the store is unavailable.
  • Confuses rotation with withdrawal when judging whether the copy is usable.
  • Accepts a cached copy without recording or bounding its age.
  • Thinks the on-call engineer can flip the behaviour during the outage.
  • Starts on a copy silently, with no alert that it happened.