skip to content

When a whole fleet must restart during a dependency outage, how should its startup sequence and readiness be designed?

level: principalimportance: should knowfreq 38%

answer

  1. boot checks are dependencies in disguise
  2. cold fleet, empty caches, aligned retries
  3. exit only on what waiting cannot fix
  4. not-ready beats exiting for transient failures
  5. jitter and stagger the recovery

basics

~20 s

Design the boot path to complete without its dependencies: verify settings and wiring eagerly, but retry reachability in the background while reporting not-ready. A service that exits when a dependency is down cannot rejoin on its own.

solid answer

~50 s

Every eager check on the start path is a **boot-time dependency**: it decides that your service cannot start while that system is down. One such check is reasonable; a fleet of them produces a start-up ordering graph nobody wrote down, and a cold start becomes a sequencing exercise during the worst hour of the year. The policy that survives is to split verification by **what can heal**. Settings, wiring and credentials formats are deterministic — verify them eagerly and exit, because no amount of waiting fixes them. Reachability of another system is transient — attempt it with a bounded, jittered retry while the process stays up and reports **not ready**, so it joins the fleet the moment the dependency returns, with no redeploy and no human in the loop. Then keep the boot itself cheap, because recovery time is dominated by how long one instance takes to become useful.

go deeper

for a junior

Understand that a service which refuses to start when something else is down cannot come back on its own. Starting up and waiting is usually safer than exiting.

for a middle

Be able to sort checks into those a retry can fix and those it cannot, and explain why the second group belongs on the exit path and the first does not.

for a senior

Talk about what a cold fleet does to downstream load, why aligned retries stampede, and how jitter, staggered waves and a warm-up floor shorten real recovery time.

for a principal

Own it as policy: the boot-time dependency graph, the cycles it can contain, the deploy gate and alerting that keep a soft failure loud, and the explicit trade of deploy-time loudness for recoverability.

Most startup-lifecycle advice is written for the happy case: one instance, starting, while everything else is healthy. The decisions that actually matter are the ones that show up when **the whole fleet is cold at once** — a regional failure, a bad release rolled back, a platform incident that restarted every instance. At that moment the startup sequence you designed becomes your recovery plan. ## The failure mode to design against A cold fleet is different from a rolling deploy in three ways: - **Nothing is warm.** Caches are empty and pools are unconnected, so downstream systems see a load pattern they never see in steady state — often several times normal read volume. - **Everything starts simultaneously.** Retries and first connections align, and aligned retries are a stampede. - **Dependencies may still be down.** The event that caused the restart is often the one that broke the dependency, so boot happens under exactly the conditions an eager check fails in. An eager check that exits the process turns that third point into a loop: the instance exits, is replaced, checks again, exits again — adding restart churn and, worse, adding load to the very system that is struggling. ## Split verification by whether waiting can fix it | Class of check | Can waiting fix it? | Boot policy | |---|---|---| | Missing or incoherent settings | no | verify eagerly, exit with the offending key named | | Missing wiring, unresolvable component | no | resolve the graph eagerly, exit | | Malformed credential or key material | no | verify format eagerly, exit | | A required system unreachable | yes | bounded retry in the background, stay up, report not-ready | | Cold cache, empty pool | yes | warm up after the bind, report ready when useful | The rule is one sentence: **exit only on defects that no amount of time will repair.** Everything else becomes a not-ready state, which is strictly more capable — no traffic is routed, the condition is visible, and recovery is automatic. ## Keep the start path cheap Recovery time for the fleet is roughly the time for one instance to become useful, plus however much you have to stagger. That makes boot duration a capacity decision, not a developer-comfort one: 1. **Move optional work off the start path.** Anything the service can do lazily on first use should not block readiness. 2. **Warm up to a floor, not to steady state.** Open a minimum pool and load the hot slice of cache; let the rest fill under real traffic. 3. **Jitter everything.** Randomise retry delays and warm-up start so a thousand instances do not align on the same second. 4. **Stagger the restart itself.** Bringing a fleet back in waves keeps downstream load survivable and gives caches time to become useful for the next wave. ## The ordering graph nobody wrote down Every boot-time dependency is an implicit edge in a start-up ordering graph. Two or more edges can form a cycle — service A refuses to start without B, B without A — and a cycle means the system cannot cold-start at all without someone manually breaking it. The cheapest defence is architectural rather than operational: make the edges *soft*. A service that starts, reports not ready, and retries has an edge that resolves itself; a service that exits has an edge that requires sequencing. Auditing which of your services can complete a boot with every dependency unavailable is a short exercise that pays for itself the first time it matters. ## What to trade away, deliberately This is a judgment question, so name the costs rather than pretending there are none. Soft edges mean a genuinely broken deployment can sit not-ready for a long time instead of failing a deploy loudly — so pair them with alerting on “ready never reached” and with a deploy gate that refuses to continue a rollout when new instances are not becoming ready. Lazy warm-up means the first minutes after recovery are slower — acceptable when the alternative is no capacity at all. And a bounded retry needs a real bound, because a process that retries forever while silently serving nothing is indistinguishable from one that is stuck. ## What an interviewer is listening for That you treat startup checks as a policy with fleet-wide consequences rather than a per-service preference; that you can state the exit-versus-not-ready rule and defend it; and that you mention the unglamorous parts — jitter, staggering, and an alert for instances that never become ready.

  • How would you find out whether your services can actually cold-start during an outage?
    Test it rather than reason about it: start a service in an environment where its dependencies are unreachable and see whether it completes boot and reports not-ready, or exits. Doing this per service produces the boot-time dependency graph directly, and any cycle it reveals is a system that cannot recover unattended.
  • What is the risk of letting a service start and stay not-ready instead of exiting?
    A genuinely broken build can sit quietly not-ready instead of failing the deploy. Guard it with an alert on instances that never reach ready within a bound, and a rollout gate that stops advancing when new instances are not becoming ready, so a soft failure is still loud.
  • Why does an empty cache change the capacity calculation during recovery?
    Steady-state load on downstream systems is measured with caches already filled. A cold fleet asks for everything it would normally have cached, often several times normal volume, at the moment the dependency is least healthy — which is why staggered waves and a warm-up floor matter more than raw start speed.

saying these in an interview costs you the question

  • Adds eager reachability checks for every dependency by default
  • Exits the process whenever any dependency is unreachable at boot
  • Assumes steady-state capacity numbers apply to a cold fleet
  • Restarts an entire fleet at once with no stagger or jitter
  • Never tests whether a service can boot with dependencies down
  • Treats boot-time ordering between services as acceptable forever