skip to content

A long-lived component supervises a set of worker tasks that can crash. Walk through the strategies available when a child dies — resume, restart, stop, escalate — and explain how you decide among them.

level: principalimportance: should knowfreq 38%

answer

  1. resume / restart / stop / escalate
  2. restart = collapse unknown state to known
  3. one-for-one vs all-for-one
  4. budget N in T, backoff + jitter, quarantine poison
  5. error kernel: state high, risk in leaves

basics

~20 s

A supervisor watches its children's outcomes and applies a policy: resume (keep state, log), restart (discard possibly corrupt state, start from a known-good initial state), stop (the child is no longer viable), or escalate (fail itself so a higher supervisor recovers a bigger subtree). Bound restarts with a budget and backoff.

solid answer

~1 min

Supervision separates the code that does work from the code that decides what happens when work dies. - **Resume / ignore** — the child keeps its state and continues. Safe only for errors that are expected, isolated, and cannot have corrupted state; otherwise you preserve the corruption. - **Restart** — discard the child's state and recreate it from a known-good initial state. The default for stateful children, precisely because after an unexpected failure you cannot trust the state. Any state that must survive lives outside the child (durable store, or a parent that owns it). - **Stop** — the child is unrecoverable or no longer needed; remove it and adjust capacity or fail over. - **Escalate** — the supervisor cannot fix it locally, so it fails itself and lets its own parent restart the larger subtree. This is how a system converges on "restart a big enough unit" without any component knowing the global picture. Decide by scope of damage: local and transient → resume or restart; shared invariant broken → restart the group (all-for-one); repeated → escalate. Bound with a **restart budget** (max N in T, then escalate) plus exponential backoff with jitter, and quarantine poison inputs so restarts do not replay the same crash. Keep critical state high in the tree and risky work in cheap-to-restart leaves. Alert on restart *rate*: supervision restores availability, it does not fix bugs.

code

text · 13 lines
text
on child_terminated(child, outcome):
    if outcome is CANCELLED: return              # we asked for it; not a failure
    if outcome is EXPECTED_ERROR: log; resume    # contained, state intact

    record_restart(child)
    if restarts_in_window(child) > N_IN_T:
        escalate()                               # let the parent recover a bigger unit
    else:
        if siblings_share_state(child): stop_all(); start_all()   # all-for-one
        else: start_fresh(child, delay=backoff_with_jitter())     # one-for-one

    if crashed_on_same_item(child) >= K:
        quarantine(item)                          # don't replay the poison input

go deeper

for a junior

Name the four strategies and say restart is the common default because a crashed child's state can no longer be trusted.

for a middle

Add one-for-one versus all-for-one, restart budgets with backoff, and the distinction between expected domain errors handled in the child and defects that reach the supervisor.

for a senior

Cover poison-input quarantine, escalation as blast-radius search, health checks that prevent a failing start from resetting the budget, alerting on restart rate, and durable at-least-once delivery with idempotent handlers for in-flight work.

for a principal

Design the tree: place critical state in a small, trivially correct error kernel with risky work in cheap-to-restart leaves, choose group semantics from the shared-state graph, define escalation ladders up to process restart, and set the operational contract that restarts mitigate but never resolve.

## What supervision is A supervisor is a component whose job is not the work but the **policy for failure of the work**. It owns a set of children, observes each child's terminal outcome, and decides what to do next. It is failure propagation made deliberate: instead of every failure travelling to the top, a supervisor intercepts it at a chosen boundary and either contains it or passes it up. This applies to long-lived children — consumers, pollers, connection handlers, per-entity workers — as opposed to a bounded fan-out where a completion policy combines results into one answer. ## The strategies **Resume / ignore.** The child continues with its existing state; the error is logged and counted. Use only when the error is *expected and local*: a single malformed message in a stream, one bad row in a batch. The precondition is that the failure cannot have left the child's own state inconsistent — which in practice means the failure happened at a clean boundary. Resuming after an unexpected defect preserves whatever corruption caused it, producing a component that keeps failing in new ways. **Restart.** Terminate the child and create a fresh instance from a known-good initial state. This is the workhorse strategy, and the reason is epistemic: after an unexpected failure you do not know what the in-memory state means, and reasoning about every possible partially-updated configuration is impossible. Restarting collapses that unbounded set of unknown states to one known state. It has design consequences — any state that must survive a restart cannot live only inside the child; it belongs in a durable store or in the parent that recreates the child. **Stop.** The child is not viable (its configuration is invalid, its downstream is permanently gone) or is no longer needed. The supervisor removes it and adjusts: reduce declared capacity, drain its partition, hand its work to a sibling, mark the feature degraded. **Escalate.** The supervisor decides it cannot restore a good state at its level and fails itself, so its own parent applies *its* policy to the larger subtree. Escalation is how a system finds the right blast radius without any single component knowing the whole topology: restart the worker; if that keeps failing, restart the pool; if that keeps failing, restart the component; ultimately restart the process, where the OS reclaims everything. ## Group semantics When a child dies, does that affect its siblings? - **One-for-one:** restart only the failed child. Right when children are independent (one worker per connection or per partition). - **All-for-one:** restart the whole sibling group. Right when siblings share state, a protocol, or a handshake — a reader and a writer over one connection, a pipeline whose stages hold matched buffers. Restarting one alone leaves the survivors talking to a partner that no longer exists. - **Rest-for-one:** restart the failed child and everything started after it, when there is a startup ordering dependency. ## Bounding the damage Without limits, restart is a way to burn CPU forever: - **Restart budget/intensity:** at most N restarts within T seconds; exceeding it escalates. This is what converts "crash-loop" into "escalate to a bigger recovery". - **Exponential backoff with jitter:** a child whose dependency is down should retry at increasing intervals, jittered so a fleet does not synchronise into a thundering herd against a recovering dependency. - **Poison-input quarantine:** if the child crashes on a specific input, restarting replays the same input and crashes again. Track attempts per item, and after k failures move it to a dead-letter/quarantine and continue. Without this, one bad message halts a queue indefinitely. - **Health-based stopping:** a child that starts and immediately fails a readiness check should be treated as failed, not counted as a success that resets the budget. ## The error kernel A structural principle: keep the state whose loss is unacceptable **high in the supervision tree**, in components whose code is small and trivially correct, and push risky work — parsing untrusted input, talking to flaky dependencies, executing user-supplied logic — into **leaves that are cheap to restart**. Then routine failure destroys only recreatable things. A design where the risky parser holds the only copy of the session state has no viable supervision policy at all, because every strategy either loses the state or preserves the corruption. ## Expected errors versus defects Supervision is for **defects** — the unexpected, the unhandled, the "should not happen". Expected domain outcomes (item not found, validation failed, downstream returned 429) should be returned as values and handled inside the child. Crashing a worker for a validation failure turns a normal outcome into a restart storm; conversely, catching a genuine defect and continuing turns a diagnosable crash into silent corruption. ## What supervision does not give you - **Correctness.** Restarting restores availability. Whatever the child was doing at the moment it died is lost unless the work is durable and idempotent — which is why supervised workers usually sit behind a queue with at-least-once delivery and an idempotent handler, so recovery replays the item rather than dropping it. - **A fixed bug.** A supervision tree that hides a defect behind ten restarts a minute is a system with an unreported outage. Restart rate is a first-class metric and should alert; the restart is the mitigation, not the resolution. - **Consistency across components.** Restarting one child while siblings hold references to its old identity requires an explicit re-registration/reconnection protocol, or all-for-one semantics. ## Making the decision A workable decision procedure: *Is the error expected and contained?* → handle in the child, no supervision event. *Is it unexpected but the child's state is recreatable?* → restart, with budget and backoff. *Do siblings depend on that state?* → restart the group. *Is it recurring past the budget, or is the invariant broader than this child?* → escalate. *Is the child no longer viable or needed?* → stop and degrade explicitly.

  • Why is restart-from-initial-state usually preferred over catching the error and continuing with the existing state?
    Because after an unexpected failure you cannot characterise the in-memory state: any subset of a multi-step update may have applied, and enumerating those possibilities is intractable. Restarting collapses that unbounded space of unknown states to a single known-good one, which is the only state the code was actually designed and tested for. The trade is that any state that must survive has to live outside the child, in durable storage or in the parent.
  • How do you stop a supervision tree from masking a real bug?
    Treat restart rate as a first-class signal: emit a metric per child with the failure classification, and alert when restarts exceed a normal background rate rather than only when the process dies. Pair it with restart budgets so persistent failure escalates and becomes visible instead of looping quietly, and with poison-input quarantine plus dead-letter inspection so the specific bad inputs are collected for diagnosis instead of being replayed forever.
  • How does supervision interact with work that was in flight when the child died?
    It does not recover it by itself — a restarted child begins from a known-good state with no memory of the message it was processing. Durability has to come from outside: an at-least-once queue that redelivers unacknowledged items, plus an idempotent handler so redelivery is safe. Supervision restores the capacity to process work; the delivery guarantee restores the work itself.

A ship's damage-control policy: patch the leak and carry on, seal and re-flood the compartment, abandon it entirely, or call the bridge because the decision is above your pay grade.

saying these in an interview costs you the question

  • Restarts endlessly with no budget or backoff, producing a crash loop
  • Resumes after an unexpected defect, preserving whatever state corruption caused it
  • Keeps the only copy of critical state inside the child most likely to crash
  • Crashes a worker for expected domain outcomes such as validation failures
  • Replays the same poison input after every restart with no quarantine or dead-letter path
  • Assumes restarting recovers the in-flight work rather than only the capacity to do work

context