skip to content

Explain supervision trees and the 'let it crash' philosophy in actor systems: who handles a failed actor, what options a supervisor has, and what this buys compared with catching every exception in place.

level: seniorimportance: should knowfreq 38%

answer

  1. parent = supervisor, decides one level up
  2. restart / resume / stop / escalate
  3. one-for-one vs all-for-one
  4. error kernel: valuable state up, risky work down
  5. backoff + restart limits or restart storm

basics

~20 s

Each actor has a parent that supervises it. On failure the child stops and its parent decides: restart it to a known-good state, resume it, stop it, or escalate to its own parent. Recovery lives in the tree, not scattered through handlers, so failure is isolated and state cannot stay corrupt.

solid answer

~60 s

Actors are created by other actors, so they form a tree, and each parent is the **supervisor** of its children. When a handler throws, the actor does not patch itself up: it stops, and the failure is reported to the parent, which chooses a strategy — **restart** (discard the possibly corrupt state and re-initialize to a known-good baseline), **resume** (keep state, skip the bad message), **stop** (this child is not needed), or **escalate** (I cannot decide; my own parent should handle it). Strategies also come in *one-for-one* (only the failing child) and *all-for-one* (siblings share fate because they are interdependent) flavours. 'Let it crash' means handlers cover the expected business cases and let unexpected failures kill the actor, because after an unexpected exception you cannot trust the state. The gains: recovery logic is centralized and explicit instead of scattered defensive catch blocks; failure is contained to one actor's state; and the **error kernel** pattern lets you keep critical state high in the tree and push risky work into cheap, disposable leaf actors. The costs to manage are restart storms (use backoff and limits), lost in-flight state (persist or re-request), and poison messages.

go deeper

for a junior

Know that actors form a parent-child tree, that a failing actor is handled by its parent, and that restart returns it to a clean known state.

for a middle

List the four strategies plus one-for-one versus all-for-one, and explain 'let it crash' as untrusted-state reasoning rather than skipping error handling.

for a senior

Add the error kernel pattern, restart storms with backoff and limits, poison-message handling, and how state is recovered when it must survive.

for a principal

Treat the supervision hierarchy as an availability design: what is in the error kernel, what the blast radius of each subtree is, when to escalate versus degrade, and how restart rate is monitored so self-healing does not hide an outage.

## The structure: a tree, not a flat pool Actors are created by other actors, so the runtime naturally forms a hierarchy. The creator is the **parent** and **supervisor** of its children. Supervision is not an optional add-on — it is how failure is defined in the model. An actor is responsible for its own work; deciding what to do when it fails is *someone else's* job, specifically its parent's. That single inversion is the whole idea. ## 'Let it crash' In a lock-based, exception-catching style, code defends locally: wrap the risky call, log, try to continue. The problem is that after an unexpected exception you generally do not know what your state is — a half-updated aggregate, a partially consumed stream, an invariant broken between two fields. Continuing means operating on data you cannot trust. 'Let it crash' says: handle the failures you *expect* and modelled (validation errors, a business rule violation, a known remote 404) as ordinary results; let the ones you did not expect terminate the actor. Because the actor's state is private and isolated, its death destroys exactly the suspect state and nothing else. Reinitialization then restores a state you *can* trust. This is the same reasoning as restarting a process, applied at a granularity of microseconds and kilobytes. ## What a supervisor can do When a child fails, the supervisor receives a failure notification and chooses: - **Restart** — stop the child, discard its state, create a fresh instance with the same address so senders are unaffected. The default for transient corruption. - **Resume** — keep the existing state and continue with the next message, skipping the one that failed. Appropriate only when you are confident the failure did not touch state, e.g. a bad input. - **Stop** — terminate the child permanently; the work is not needed or cannot succeed. - **Escalate** — the supervisor declares itself unable to handle this class of failure and fails in turn, handing the decision to *its* parent. This is how a leaf-level 'database credentials invalid' turns into a system-level decision. Orthogonally, the strategy has a scope: **one-for-one** affects only the failing child, while **all-for-one** restarts the whole sibling group, which is right when children share a dependency or a session and a partial restart would leave the group inconsistent. The failing message itself is normally *not* redelivered after a restart, precisely so a poison message cannot loop the actor to death. ## The error kernel pattern The design consequence is architectural: keep valuable, hard-to-reconstruct state near the **root** of the tree, where nothing risky is executed, and push risky work (parsing untrusted input, calling flaky remote systems, running plug-ins) down into disposable leaf actors. Then a failure destroys a leaf whose state costs nothing to rebuild, while the important state never sits in the blast radius. Choosing what belongs in the error kernel is the real design work. ## What it buys - **Centralized, explicit recovery policy.** Instead of hundreds of ad-hoc catch blocks, the policy is declared in one place per subtree and can be reviewed. - **Containment.** One actor's failure does not corrupt others because there is no shared mutable state to corrupt. - **Self-healing.** Transient failures — a dropped connection, a temporary dependency outage — resolve without operator involvement. - **Honest code.** Handlers describe the success path plus modelled errors, which is far more readable than defensive code that pretends to recover from the unknown. ## What it costs, and how to manage it - **Restart storms.** A permanently broken dependency makes an actor crash and restart in a tight loop, burning CPU and flooding logs. Supervisors must apply *exponential backoff* and a limit ('N restarts within T, then escalate or stop'). - **Lost in-flight state.** Restarting discards the state, which is the point, but if that state was needed it must come from somewhere: event sourcing / snapshot recovery, a re-request to the source of truth, or a design where the state is reconstructible from the next messages. - **Poison messages.** A message that reliably kills the actor will keep doing so if it is redelivered. Drop the failing message, or route repeat offenders to a dead-letter destination for inspection. - **Observability.** Crash-and-restart is invisible if nobody counts it. Restart rate is a first-class metric; a silently self-healing system can hide a real outage. ## Interview delivery Describe the parent-child tree, state that recovery is the parent's decision, list the four strategies plus one-for-one versus all-for-one, explain 'let it crash' as a state-trust argument rather than laziness, then show the error kernel design consequence and finish with the operational caveats: backoff, restart limits, state recovery, poison messages and restart-rate alerting.

  • When is 'resume' the right supervisor strategy instead of 'restart'?
    Only when you are confident the failure could not have left the actor's state inconsistent — typically a malformed input that was rejected before any state mutation. Resume keeps the accumulated state and simply moves on to the next message, which is valuable when that state is expensive to rebuild. If there is any doubt about which mutations completed, restart is the safe choice.
  • How do you stop a supervised actor from restarting in a tight loop when its dependency is down?
    Give the supervisor a restart limit with a time window — for example at most five restarts in a minute — and exponential backoff between attempts, so retries slow down instead of spinning. When the limit is exceeded, the supervisor escalates or stops the subtree so the failure becomes visible rather than being absorbed. Restart rate should also be a monitored metric with an alert.
  • An actor holds state that took ten minutes to build. Does 'let it crash' throw that away?
    By default yes, which is why such state should not sit in a risky actor. Either move it up into the error kernel and push the risky work into a child, or make it recoverable — event sourcing with periodic snapshots, or rebuilding from a source of truth on restart. Choosing between those is exactly the design decision the supervision hierarchy forces you to make explicitly.

A crashed actor is like a jammed machine on a production line: the operator does not disassemble it mid-shift, they stop it and swap in a known-good unit. The supervisor decides whether to swap one machine, the whole cell, or call the plant manager.

saying these in an interview costs you the question

  • Treating 'let it crash' as an excuse to skip error handling entirely, including for expected business errors
  • Believing a restart preserves the actor's in-memory state
  • Configuring unlimited restarts with no backoff
  • Catching every exception inside the handler and continuing on state that may be inconsistent
  • Assuming the failing message is redelivered after a restart

context