skip to content

In Prefect, what distinguishes a Crashed state from a Failed state?

level: middleimportance: must knowfreq 60%

answer

  1. ask who ended the run
  2. one has a traceback, one has an exit code
  3. the process never got to say goodbye
  4. retries live in a process that no longer exists

basics

~20 s

Failed means the run executed and its code raised or reported an error, so Prefect recorded the exception. Crashed means the run was cut off by its environment — process killed, container evicted, out of memory — so it never reported its own outcome.

solid answer

~50 s

Both are terminal states, but they say different things about *who* ended the run. **Failed** is an in-band outcome: the task or flow actually ran, an exception propagated, and Prefect captured it — so the retry logic you configured applies, and `on_failure` hooks fire. **Crashed** is out-of-band: the run's process disappeared or was killed by infrastructure — SIGKILL, an OOM kill, a pod eviction, a spot instance reclaim — so there was no orderly reporting. Because the retry machinery lives in the process that vanished, a crash usually is not retried by the run itself; recovery means an external mechanism re-running it, and the hook that applies is `on_crashed`, not `on_failure`. Operationally the distinction is a triage shortcut: a wall of Failed runs points at your code or the data, while Crashed runs point at memory limits, node churn, or the platform.

code

text · 7 lines
text
Flow run 'amber-otter'  Running -> Failed
  ValueError: column 'order_ts' not found
  -> AwaitingRetry (1/3), retry in 30s

Flow run 'brave-newt'   Running -> Crashed
  Flow run infrastructure exited with non-zero status code 137
  (no traceback; container OOM-killed)

go deeper

for a junior

Recall that both states are terminal, that Failed means the code raised an error Prefect captured, and that Crashed means the environment killed the run before it could report anything.

for a middle

Explain the consequences: retries and on_failure hooks key off Failed, crashes bypass in-process recovery entirely, and Prefect state names such as TimedOut or Cached sit on top of a smaller set of state types.

for a senior

Demonstrate triage — a batch of Crashed runs sends you to memory limits, node churn and image pulls, not tracebacks — plus how zombie runs and leaked concurrency slots are detected and cleared.

for a principal

Own the reliability posture: where alerting must live so it survives the process dying, how much work a re-run may safely repeat, and what infrastructure sizing and preemption policy the crash rate is really telling you.

## Prefect's state model in one paragraph Every flow run and task run in Prefect is represented by a **state** object, and orchestration is just the server deciding which state a run moves to next. A state has a *type* (a small enum: Scheduled, Pending, Running, Completed, Failed, Crashed, Cancelling, Cancelled, Paused) and a *name*, which can be more specific than the type. `Cached` is a state name whose type is Completed; `AwaitingRetry` is a Scheduled state; `Retrying` is a Running state; `TimedOut` is a Failed state; `Late` is a Scheduled run that missed its expected start. Reading UI badges as "types with nicknames" removes most of the confusion. ## Failed: the run told you it went wrong A run is Failed when the code executed and produced an error in band. Concretely: your task raised an exception, or a flow returned/raised something Prefect interprets as failure. Prefect captures the exception, attaches it to the state's message, and the normal machinery applies: - If the task declares `retries`, the run moves to `AwaitingRetry` and is retried after `retry_delay_seconds`. - Once retries are exhausted, the state settles as Failed. - `on_failure` hooks run. - A `TimedOut` run (a task or flow that blew past `timeout_seconds`) is a Failed-type state, so it participates in the same handling. Failed is the state you *want* for anything recoverable, because everything Prefect knows how to do about errors keys off it. ## Crashed: the environment ended the run A run is Crashed when the execution environment removed it without letting it report. Typical causes: - The container hit its memory limit and the kernel OOM-killed the process. - A Kubernetes pod was evicted, or the node was drained. - A spot/preemptible VM was reclaimed. - Someone killed the worker process, or the machine rebooted. - The infrastructure never came up at all — an image pull failure or a bad job definition — so the flow-run process could not start. The defining property is that nothing in your process got a chance to run cleanup. That has consequences people find surprising: 1. **Your retries probably did not apply.** Task retries are executed by the process that owns the run. If that process is gone, there is nobody left to sleep and try again. Recovery for crashes comes from outside — re-running the flow run or deployment, or an automation that reacts to the Crashed event. 2. **Local hooks are unreliable.** An `on_crashed` hook is still client-side code. If Prefect is able to detect and record the crash while some client is alive, it can fire; if the machine vanished, nothing local runs and the Crashed state is set by the API side of the system. 3. **Held resources can leak.** A run holding a concurrency slot when it is killed may leave that slot occupied until it is reset. ## How a crash actually gets recorded Sometimes the run's own process sees the signal (a SIGTERM handled during shutdown) and reports Crashed itself. Sometimes there is no process left, and the state has to come from the outside: Prefect 3 flow runs emit heartbeats, and you can build an automation that marks a run whose heartbeats stopped as Crashed, so it does not sit in Running forever. Without something like that, an abruptly killed run is a *zombie*: the database says Running, reality says nothing is executing. ## Cancelled and Paused, for completeness Cancellation is deliberate: a run moves to Cancelling while Prefect tries to stop the infrastructure, then Cancelled. It is neither Failed nor Crashed, and `on_failure` does not fire for it — you need `on_cancellation`. Paused runs are suspended awaiting input or a resume, and they can time out back into a terminal state. ## Why interviewers ask this The distinction is a fast diagnostic. If a deployment's runs are mostly Failed, read the tracebacks: it is your code, a schema change, a credential, bad data. If they are mostly Crashed, stop reading tracebacks — there usually isn't one — and look at memory limits, node lifecycle, image availability, and worker health. Candidates who treat the two as synonyms tend to spend an afternoon debugging application code for what was an OOM kill. ## Practical hardening Give tasks realistic `timeout_seconds` so a hang becomes a TimedOut failure you can retry rather than a hung slot; size memory for the real payload rather than the happy path; persist results so an externally re-run flow can reuse completed work instead of recomputing everything; and put the alerting that must survive a crash on the server side, as an automation on state-change events, rather than only in in-process hooks.

  • Does a Crashed task get retried by its retries setting?
    Generally no. Retry scheduling is carried out by the process that owns the run, and a crash means that process is gone, so nothing local is left to wait and try again. Recovery comes from outside: re-running the flow run or deployment, or an automation that reacts to the Crashed event. Design for it by persisting results so a re-run resumes rather than recomputing.
  • What is a zombie run and how do you stop the UI filling with them?
    A zombie is a run the database still shows as Running while nothing is actually executing, because the process died without reporting. Prefect 3 flow runs emit heartbeats, so an automation can mark runs whose heartbeats stopped as Crashed. Combining that with realistic timeout_seconds keeps stuck and dead runs from occupying dashboards and concurrency slots indefinitely.
  • Which state type does a TimedOut run have?
    TimedOut is a state name whose type is Failed, so it behaves like any other failure: retries apply if configured, and on_failure hooks fire. That is deliberate — a task that blew past timeout_seconds was still under Prefect's control, unlike a Crashed run where the environment removed the process outright.

saying these in an interview costs you the question

  • Treats Crashed as a synonym for Failed
  • Expects task retries to recover an OOM-killed run
  • Hunts for a traceback on a Crashed run
  • Thinks on_failure hooks cover crashes and cancellations
  • Believes a Running badge proves something is executing

context