skip to content

In Prefect, why might an on_failure hook on a flow never fire in production?

level: seniorimportance: should knowfreq 40%

answer

  1. it runs where the run runs
  2. one hook, one family of states
  3. a killed process cannot phone home
  4. the failure you most need told about

basics

~20 s

State hooks are client-side code that runs in the process owning the run. If that process is killed the hook never executes, and on_failure only covers Failed states — a Crashed or Cancelled run needs on_crashed or on_cancellation instead.

solid answer

~50 s

An `on_failure` hook is a normal Python callable that Prefect invokes **in the run's own process** after the run enters a Failed state, with `(flow, flow_run, state)`. Three things stop it firing. First, state coverage: `on_failure` is bound to Failed-type states only, so a run that was OOM-killed lands in Crashed and a cancelled run lands in Cancelled — both slip past unless you also register `on_crashed` and `on_cancellation`. Second, process liveness: if the container is SIGKILLed there is no interpreter left to run the callback at all, whatever hook you registered. Third, hooks are best-effort — an exception raised inside a hook is logged and does not change the run's state, so a broken alerting call fails invisibly. The durable answer for alerting is server-side automations reacting to state-change events, with hooks reserved for in-process cleanup and enrichment.

code

python · 16 lines
python
from prefect import flow, task

def notify(flow, flow_run, state):
    # best-effort: exceptions here are logged, not raised
    post(f"{flow_run.name} ended as {state.name}: {state.message}")

@task(retries=2, timeout_seconds=600)
def load(): ...

@flow(
    on_failure=[notify],
    on_crashed=[notify],       # OOM kills and evictions land here
    on_cancellation=[notify],  # deliberate stops land here
)
def pipeline():
    load()

go deeper

for a junior

Know that hooks are functions attached to a flow or task for a state transition, that they take a list of callables, and that a flow hook receives the flow, the flow run and the state.

for a middle

Explain the coverage map: on_failure fires for Failed-type states including TimedOut, while Crashed and Cancelled need their own hooks, and hook exceptions are logged rather than propagated.

for a senior

Demonstrate the operational reasoning — hooks run inside the process you are trying to be warned about, so paging belongs in server-side automations and hooks stay for artifacts, metrics and cleanup that may be skipped.

for a principal

Own the alerting architecture: which signals must survive infrastructure death, how stuck runs are detected at all, and the rule that nothing correctness-critical may live in a best-effort callback.

## What a hook is Prefect lets you attach callables to state transitions on a `@flow` or `@task`: ```python def alert(flow, flow_run, state): post_to_slack(f"{flow_run.name} -> {state.name}: {state.message}") @flow(on_failure=[alert], on_crashed=[alert], on_cancellation=[alert]) def pipeline(): ... ``` Flow hooks receive `(flow, flow_run, state)`; task hooks receive `(task, task_run, state)`. The available hooks mirror the state model — completion, failure, crash, cancellation, and running — and each takes a **list** of callables, so several can be registered. ## Reason one: the hook is bound to a state type `on_failure` fires for Failed-type states. That includes `TimedOut`, which is a Failed state under a more specific name — good news, since a hung task that blows its `timeout_seconds` does reach your alert. It does **not** include: - **Crashed** — the environment ended the run (OOM kill, pod eviction, spot reclaim, infrastructure that never started). Needs `on_crashed`. - **Cancelled** — someone stopped the run deliberately. Needs `on_cancellation`. This is the single most common production surprise: alerting is wired to `on_failure`, memory pressure starts killing containers, the runs go Crashed, and the on-call channel stays silent while the pipeline is down. ## Reason two: hooks run in the process that is dying Hooks are client-side. Prefect executes them in the same process that owns the run, after the state transition. That means: - A `SIGKILL` (OOM killer, `kill -9`, an evicted pod given no grace) leaves nothing to execute the callback. Even `on_crashed` cannot run locally if the machine is gone; in that case the Crashed state is recorded from outside, and no local code ran. - A graceful `SIGTERM` with enough shutdown time can let cancellation or crash handling run — which is why grace periods and handling shutdown properly are worth configuring. - A worker that loses network connectivity may execute the hook but fail to report anything. So the reliability of a hook is bounded by the reliability of the very process whose death you are trying to be told about. That is a circular dependency, and it is why crash-proof alerting cannot live only in hooks. ## Reason three: hooks fail quietly An exception inside a hook is logged and swallowed: it does not fail the run, does not retry, and does not change the state. That is deliberate — you do not want a flaky Slack webhook turning a successful pipeline into a failed one — but it means a hook that has been broken for a month looks exactly like a hook that never had anything to report. Hooks that matter deserve their own error handling, a timeout on any network call they make, and an occasional deliberate test. Hooks are also not retried and are not guaranteed to run exactly once in every pathological case, so treat them as best-effort notification and enrichment, never as the step that maintains correctness. Never put "release the lock", "commit the transaction" or "write the watermark" in a hook; put it in the task body where failure is visible and retried. ## What to do instead for alerting Prefect's server side emits **events** on state changes, and **automations** react to them: notify a channel, call a webhook, cancel or re-run something. Because that logic lives in the API rather than in the run's process, it still triggers when the process is gone — which is exactly the case you most need alerting for. The healthy split is: - **Automations** — "this deployment's run failed or crashed", "a run has been Running longer than N minutes", "no run completed in the last day". This is your paging path. - **Hooks** — in-process work that needs the run's own context: publishing a markdown artifact summarising what the failure found, closing a client, emitting a metric with local variables, writing a debug bundle. In Prefect 3, flow-run heartbeats give automations a way to notice a run that stopped reporting and mark it Crashed, closing the "stuck in Running forever" gap that hooks structurally cannot cover. ## Reviewing a hook setup Ask: which state types are actually covered? Does anything paging depend on the dying process? Would a hard kill of the container be noticed at all, and how fast? Is any correctness-critical step hiding in a hook? Does the hook itself have a timeout and a failure path? Most incidents in this area are not exotic — they are one hook, registered on one state type, in a system where the interesting failures happen in a different one.

  • Where should alerting live if hooks are best-effort?
    On the server side, as automations reacting to state-change events: run failed or crashed, run stuck in Running past a threshold, deployment produced no completed run today. That logic executes in the API rather than in the run's process, so it still fires when the process is gone — which is precisely the failure mode hooks cannot cover. Keep hooks for in-process enrichment.
  • What happens if the hook itself raises?
    The exception is logged and swallowed: the run's state is unchanged, the hook is not retried, and nothing else fails. That protects pipelines from flaky webhooks, but it also means a hook broken for weeks is indistinguishable from a quiet one. Give network calls in hooks a timeout and test them deliberately rather than assuming silence means health.
  • Is a hook a safe place to release a lock or write a watermark?
    No. Hooks are best-effort notification and enrichment, not part of the correctness path — they can be skipped entirely when a process is killed and they never retry. Anything that must happen belongs in a task body where failure is visible, retried and reflected in the run's state, or in an idempotent step the next run can repair.

saying these in an interview costs you the question

  • Wires all alerting to on_failure and calls it done
  • Expects a hook to run after a SIGKILL
  • Puts cleanup that must happen inside a hook
  • Thinks a raising hook will fail the run
  • Assumes cancellation triggers the failure hook

context