Which Airflow task instance states do you check first when a DAG run appears stuck?
answer
- read the colour before the log
- one state means no slot, another means no permission
- a queued task has no log yet
- sensors have a state of their own
- the earliest blocked task, not the last
basics
~20 sRead the task instance state in Airflow's grid view: scheduled or queued means no worker slot, running means the work is live, up_for_retry means it is waiting between attempts, and none or upstream_failed means it was never eligible. The state names the layer to investigate.
solid answer
~50 sAirflow's task instance state is the fastest diagnostic because each state points at a different layer. `none`/`scheduled` means the scheduler has not released it — check `depends_on_past`, `wait_for_downstream`, trigger rules and whether the DAG is paused. `queued` means the scheduler released it but no worker picked it up — check pool slots, `max_active_tasks`, queue routing and worker health. `running` with no log progress means the work is genuinely live but blocked downstream, or it is about to be a zombie. `up_for_retry` means it failed and is waiting out `retry_delay`. `up_for_reschedule` means a sensor in reschedule mode is sleeping between pokes. `upstream_failed` and `skipped` mean it was never eligible, which redirects you to the predecessor. So the drill is: read the state, then look at the layer that state implicates — scheduler constraints, capacity, or the process itself.
code
text · 9 linesgrid view, run scheduled__2024-03-11:
extract_orders success
wait_for_partner running <- poke-mode sensor, 5h, holds a slot
transform_orders queued <- pool 'default_pool' has 0 free slots
publish_mart scheduled <- upstream not done
notify none
diagnosis: the sensor is not stuck; it is occupying the slot
transform_orders needs. mode='reschedule' frees it.go deeper
Recall the main states — scheduled, queued, running, up_for_retry, failed, skipped, upstream_failed — and that the grid view shows them per task instance.
Explain what each state implies about which component is responsible: scheduler eligibility, executor capacity, the running process, or an upstream decision.
Run the triage for real: start at the earliest blocked instance, read scheduler logs for a queued task, spot poke-mode sensors eating a pool, and connect depends_on_past to a stalled chain of runs.
Own the instrumentation that removes the drill — state-count and queued-duration metrics, pool utilisation dashboards, a mandatory execution_timeout policy, and capacity planning so queued is never the normal state.
## Why state-first triage works A 'stuck DAG' is not one problem. It could be scheduler-side (the task was never released), capacity-side (released but nothing ran it), execution-side (running but blocked), or dependency-side (never eligible in the first place). Airflow's task instance state distinguishes all four before you read a single log line, which is why experienced operators open the grid view and read colours rather than opening logs. ## The states and what each implicates **`none` (no state yet) / `scheduled`** — the scheduler either has not considered the instance or has decided it is eligible but has not queued it. If instances sit here: - Is the DAG **paused**? A paused DAG creates no new runs. - Is `depends_on_past=True` with the previous run's instance failed or still running? That pins every subsequent instance of that task. - Is `wait_for_downstream=True` holding it behind the previous run's downstream tasks? - Is the `trigger_rule` unsatisfiable given upstream states? - Is `max_active_runs` already saturated by earlier runs, so this run cannot start? **`queued`** — the scheduler handed it to the executor and nothing has claimed it. This is a **capacity or routing** problem: - Is the task's **pool** out of slots? A pool with 4 slots and a task requesting `pool_slots=4` blocks behind anything else in that pool. - Is `max_active_tasks` (DAG-level concurrency) or the task's own concurrency limit reached? - Is the task assigned to a **queue** no running worker is consuming? A Celery task routed to `queue='gpu'` with no GPU worker online waits forever with no error. - Are the workers alive at all? Celery workers down, or a Kubernetes cluster that cannot schedule the pod (insufficient resources, an image pull failure) both present as an instance parked in queued. **`running`** — the process exists and is heartbeating. Now the question is what it is waiting on: a lock in the warehouse, a query that has been running for hours, an HTTP call with no timeout, a sensor in poke mode occupying its slot. Read the task's logs and, separately, look at the *downstream* system — a task 'stuck' in running is usually a symptom of something outside Airflow. If the heartbeats have actually stopped, the scheduler's zombie sweep will fail it shortly. **`up_for_retry`** — the attempt failed and the instance is waiting out its `retry_delay` before the next try. Not stuck; working as configured. Check `try_number` to see how many attempts have burned, and read an *early* attempt's log, since the first failure often carries the real cause. **`up_for_reschedule`** — a sensor running in `mode='reschedule'` has released its worker slot and will be re-queued at the next poke interval. Looks idle in the grid, and is the correct state for a long wait. A sensor in the default `mode='poke'` would instead sit in `running` and hold its slot, which is the classic way to deadlock a pool with sensors. **`deferred`** — a deferrable operator has handed a trigger to the triggerer component and released its worker slot. If tasks pile up here, check that the **triggerer** process is actually running; without it, deferred tasks never resume. **`upstream_failed` / `skipped`** — the instance was never eligible. `upstream_failed` means a predecessor failed under an `all_success`-style trigger rule; `skipped` usually means a branch operator chose a different path, or a `ShortCircuitOperator` cut the branch. Both redirect the investigation upstream; nothing is wrong with this task. **`failed` / `success`** — terminal. ## A worked triage order 1. **Open the grid view** and identify the state of the earliest non-successful instance in the run. The first blocked task, not the last, is the one to diagnose. 2. **`scheduled`/`none`** → look at DAG-level gates: paused, `max_active_runs`, `depends_on_past`, trigger rules. 3. **`queued`** → look at capacity: pool slots, `max_active_tasks`, queue-to-worker mapping, worker/pod health. Check the **scheduler** logs, not the task logs — a queued task has no task log yet, which itself is a useful tell. 4. **`running`** → read the task log and the downstream system. Check `execution_timeout` exists; if not, this can hang indefinitely. 5. **`deferred`** → is the triggerer up? 6. **`up_for_retry`** → read attempt 1's log, not the latest. 7. **`upstream_failed`/`skipped`** → move upstream and start again. ## The two states people misread The pair worth calling out in an interview is **`queued` versus `running`**. A queued task has produced no task log at all, and engineers who go straight to the log see an empty page and conclude Airflow is broken. Queued means capacity — pools, concurrency, queue routing, dead workers — and the evidence is in the scheduler logs and the pool view. The other is **`up_for_reschedule` versus `running`** for sensors. A sensor sitting in `running` for six hours in poke mode is silently consuming a worker slot the whole time; the same sensor in reschedule mode alternates between `up_for_reschedule` and brief `running` moments and costs nothing while waiting. A pool exhausted by poke-mode sensors is one of the most common self-inflicted 'stuck DAG' causes, and it presents exactly as other tasks parked in `queued`. ## What to instrument so you do not need this drill State-based triage is reactive. The proactive versions are: a dashboard of task instances by state over time (Airflow can emit these as StatsD or OpenTelemetry metrics), an alert on queued-duration exceeding a threshold, pool utilisation graphs, and `execution_timeout` on every task so nothing can hang unbounded in `running`.
- Tasks sit in queued for an hour and their logs are empty. What does that tell you?That no worker ever claimed them, which is why there is no task log to read. Look at capacity and routing rather than the DAG: exhausted pool slots, `max_active_tasks` reached, a `queue` with no worker consuming it, dead Celery workers, or pods that cannot be scheduled. The evidence lives in the scheduler logs and the pool view.
- Why does a sensor in the default poke mode make other tasks look stuck?Because a poke-mode sensor sits in `running` and holds its worker slot for the whole wait. A handful waiting hours can exhaust a pool, and everything else then parks in `queued` with no obvious cause. Switching to `mode='reschedule'` releases the slot between pokes — the sensor shows `up_for_reschedule` instead — or use a deferrable operator.
- Every instance of one task is stuck in scheduled across several consecutive runs. What would you check first?`depends_on_past=True` with an unresolved earlier instance. That setting pins the task in every later run until the previous run's instance succeeds, so one old failure silently blocks a queue of runs. `wait_for_downstream`, a saturated `max_active_runs`, and an unsatisfiable trigger rule are the other gates to rule out.
- Tasks are accumulating in the deferred state and never resuming. What is missing?The triggerer. Deferrable operators release their worker slot and hand a trigger to Airflow's separate triggerer component; if that process is not running, nothing ever fires the resume and the instances sit deferred indefinitely. It is a distinct component from the scheduler, webserver and workers and is easy to omit from a deployment.
saying these in an interview costs you the question
- Opens task logs first when the instance is still queued
- Treats queued and running as interchangeable
- Blames the DAG code when the pool is out of slots
- Does not know poke-mode sensors hold a worker slot
- Ignores depends_on_past when consecutive runs stall