skip to content

Airflow task instances sit in the queued state for hours and never start — how do you diagnose it?

level: seniorimportance: should knowfreq 55%

answer

  1. dependencies are already satisfied at this point
  2. idle workers plus queued tasks means configuration
  3. someone may be publishing to a queue nobody reads
  4. for pods, ask the cluster why it will not schedule

basics

~20 s

Queued means Airflow's scheduler handed the task to the executor but nothing picked it up. Check in order: pool and concurrency limits, whether any worker consumes the task's queue, worker or broker health, and for Kubernetes whether the pod is stuck Pending.

solid answer

~50 s

`queued` in Airflow means dependencies passed and the executor accepted the task, but no worker has started it. Work outward from Airflow's own throttles: is the task's **pool** out of free slots, or is the DAG at `max_active_tasks` / `max_active_runs`, or the scheduler at `[core] parallelism`? Those show as queued tasks while workers sit idle. If limits are clear, suspect the hand-off. With `CeleryExecutor` the classic cause is a **queue nobody consumes** — the task sets `queue="gpu"` and no worker was started with `--queues gpu` — followed by a dead worker fleet or an unreachable broker. With `KubernetesExecutor`, look at the pods: `Pending` means the cluster cannot schedule them (resources, quota, node selector, taints) and `ErrImagePull` means the worker image is wrong. Also check the scheduler logs and whether the scheduler is alive at all; recent Airflow 2 releases will eventually time a task out of `queued` and retry it, which masks the cause if you do not look.

code

bash · 9 lines
bash
# what queues does this worker actually consume?
airflow celery worker --queues gpu,default

# Kubernetes executor: why has the worker pod not started?
kubectl -n airflow get pods
kubectl -n airflow describe pod <worker-pod>

# force a clean re-dispatch once the cause is fixed
airflow tasks clear my_dag -t load_orders -s 2026-08-20 -e 2026-08-21

go deeper

for a junior

Know what the queued state means: the scheduler is done with it and something downstream has not picked it up yet.

for a middle

Explain the ordered checks — pool and concurrency limits first, then queue routing, then worker and broker health.

for a senior

Show real diagnosis: correlating idle workers with queued tasks, inspecting Celery subscriptions or pending pods, and clearing orphans after a scheduler restart.

for a principal

Argue for the guardrails that prevent recurrence: queued-duration alerting, CI validation of queue names, capacity policy, and who owns the broker or cluster.

## What queued means An Airflow task instance moves `scheduled` → `queued` → `running`. `queued` is the boundary between the scheduler's world and the execution world: the scheduler has verified upstream dependencies, the trigger rule, pool availability and concurrency limits, and has handed the task to the executor. A task stuck in `queued` therefore is *never* a dependency problem — it is a capacity, routing or infrastructure problem. ## Step 1: are Airflow's own limits binding? The most common and most benign cause is that the task is legitimately waiting for a slot. - **Pool exhausted.** Open the Pools page: if used slots equal total slots, everything else assigned to that pool queues. A task with `pool_slots=4` in a 4-slot pool waits for the pool to be completely empty. - **DAG limits.** `max_active_tasks` and `max_active_runs` on the DAG cap it independently of cluster capacity. - **`[core] parallelism`.** The scheduler-wide ceiling. The tell for all of these is queued tasks *while workers are idle*. ## Step 2: is the hand-off actually reaching a worker? If limits are not binding, the message went somewhere nobody is listening. With **`CeleryExecutor`**, check the task instance's `queue` value in the UI, then check what your workers subscribe to. `airflow celery worker --queues gpu` is a *subscription*, not a validated assignment: a task routed to a queue that no running worker consumes waits forever with no error anywhere. This happens after renaming a queue, decommissioning a worker pool, or a typo in `queue=`. Next, confirm workers are alive and connected — Flower, `celery inspect active`, or simply the worker logs — and that the broker is reachable and not out of memory. A Redis broker that has evicted keys, or a RabbitMQ node in an alarm state, will silently drop or stall messages. With **`KubernetesExecutor`**, the diagnosis moves to the cluster: `kubectl get pods -n <airflow-ns>` and `kubectl describe` the worker pod. `Pending` with `FailedScheduling` means insufficient CPU/memory, an unschedulable node selector or an untolerated taint, or a namespace ResourceQuota that is full. `ErrImagePull`/`ImagePullBackOff` means the worker image tag or registry credentials are wrong. If no pod exists at all, the scheduler could not create it — look for Kubernetes API errors or RBAC denials in the scheduler log. ## Step 3: is the scheduler healthy? If the scheduler died after queueing tasks, nothing will progress them. On restart, Airflow tries to adopt or re-queue orphaned task instances; if adoption fails you can see tasks queued against an executor that no longer knows about them. Check the scheduler's heartbeat, its logs for exceptions in the loop, and — in Celery deployments — whether the executor believes it has more running slots than it does. Clearing the affected task instances forces a clean re-dispatch and is the standard remedy once the root cause is fixed. ## Step 4: the metadata database A saturated or locked metadata database slows the scheduling loop enough that tasks appear stuck. Long-running queries, missing autovacuum on PostgreSQL, or an undersized instance all show up as an Airflow that "queues but does not run". ## Guardrails so it does not happen silently - Alert on queued-state duration, not just on failure. A task queued for an hour is an incident even though nothing has failed. - Recent Airflow 2 releases will fail or retry a task that has been queued past a configured timeout, which converts a silent hang into a visible failure — useful, but make sure you still capture the reason. - Validate `queue` names in CI against the queues your workers actually subscribe to; this single check eliminates the most common Celery instance of the problem. - Monitor broker depth and worker count together: rising depth with flat worker count is the signature of an under-scaled or dead fleet.

  • How do you tell a pool-starved task from a task routed to an unconsumed Celery queue?
    Look at the task instance details. If it names a pool and the Pools page shows zero free slots, it is throttled and will start on its own once slots free up. If the pool has free slots but the task names a queue, list your running workers and their `--queues` subscriptions; a queue with no subscriber will never drain. The second case never resolves by waiting, which is the practical difference.
  • After a scheduler crash, some task instances are queued but no worker is running them. What do you do?
    Confirm the scheduler is back and heartbeating, then check whether it adopted the orphans — the scheduler log reports adopting or resetting orphaned task instances. If they remain stuck, clear those task instances in the UI or with `airflow tasks clear`, which returns them to a schedulable state so the executor dispatches them cleanly. Fix the crash cause first, or they will simply pile up again.
  • What would you alert on so this is caught automatically?
    Alert on the age of the oldest task instance in the `queued` state, per queue or pool, with a threshold well below your SLA. Pair it with broker depth versus active worker count for Celery, or pending worker-pod count for Kubernetes. Failure-only alerting misses this entirely, because a stuck queued task never fails until a queued timeout or an SLA miss fires.

saying these in an interview costs you the question

  • Starts by checking upstream task dependencies
  • Restarts the scheduler without looking at pools or queues
  • Assumes queued always means the cluster is out of capacity
  • Ignores broker or Kubernetes API health entirely
  • Only alerts on failures, never on queued duration

context