What is a zombie task in Airflow, and what causes the scheduler to mark one?
answer
- the database says one thing, the process another
- liveness is proved by a periodic signal
- OOM killer and evicted pods produce them
- the scheduler sweeps and fails them
- there is an opposite condition too
basics
~20 sAn Airflow zombie task is a task instance the metadata database still shows as running while no live worker process is heartbeating for it. The scheduler detects the missing heartbeats and fails the instance so its retries or downstream handling can proceed.
solid answer
~50 sIn Airflow, a running task instance's process periodically heartbeats into the metadata database. A **zombie** is an instance the database still marks `running` whose heartbeats have stopped — the worker process died without getting to report failure. Typical causes are the OOM killer reaping the process, a Kubernetes pod evicted or preempted, a Celery worker node terminated, or the machine losing power. The scheduler runs a periodic zombie check; when an instance has not heartbeated within the configured threshold, it fails the instance, logs a zombie message, and lets the normal retry path take over. The mirror image is the **undead** task: a process still running for an instance the database has been told to clear or mark failed, which Airflow kills on detection. Zombies are an infrastructure signal — if you see them regularly, look at worker memory limits and node stability, not at your DAG code.
code
text · 12 lines[2024-03-12 04:17:02] INFO - Running task: transform_events, attempt 1
[2024-03-12 04:19:44] INFO - Loaded 8,412,000 rows into memory
-- log ends here, no traceback, no exit line --
scheduler log:
[2024-03-12 04:24:10] ERROR - Detected zombie job: {'dag_id': 'events',
'task_id': 'transform_events', 'run_id': 'scheduled__2024-03-11T00:00:00+00:00'}
kubectl describe pod ...:
Last State: Terminated
Reason: OOMKilled
Exit Code: 137go deeper
Recall the definition: the metadata database still says the task is running but the process that was executing it is gone, so the scheduler cleans it up.
Explain the heartbeat mechanism, the scheduler's periodic sweep, and that resolution goes through the ordinary failure path so retries and failure callbacks still fire.
Diagnose one for real: read the pod termination reason or kernel log rather than the task log, tie it to OOM or preemption, and fix the memory profile or the capacity class rather than the threshold.
Own the platform posture — worker memory sizing and limits, which workloads may run on preemptible capacity, and how much slot starvation you accept between a worker dying and the scheduler noticing.
## The heartbeat contract When Airflow runs a task, two things are true at once: a row in the metadata database records the task instance as `running`, and a process somewhere — a Celery worker's child, a Kubernetes pod, a LocalExecutor subprocess — is actually executing the code. That process periodically writes a heartbeat timestamp back to the database. The heartbeat is the only evidence the scheduler has that the work is still alive, because the scheduler has no other channel to the worker. A **zombie task** is the state where the database says `running` but the heartbeats have stopped. The process is gone and nothing told the database, so without intervention the row would sit in `running` forever, holding a concurrency slot and blocking anything that waits on it. ## What kills a process without letting it report Normal failures never produce zombies: an exception propagates, the task process catches it, writes `failed` or `up_for_retry`, and exits cleanly. Zombies come from failures that deny the process that final write: - The Linux **OOM killer** reaps the process because the task allocated more memory than the container or node allowed. This is the single most common cause, and it is why a pandas task that works on a sample and dies on the full dataset shows up as a zombie rather than a `MemoryError`. - A **Kubernetes pod is evicted or preempted** — node pressure, a spot/preemptible instance reclaimed, a node drained for an upgrade. - The **worker node dies**: hardware failure, an operator terminating an instance, a network partition that isolates the worker from the database long enough to miss the threshold. - A `SIGKILL` from any source — someone running `kill -9`, a container runtime enforcing a limit. ## How the scheduler detects and resolves them The scheduler periodically scans for task instances in a running state whose last heartbeat is older than a configured threshold, controlled by settings in the `[scheduler]` section of `airflow.cfg` (the zombie detection threshold and how often the check runs). When it finds one, it marks the instance failed and writes a log line naming it as detected as a zombie. Crucially, this goes through the ordinary failure path — so if the task had `retries` configured, it becomes `up_for_retry` and gets another attempt, and any `on_failure_callback` fires. This is exactly what you want: a preempted spot instance should just be retried elsewhere. The detection interval matters operationally. Between the process dying and the scheduler noticing, the instance still counts against DAG-level and pool concurrency, so a batch of zombies can quietly starve a busy Airflow of slots for minutes. ## Undead tasks — the opposite failure The complementary condition is the **undead** task: a process that is still running for a task instance whose database row no longer says it should be. The usual way to create one is clearing a task in the UI while its previous attempt is still executing, or manually marking a running instance as failed or success. Airflow detects the mismatch and kills the stray process, precisely so you do not end up with two copies of the same task writing to the same table. ## Diagnosing a recurring zombie Because a zombie task by definition never wrote a final log line, the task log usually ends mid-work with no traceback — that abrupt truncation is itself the tell. The investigation lives outside Airflow: 1. Check the **worker's** system logs or the pod's termination reason. `OOMKilled` in a Kubernetes pod status, or an oom-kill entry in the kernel log, settles it immediately. 2. Correlate the timing with node events — spot reclamation, autoscaler scale-down, a rolling node upgrade. 3. If it is memory, the fix is either raising the container's memory limit or changing the task so it does not hold the whole dataset in memory. Pushing the aggregation into the warehouse and having the task only orchestrate it is usually the durable answer. 4. If it is preemption, the fix is retries with backoff plus, where it matters, running that task on a queue backed by non-preemptible capacity. 5. If heartbeats are being missed while the process is fine — a badly overloaded metadata database, or a task that blocks the process so hard it cannot heartbeat — you are looking at a false positive, and the answer is fixing the database contention rather than raising the threshold. ## Why interviewers ask Zombie detection is a good probe because it separates people who have only authored DAGs from people who have operated a cluster. The right answer says three things: it is a heartbeat-liveness mechanism, the usual root causes are memory and node lifecycle rather than DAG code, and detection routes through the normal failure/retry path so a well-configured task recovers on its own. A candidate who says a zombie is 'a task that got stuck' has not operated one.
- After the scheduler marks an instance as a zombie, does the task get another chance?Yes, if it has retries. Zombie resolution goes through the normal failure path: the instance is failed, so a configured retry moves it to `up_for_retry` and it runs again. Any `on_failure_callback` fires too. That is deliberate — a preempted spot node should simply be retried elsewhere without human involvement.
- What is an undead task, and how do you create one by accident?An undead task is the inverse: a process still executing for a task instance the database no longer considers running. The usual cause is clearing or manually marking a task in the UI while its attempt is still live. Airflow detects the mismatch and kills the stray process, preventing two copies writing the same output.
- A task's log just stops mid-work with no traceback. Where do you look next?Outside Airflow. A zombie never got to write a final log line, so the evidence is in the worker's environment: the pod's termination reason (`OOMKilled`), the node's kernel log for an oom-kill entry, or cloud events for spot reclamation and autoscaler scale-down at that timestamp.
- Would raising the zombie detection threshold be a reasonable fix for frequent zombies?Almost never. If the process is genuinely dead, a longer threshold only means the slot stays occupied longer before recovery. Raising it is defensible in exactly one case: false positives caused by metadata-database contention so severe that healthy tasks miss heartbeats — and then the real fix is the database.
saying these in an interview costs you the question
- Calls any long-running or stuck task a zombie
- Thinks the DAG code raised an error that Airflow failed to catch
- Proposes raising the detection threshold as the fix
- Believes a zombie bypasses retries and callbacks
- Never mentions memory limits or pod eviction as causes