skip to content

An Airflow ExternalTaskSensor in a 06:00 DAG times out waiting on a task in a 02:00 DAG. Why?

level: seniorimportance: should knowfreq 40%

answer

  1. it is not waiting for "the latest run"
  2. which run is it actually looking for?
  3. two schedules, two different stamps
  4. one argument translates between the two calendars
  5. execution_delta, plus check_existence to fail fast

basics

~20 s

By default ExternalTaskSensor looks for a run of the other DAG at exactly its own logical date. Two DAGs on different schedules never share a logical date, so the run it waits for does not exist and it waits until timeout.

solid answer

~50 s

`ExternalTaskSensor` does not wait for "the last successful run" of the other DAG — it waits for a run whose **logical date matches its own**, and then for `external_task_id` within that run to be in `allowed_states`. A DAG scheduled at 06:00 and a DAG scheduled at 02:00 produce runs stamped at different times, so the sensor queries for a run that will never exist and simply keeps poking until `timeout`. The fix is to translate between the two calendars with `execution_delta` — here a four-hour delta pointing back at the 02:00 run — or with the callable form for irregular mappings. Two hardening arguments belong in the same answer: `check_existence=True` so a typo'd `external_dag_id` fails fast instead of hanging, and `failed_states` so the sensor fails when the upstream task failed rather than waiting out its whole timeout on a run that is already dead.

code

python · 16 lines
python
from datetime import timedelta
from airflow.sensors.external_task import ExternalTaskSensor

# 06:00 DAG waiting on the 02:00 DAG's run of the same morning
wait = ExternalTaskSensor(
    task_id="wait_for_ingest",
    external_dag_id="ingest_events",
    external_task_id="publish",
    execution_delta=timedelta(hours=4),
    allowed_states=["success"],
    failed_states=["failed", "skipped"],
    check_existence=True,
    mode="reschedule",
    poke_interval=300,
    timeout=60 * 60 * 3,
)

go deeper

for a junior

Recall that this sensor matches the upstream run by logical date rather than picking the latest one, so DAGs on different schedules need an explicit offset.

for a middle

Explain what the sensor queries on each poke, how execution_delta shifts the date it looks for, and why a mismatch shows up as a timeout instead of an error.

for a senior

Diagnose it end to end: date mismatch versus renamed task versus failed upstream, and harden with check_existence, failed_states and reschedule or deferrable mode. Mention what a backfill does to it.

for a principal

Own the cross-DAG dependency pattern: whether polling another DAG's history is allowed at all, versus upstream-triggers-downstream, one merged pipeline, or asset-driven scheduling.

## What the sensor actually queries `ExternalTaskSensor` is cross-DAG dependency by lookup, not by subscription. On each poke it asks the metadata database, in effect: *does a run of `external_dag_id` exist with logical date X, and is `external_task_id` in that run in one of `allowed_states`?* The key detail is where X comes from. By default X is **the sensor's own run's logical date** — not "the most recent run", not "today's run". That default is deliberate and correct for the intended case: two DAGs on the *same* schedule, where matching intervals really do correspond. It is wrong for everything else, and the failure is silent: no error, no missing-object exception, just a sensor that pokes patiently until it times out, every single day. ## Making the calendars line up - **`execution_delta`** — a `timedelta` *subtracted* from the sensor's logical date to produce the one it looks for. A 06:00 DAG waiting on that morning's 02:00 run uses `execution_delta=timedelta(hours=4)`. Getting the sign backwards is the most common follow-up bug, and it presents identically: perpetual timeout. - **The callable form** — a function mapping your logical date to the target's, for cases a fixed delta cannot express: a daily DAG waiting on the last hourly run, a weekday DAG waiting on Friday's run, a monthly roll-up waiting on the month's final daily run. Return a single date or a list of dates when you must wait for several runs. - You cannot pass both a delta and a callable — pick one. ## Hardening arguments people forget - **`check_existence=True`** makes the sensor verify that the referenced DAG and task actually exist and fail immediately if not. Without it, a renamed task or a typo in `external_dag_id` produces the same symptom as everything else — a long silent wait. - **`failed_states`** lets the sensor treat upstream failure as failure. Left unset, an upstream run that failed at 02:05 leaves the sensor poking until its own timeout hours later, delaying the alert and burning capacity for a result that will never come. - **`allowed_states`** defaults to success but can include others; and `external_task_id=None` waits on the whole DAG run rather than a single task. - **`mode='reschedule'` or `deferrable=True`** — this is a long wait by construction, so it should not hold a worker slot in poke mode. ```python ExternalTaskSensor( task_id="wait_for_ingest", external_dag_id="ingest_events", external_task_id="publish", execution_delta=timedelta(hours=4), failed_states=["failed", "skipped"], check_existence=True, mode="reschedule", poke_interval=300, timeout=60 * 60 * 3, ) ``` ## Backfills are where this really bites Because the match is on logical date, a backfill of the downstream DAG looks for upstream runs at the corresponding historical dates. If the upstream DAG was created later, was renamed, or has had its history cleaned up, those runs do not exist and every backfilled sensor times out — hours of poking per run before anything reports a problem. Anyone who has backfilled a month of a cross-DAG pipeline has felt this; saying so is what marks the answer as lived rather than read. ## The design point Even configured perfectly, this sensor couples two DAGs through a polling loop over the metadata database and inherits the fragility of a shared calendar. The alternatives worth naming: have the upstream DAG trigger the downstream one explicitly at the end of its work; merge the two DAGs if they are genuinely one pipeline with one cadence; or move to data/asset-driven scheduling so the downstream is started by the upstream's output rather than by a clock plus a poll. A senior answer fixes the timeout *and* asks whether the dependency should have been expressed by polling at all.

  • When would you use the callable form instead of execution_delta?
    When the relationship between the two calendars is not a fixed offset: a daily DAG waiting on the last hourly run of the day, a weekday DAG waiting on Friday's run, a monthly roll-up waiting on the final daily run of the month. The callable maps your logical date to the target's, and can return several dates when you must wait for multiple runs.
  • Why does setting failed_states matter so much here?
    Without it, an upstream task that failed early leaves the sensor poking for hours until its own timeout, so the alert arrives long after the real failure and capacity is spent waiting for something that can never succeed. With failed_states the sensor fails as soon as the upstream is known dead.
  • What is a better way to express a cross-DAG dependency than polling?
    Let the producer drive the consumer: have the upstream DAG trigger the downstream one when its work is genuinely complete, use asset- or data-driven scheduling so the downstream starts on the upstream's output, or merge the two DAGs if they are one pipeline on one cadence. Each removes the shared-calendar coupling and the poll entirely.

saying these in an interview costs you the question

  • Believes the sensor waits for the most recent successful upstream run
  • Thinks a timeout means the upstream DAG is slow
  • Gets the execution_delta sign backwards and blames the sensor
  • Leaves failed_states unset so upstream failure waits out the full timeout
  • Runs this long cross-DAG wait in poke mode, holding a worker slot

context