In Airflow, how would you choose between ExternalTaskSensor, TriggerDagRunOperator and Dataset scheduling for cross-DAG dependencies?
answer
- ownership decides more than mechanics
- one pulls, one pushes, one declares
- the pull mechanism matches on logical date
- think about what failure looks like for each
- ask first whether it should be one DAG
basics
~20 sMatch the mechanism to who owns the dependency. Datasets suit a declared data contract between separately owned DAGs; TriggerDagRunOperator suits a producer that deliberately fans out to a downstream it owns; ExternalTaskSensor suits waiting on a DAG you cannot modify, and is the most fragile.
solid answer
~50 sThree mechanisms with three different ownership models. **Dataset (Asset) scheduling** is declarative: the producer names what it publishes via `outlets`, the consumer lists it in `schedule`, neither knows the other's cadence, and the dependency is visible in the dataset graph. **`TriggerDagRunOperator`** is imperative push: the producer explicitly starts a named downstream DAG, optionally with `conf` and `wait_for_completion`. **`ExternalTaskSensor`** is pull: the consumer polls another DAG's task, matching on logical date via `execution_delta` or `execution_date_fn`. I default to Datasets when the dependency is really about data readiness, use `TriggerDagRunOperator` when one team owns both sides and needs to pass parameters, and use `ExternalTaskSensor` only when I cannot change the upstream — with `mode="reschedule"` or a deferrable variant so it does not hold a worker slot. The deciding factors are team boundaries, whether the schedules must align, and whether the dependency has to survive a backfill.
code
python · 22 linesfrom datetime import timedelta
from airflow.sensors.external_task import ExternalTaskSensor
from airflow.operators.trigger_dagrun import TriggerDagRunOperator
# pull: consumer waits on an upstream task, matched by logical date
wait = ExternalTaskSensor(
task_id="wait_for_orders",
external_dag_id="daily_orders",
external_task_id="load",
execution_delta=timedelta(hours=1),
failed_states=["failed", "skipped"],
timeout=60 * 60,
mode="reschedule", # do not hold a worker slot while waiting
)
# push: producer explicitly starts the downstream and passes the window
fire = TriggerDagRunOperator(
task_id="start_mart",
trigger_dag_id="orders_mart",
conf={"window_start": "{{ data_interval_start }}"},
wait_for_completion=False,
)go deeper
Know that Airflow offers more than one way for one DAG to depend on another, and that hardcoding a later start time and hoping upstream finished is not one of them.
Explain each mechanism concretely: the sensor polls and matches on logical date, the trigger operator pushes and can pass conf, dataset scheduling declares producer outlets and consumer schedules.
Demonstrate operational judgment — reschedule or deferrable sensors, timeouts and failed_states, staleness alerts for dataset consumers, and how each mechanism behaves during a backfill.
Own the coupling policy across teams: where declared data assets replace cross-repository schedule negotiations, when a chain should simply be one DAG, and how you keep the silent failure mode observable at organisational scale.
## Why this is a judgment call Every orchestration platform of any size eventually has more than one DAG, and the moment it does, someone has to express "B needs A's output". Airflow offers several ways, and the choice is less about mechanics than about **who owns which side** and **how the dependency behaves when things go wrong**. ## `ExternalTaskSensor` — pull, date-matched A task in the consumer polls the metadata database for a specific task instance in another DAG: ```python ExternalTaskSensor( task_id="wait_for_orders", external_dag_id="daily_orders", external_task_id="load", execution_delta=timedelta(hours=1), allowed_states=["success"], failed_states=["failed", "skipped"], mode="reschedule", ) ``` Its defining property — and its defining weakness — is that it matches on **logical date**. By default it looks for the upstream run with the *same* logical date; when schedules differ you must express the offset with `execution_delta` or a custom `execution_date_fn`. That offset is a piece of coupling written on the consumer side, silently invalidated whenever the upstream's schedule changes. The classic production failure is an upstream DAG moved from hourly to every two hours; the sensor now waits for a run that will never exist and times out every day. It is also the only one of the three that consumes resources while waiting. In the default poke mode it occupies a worker slot for the whole wait; `mode="reschedule"` frees the slot between checks, and a deferrable variant hands the wait to the triggerer process. On a busy cluster, a fleet of default-mode sensors is a self-inflicted capacity outage. Use it when you cannot modify the upstream DAG — a different team's repository, a vendor-supplied DAG — and always set `timeout`, `failed_states`, and a non-poke wait mode. ## `TriggerDagRunOperator` — push, explicit The producer names the downstream and starts it: ```python TriggerDagRunOperator( task_id="start_mart", trigger_dag_id="orders_mart", conf={"window_start": "{{ data_interval_start }}"}, wait_for_completion=False, ) ``` Strengths: the dependency is unambiguous, it can pass parameters through `conf`, and there is no schedule alignment to get wrong. With `wait_for_completion=True` the producer even reflects the downstream's outcome, which makes a multi-DAG chain behave like one unit. Weaknesses: it is a *hard-coded fan-out*. Adding a second consumer means editing the producer, so the producer accumulates knowledge of every downstream — which is exactly wrong when the downstreams belong to other teams. And `wait_for_completion=True` means the producer's task occupies a slot for the downstream's entire duration, so it inherits the sensor's capacity problem. Use it when one team owns both DAGs, when parameters must flow, or when the split is really about DAG size rather than ownership. If the split exists purely because the DAG got long, consider whether the honest answer is one DAG with task groups. ## Dataset / Asset scheduling — declarative The producer declares what it publishes; the consumer declares what it needs; the scheduler connects them. Neither side names the other's schedule, adding a second consumer requires no producer change, and the dependency graph is inspectable in the UI. This is the closest Airflow gets to a **data contract**, and it is the right default when the DAGs belong to different teams. The costs are real. A dataset-triggered run has no data interval, so a consumer that needs a defined window per run must obtain it another way. There is no catchup over dataset events, so historical replay through the same path does not exist — you backfill the consumer explicitly. Events are recorded on task *success*, not verified writes, so a producer that fails before emitting leaves the consumer silently idle with nothing failing anywhere. And datasets link DAGs inside a single Airflow deployment. ## The decision framework Ask, in order: 1. **Is this really one pipeline?** If one team owns every DAG in the chain and they always run together, the split may be artificial. One DAG with task groups gives you a single run, a single retry story and a single backfill. 2. **Who owns each side?** Cross-team dependencies want a declared contract (Datasets), because a contract survives one side reorganising its internals. Same-team chains tolerate direct coupling. 3. **Does the consumer need a data window per run?** If yes, dataset triggering fights you; interval-based scheduling with a sensor, or passing the window through `conf`, fits better. 4. **Must the dependency survive a backfill?** Replaying six months through a dataset-triggered chain does not work like the live path. If historical replay is routine, prefer interval-aligned DAGs where every run has a window. 5. **What does failure look like?** Sensors fail loudly on timeout; datasets fail *silently* by never firing. Choose the failure mode you can actually detect, and add the matching alert — timeout alerts for sensors, staleness alerts for dataset consumers. 6. **What does it cost while waiting?** Never leave a default poke-mode sensor in a busy cluster; use reschedule or deferrable so waiting is nearly free. ## The organisational view At scale the real decision is about coupling, not syntax. A platform that standardises on declared data assets lets teams change their internal schedules without a cross-repository coordination meeting; a platform built on date-matched sensors makes every schedule change a negotiation. Against that, dataset-driven graphs are harder to reason about during an incident because there is no clock to compare against — "it should have run by now" requires you to have written down what "by now" means. The mature answer is usually both: declared assets for cross-team edges, direct coupling inside a team's own chain, and staleness monitoring so the silent failure mode is not silent.
- Why is ExternalTaskSensor considered the most fragile of the three?Because it matches on logical date. Any schedule change upstream invalidates the execution_delta written on the consumer, and the sensor then waits for a run that will never exist until it times out. It also encodes upstream knowledge in the downstream repository, and in default poke mode it burns a worker slot for the entire wait.
- When would you argue the DAGs should be merged into one instead?When one team owns every DAG in the chain, they always run together on the same cadence, and the split exists only because the file got long. A single DAG gives one run to monitor, one retry and clearing story, and a backfill that replays the whole chain coherently — none of which a cross-DAG mechanism reproduces for free.
- How do the three mechanisms differ in how a failure surfaces?TriggerDagRunOperator fails visibly on the producer's task. ExternalTaskSensor fails loudly at its timeout, and can fail fast if you set failed_states. Dataset scheduling fails silently: an upstream that never succeeds emits no event, so the consumer simply never runs and nothing anywhere is red. Dataset consumers therefore need staleness alerts rather than failure alerts.
- What breaks first when a cross-DAG chain has to be backfilled six months?Date matching and event semantics. A sensor replaying history must find upstream runs at the matching logical dates, which requires backfilling both DAGs in the right order. Dataset triggering has no historical events to replay at all, so the consumer must be backfilled explicitly over a range. If replay is routine, prefer interval-aligned DAGs where every run has a window.
saying these in an interview costs you the question
- Picks a mechanism on syntax preference without asking who owns each DAG
- Leaves ExternalTaskSensor in default poke mode on a busy cluster
- Assumes dataset-driven chains can be backfilled like scheduled ones
- Adds a TriggerDagRunOperator per consumer, making the producer know every downstream
- Ignores that dataset consumers fail by going silent rather than turning red