In Airflow, what do the >> and << operators do between two tasks in a DAG?
answer
- the arrow is an edge, not a pipe
- it says after, not immediately after
- same thing as set_downstream
- lists broadcast, list-to-list does not
- order is decided at parse time
basics
~10 sIn Airflow, a >> b puts b downstream of a: b is not started until a has finished successfully. The arrow declares execution order only. It moves no data between the tasks.
solid answer
~40 s`a >> b` is shorthand for `a.set_downstream(b)`, and `b << a` writes the same edge the other way. It adds a directed edge that the scheduler uses to decide eligibility: with the default `trigger_rule='all_success'`, `b` becomes schedulable only once every upstream task has succeeded. Lists give you fan-out and fan-in in one line — `extract >> [clean_a, clean_b] >> load` — but list-to-list is not supported by the operators; use the `chain()` helper for that. Two things the arrow does **not** do: it does not pass data (that is XComs or shared storage), and it does not promise `b` runs immediately or on the same worker — only that it runs after. Edges are built when the DAG file is parsed, so a task cannot add one at run time.
code
python · 7 linesextract = PythonOperator(task_id="extract", python_callable=pull_orders)
clean_a = PythonOperator(task_id="clean_orders", python_callable=clean_orders)
clean_b = PythonOperator(task_id="clean_customers", python_callable=clean_customers)
load = PythonOperator(task_id="load", python_callable=load_warehouse)
# fan-out then fan-in: load waits for BOTH cleaners
extract >> [clean_a, clean_b] >> loadgo deeper
Be able to read and write both forms fluently and to sketch the resulting graph on a whiteboard, including the fan-out and fan-in forms with a list on one side.
Explain that the edge only encodes ordering evaluated against a trigger rule, that it carries no data, and that the graph is fixed when the file is parsed rather than while a run executes.
Point out the operational consequences: tasks joined by an arrow can run minutes apart on different workers, so shared local state is a bug, and run-time-dependent shape needs branching or dynamic mapping instead.
Push for a house style — one direction, helpers like chain for long sequences, dependencies declared in one readable block — because a graph that reviewers cannot read at a glance is where incorrect ordering hides.
## What the arrow declares An Airflow DAG is a directed acyclic graph whose nodes are tasks and whose edges are ordering constraints. The bitshift operators are Airflow's syntax for adding those edges. `a >> b` reads "a then b" and is exactly equivalent to `a.set_downstream(b)`; `b << a` is the mirrored form, equivalent to `b.set_upstream(a)`. Airflow implements this by overriding `__rshift__` and `__lshift__` on `BaseOperator` (and on the `XComArg` objects the TaskFlow `@task` decorator returns), so the syntax works on operator instances, on lists of them, and on `TaskGroup` objects. The edge is graph metadata, nothing more. When a DAG run is created, every task instance starts in the `none` state and the scheduler repeatedly asks whether each task's dependencies are met. The default answer rule is `trigger_rule='all_success'`: all direct upstream tasks must be in `success` before the task is queued. Change the trigger rule and the same edge is still there but the condition on it changes — the edge says *ordering*, the trigger rule says *under what upstream outcome*. ## The four ways to write the same thing ```python extract >> transform # bitshift, the idiomatic form transform << extract # identical edge, reversed reading extract.set_downstream(transform) transform.set_upstream(extract) ``` All four produce the same serialized DAG. Teams pick one style and stay with it; mixing `>>` and `<<` in one file is a readability complaint reviewers raise often, because the reader has to re-establish direction on every line. ## Fan-out, fan-in, and the list rule A list on either side broadcasts: ```python extract >> [clean_orders, clean_customers] >> load ``` That is four edges: extract to each cleaner, and each cleaner to load. `load` waits for **both** cleaners under the default trigger rule. What does *not* work is list on both sides — `[a, b] >> [c, d]` raises a Python error, because a plain list has no `>>` behaviour to fall back on. When you genuinely want that, use the helpers: `chain(a, b, c)` wires a sequence, `chain([a, b], [c, d])` pairs same-length lists element-wise, and `cross_downstream([a, b], [c, d])` builds the full cross product. ## What the arrow is not **It is not data flow.** `extract >> transform` gives `transform` no access to whatever `extract` computed. Values travel through XComs (small metadata) or, for anything of size, through object storage with only the path passed along. Candidates who say "the arrow passes the return value" are conflating the dependency with TaskFlow's implicit XCom, which is a *separate* mechanism that happens to create the edge as a side effect. **It is not "runs immediately after".** Once dependencies are met the task is queued; when it actually starts depends on the executor, free slots, pool capacity and DAG-level concurrency limits. Two tasks joined by an arrow routinely run on different workers, in different processes, possibly minutes apart. Anything that relies on shared local state — a file written to `/tmp`, an open database session, a Python variable — breaks the moment the two tasks land on different machines. **It is not dynamic.** Dependencies are declared while the DAG file is being parsed, before any run exists. Writing `if result > 10: a >> b` inside a task's Python body has no effect on the graph; the arrow ran on a worker, long after the scheduler built its picture of the DAG. Run-time shape changes need the mechanisms designed for it: branching operators to *skip* paths, or dynamic task mapping to fan out over a list that is only known at run time. ## Cycles The A in DAG is enforced. If your declarations create a loop, Airflow raises a cycle error when the file is parsed and the DAG is marked as broken rather than being scheduled with an arbitrary order. The most common accidental cycle comes from a loop that wires each element to the next without breaking on the last item, or from wiring a TaskGroup back into one of its own members. ## Reading a graph in an interview Be fluent at translating both directions. Given `a >> [b, c] >> d >> e`, you should be able to say instantly: a runs first; b and c run in parallel once a succeeds; d waits for both; e waits for d. And given that description you should be able to write the two lines back.
- If `a >> b` does not move data, how does `b` get a value that `a` computed?Through XCom: `a` pushes a small value into the metadata database and `b` pulls it, either explicitly with `xcom_pull` or implicitly by taking a TaskFlow function's return value as an argument. XCom is sized for identifiers and small metadata, not datasets — for anything large, write to object storage in `a` and pass only the path.
- What happens if the declarations in a DAG file create a cycle?Airflow detects it while parsing the file and raises a cycle error, so the DAG shows as broken with an import error instead of being scheduled. There is no partial run and no arbitrary tie-break: acyclicity is a precondition for the scheduler being able to answer "are this task's dependencies met" at all.
- Why does `[a, b] >> [c, d]` fail, and what do you use instead?The left operand is a plain Python list, which has no bitshift behaviour, so there is nothing for Airflow to hook into. Use `chain([a, b], [c, d])` to pair the two lists element-wise, or `cross_downstream([a, b], [c, d])` to connect every task on the left to every task on the right.
saying these in an interview costs you the question
- Says the arrow passes the upstream task's return value
- Claims the downstream task starts immediately after the upstream one
- Assumes both tasks share a worker, filesystem or process
- Adds dependencies inside a task body and expects the graph to change
- Thinks >> and << mean different kinds of dependency