skip to content

XComs and Data Passing

The small side-channel for passing values between tasks, stored in the metadata database unless you configure a backend. Interviewers ask what you do with a large dataset, and the expected answer is a reference to object storage rather than the payload itself.

on this pageshow

questions

6

In Airflow, what is an XCom and how does one task push a value another task pulls?

level: juniorimportance: must knowfreq 80%

answer

  1. a side channel, not a data pipe
  2. tasks are separate processes on separate machines
  3. the value is a row in the metadata database
  4. push writes a key, pull reads it
  5. the implicit key is return_value

basics

~10 s

An XCom is Airflow's small cross-task message. A task pushes a keyed value into the metadata database, and a downstream task in the same DAG run pulls it back by task id and key.

solid answer

~50 s

XCom — short for cross-communication — is Airflow's built-in side channel for passing **small** values between tasks in the same DAG run. A task pushes explicitly with `ti.xcom_push(key='row_count', value=42)`, or implicitly: any operator left at the default `do_xcom_push=True` stores its return value under the key `return_value`. A downstream task reads it with `ti.xcom_pull(task_ids='extract', key='row_count')`; omit `key` and you get `return_value`. The same call works inside a Jinja template: `{{ ti.xcom_pull(task_ids='extract') }}`. Values live as rows in the `xcom` table of the metadata database, scoped to dag id, task id, run and map index, so a pull normally sees only the current DAG run. Two things trip people up: `xcom_pull` does **not** create a dependency — you still write `extract >> load` — and XCom is for small facts like counts, ids and object paths, not for the data itself.

code

python · 20 lines
python
from airflow import DAG
from airflow.operators.python import PythonOperator
import pendulum

def extract(ti):
    ti.xcom_push(key="row_count", value=42)

def load(ti):
    count = ti.xcom_pull(task_ids="extract", key="row_count")
    print(f"loading {count} rows")

with DAG(
    dag_id="xcom_demo",
    start_date=pendulum.datetime(2024, 1, 1, tz="UTC"),
    schedule="@daily",
    catchup=False,
) as dag:
    a = PythonOperator(task_id="extract", python_callable=extract)
    b = PythonOperator(task_id="load", python_callable=load)
    a >> b  # the pull does NOT create this edge

go deeper

for a junior

Be ready to define XCom in one sentence and write the push/pull pair from memory, including that an operator's return value lands under the key return_value.

for a middle

Explain the identity of an XCom row — dag, run, task, key, map index — and why that scoping means a pull normally sees only the current run's value.

for a senior

Show the judgment about payload size and secrets: XCom carries pointers and small facts, and a task that needs to hand over a dataset writes it to storage and pushes the URI.

for a principal

Own the boundary: XCom is control-plane metadata for the orchestrator, and treating it as a transport turns your metadata database into a data store the whole platform depends on.

## What an XCom is Airflow tasks do not share memory. Each task instance runs as its own process, often on a different worker machine or in a different Kubernetes pod, so a Python variable set in one task simply does not exist in the next. XCom ("cross-communication") is the mechanism Airflow provides to close that gap: a small, keyed value written by one task and read by another within the same DAG run. An XCom row is identified by the DAG id, the run, the task id, the key, and (for mapped tasks) the map index. That composite identity is why the same task pushing on Monday and Tuesday produces two separate values rather than overwriting one. ## Pushing a value There are two ways to push. **Explicitly**, from inside a Python callable, using the task instance object that Airflow passes into the context: ```python def extract(ti): ti.xcom_push(key="row_count", value=42) ``` **Implicitly**, by returning a value. Every operator inherits a `do_xcom_push` parameter that defaults to `True`; when the task finishes, its return value is stored under the reserved key `return_value`. A `PythonOperator` pushes whatever the callable returns; a `BashOperator` pushes the last line the command wrote to stdout. Setting `do_xcom_push=False` on the operator turns that off, which is what you do when an operator returns something big or uninteresting. ## Pulling a value ```python def load(ti): count = ti.xcom_pull(task_ids="extract", key="row_count") ``` If you leave `key` out, `xcom_pull` looks for `return_value` — the implicit push. If you pass a list of task ids, you get a list of values back, in the order of the ids you asked for. Templated fields can pull too, because the task instance is available in the Jinja context: `bash_command="echo {{ ti.xcom_pull(task_ids='extract') }}"`. By default a pull is scoped to the current DAG run. `include_prior_dates=True` widens the search to earlier runs, which is occasionally useful for "what was the last watermark" patterns but makes the DAG's behaviour depend on history — a trap for backfills. ## Pulling is not a dependency This is the single most common beginner error. `xcom_pull` is a database lookup executed at runtime; it tells the scheduler nothing. If you write a `load` task that pulls from `extract` but never wire `extract >> load`, the scheduler is free to run them in parallel, and `load` will very likely pull `None`. You must still declare the edge — either with the bitshift operators, or by using the TaskFlow API, where passing one task's output into another call creates the edge for you. ## What belongs in an XCom Because values are serialized into a row in Airflow's own metadata database, XCom is a control-plane channel, not a data-plane one. Good payloads are small facts: a row count, a partition date, a job id returned by an external system, an S3 or GCS URI naming where the real data was written. Bad payloads are the data itself — a DataFrame, a file's contents, a large JSON document. The accepted pattern is that the upstream task writes the dataset to object storage or a warehouse table and pushes only the pointer. XComs are also not the right place for credentials. Values are stored unencrypted and are rendered in the web UI, unlike Connections and Variables, which use Airflow's Fernet encryption. Secrets belong in a Connection or a secrets backend. ## When a pull returns None A `None` from `xcom_pull` almost always means one of: the `task_ids` or `key` string does not match what was pushed (they are plain strings, and typos fail silently); the upstream operator had `do_xcom_push=False`; the upstream callable simply returned `None`; the two tasks are not in the same DAG run; or there is no dependency, so the pull ran first. Debugging starts by opening the XCom tab for the upstream task instance in the UI and looking at exactly which keys exist. ## Related concepts you should not confuse with it Airflow Variables are deployment-wide key/value configuration read at parse or run time; Connections hold credentials and endpoints. Both are global and long-lived. An XCom is per-task, per-run, and transient — it is about *this* execution of *this* pipeline, and it disappears when the run's rows are cleaned out of the metadata database.

  • Does calling `xcom_pull` on an upstream task create a dependency between the two tasks?
    No. `xcom_pull` is just a lookup performed while the task runs; the scheduler never inspects it. You still declare `extract >> load` explicitly, or use the TaskFlow API, where passing one task's output as an argument to another creates the edge. Without the edge the pull can execute before the push and quietly return `None`.
  • `xcom_pull` returns `None` even though the upstream task shows success. What do you check first?
    Open the upstream task instance's XCom tab in the UI and see which keys actually exist. The usual causes are a mismatched `task_ids` or `key` string, an operator configured with `do_xcom_push=False`, a callable that returned `None`, a missing dependency so the pull ran first, or pulling across DAG runs without `include_prior_dates`.
  • How do you stop an operator from pushing its return value at all?
    Set `do_xcom_push=False` on the operator. Every operator inherits the parameter from `BaseOperator` and it defaults to `True`. Turning it off is worth doing whenever the return value is large or useless — a `BashOperator` whose last stdout line is noise, or a task that returns a big object you never intend to read downstream.

saying these in an interview costs you the question

  • Says XCom is how you pass DataFrames between tasks
  • Thinks xcom_pull automatically creates the task dependency
  • Believes XComs are shared across DAG runs by default
  • Confuses XCom with Airflow Variables or Connections
  • Assumes tasks share memory so a global variable works

context

open as a page

An Airflow task returns a 2 GB DataFrame via XCom — what breaks, and what should it pass instead?

level: seniorimportance: must knowfreq 66%

basics

~20 s

It fails: a DataFrame is not JSON-serializable, and even serialized it would not fit the metadata database's XCom column. Write the frame to object storage or a table and push only the URI or partition key.

open as a page

Where does Airflow store XCom values by default, and what limits their size?

level: middleimportance: should knowfreq 70%

basics

~20 s

Airflow serializes each XCom to JSON and writes it as a row in the xcom table of its metadata database. The ceiling is that column's type, so the practical budget is kilobytes to a few megabytes, not gigabytes.

open as a page

In Airflow's TaskFlow API, how does returning a value from an @task function create an XCom?

level: middleimportance: should knowfreq 62%

basics

~20 s

The @task decorator wraps the function in an operator whose return value is pushed as the return_value XCom. Calling the function in a DAG returns a lazy reference, and passing it to another task both wires the dependency and generates the pull.

open as a page

How do you configure a custom XCom backend in Airflow, and what must the class implement?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Point the core xcom_backend setting at your class, which subclasses BaseXCom and overrides serialize_value and deserialize_value. Serialize writes the payload to external storage and returns a small pointer; deserialize resolves the pointer back into the value.

open as a page

What policy would you set for XCom use across a shared multi-team Airflow platform?

level: principalimportance: should knowfreq 33%

basics

~20 s

Treat XCom as control-plane metadata only: small JSON facts and object URIs, never payloads and never secrets, since values are unencrypted and visible in the UI. Add scheduled retention, and enforce the rules with cluster policies rather than documentation alone.

open as a page