What policy would you set for XCom use across a shared multi-team Airflow platform?
answer
- it is a control channel, not a data channel
- small facts, pointers, nothing more
- unencrypted and visible in the UI
- the table needs a retention job
- cluster policies enforce what wikis cannot
basics
~20 sTreat XCom as control-plane metadata only: small JSON facts and object URIs, never payloads and never secrets, since values are unencrypted and visible in the UI. Add scheduled retention, and enforce the rules with cluster policies rather than documentation alone.
solid answer
~50 sI would write the policy around one idea: XCom is the orchestrator's control channel, not a data channel. **Allowed** — row counts, partition dates, external job ids, object URIs; small, JSON-serializable, ideally reconstructible from the run's logical date. **Forbidden** — datasets of any size, and secrets of any kind, because XCom values are stored unencrypted in the metadata database and rendered in the UI, unlike Connections and Variables which are Fernet-encrypted. **Operationally** — `airflow db clean` on a schedule with an alert on `xcom` table growth, so one team's habit cannot degrade everyone's scheduler. Then the real decision: a deployment-wide offloading backend gives ergonomics but hides data movement, creates a shared bucket and a retention bill; explicit URIs keep ownership with the team. I would default to the convention, and enforce both convention and size ceilings with a `task_policy` cluster policy so it is checked at parse time, not in review.
code
python · 11 lines# airflow_local_settings.py - enforced at parse time, deployment-wide
NOISY_OPERATORS = {"BashOperator"}
def task_policy(task):
# stop noisy operators from pushing their last stdout line
if type(task).__name__ in NOISY_OPERATORS:
task.do_xcom_push = False
def dag_policy(dag):
if not dag.owner or dag.owner == "airflow":
raise ValueError(f"{dag.dag_id}: set an explicit owner")go deeper
Take away the two absolute rules: keep XComs small, and never put a password, token or key in one because the value is visible in the UI.
Be able to explain why the limits exist — values are unencrypted rows in the scheduler's own database, and nothing removes them until a retention job does.
Show you would operationalise it: a scheduled db clean, an alert on xcom table growth, and object paths derived from the data interval so reruns survive trimmed history.
Own the tradeoff and the enforcement — transparent offloading versus explicit convention, who pays the storage bill, and cluster policies that make the rule mechanical rather than aspirational.
## Frame the decision On a single-team Airflow, XCom hygiene is a code-review habit. On a shared platform with dozens of teams, it is a capacity and security question, because every XCom lands in one metadata database that every team's scheduling depends on. The policy question is therefore: what may pass through this shared channel, who enforces it, and what happens when someone ignores it. ## Rule one: size and shape Allowed payloads are facts about a run, not the run's output: a row count, a partition date, an external system's job id, an object URI, a short status string. Small, JSON-serializable, and — the property I would push hardest — *reconstructible*. If a downstream task can derive the S3 path from `data_interval_start` instead of reading it from an XCom, the pipeline survives clearing, retries and trimmed history. Forbidden are datasets. The reason is not aesthetic. XCom values are rows in the metadata database, so payload growth becomes scheduler-loop latency, slower UI, larger backups and longer restores. Those costs land on every team, not the one that caused them, which is the textbook definition of something a platform must govern rather than suggest. I would state an explicit ceiling — say, single-digit kilobytes as the norm and anything approaching a megabyte requiring a conversation — because a number is enforceable and "keep it small" is not. ## Rule two: no secrets, ever This one is non-negotiable and is the security answer interviewers listen for. Airflow encrypts Connections and Variables with its Fernet key. XCom values are **not** encrypted: they sit in the metadata database in the clear, appear in database backups, and are rendered in the web UI to anyone who can view that task instance. A task that fetches a short-lived token and pushes it for the next task to use has effectively published it. The pattern instead is that each task retrieves the credential from a Connection or the configured secrets backend at the moment it needs it. ## Rule three: retention is a platform job XCom rows do not expire. `airflow db clean --clean-before-timestamp` must run on a schedule with a retention window the platform chooses, and the `xcom` table's size should be a monitored metric with an alert, because table growth is the leading indicator that someone is passing payloads. Retention also interacts with correctness: once an old run's XComs are trimmed, clearing a single downstream task in that run leaves it with nothing to pull. That is another reason the policy prefers values that can be recomputed from the logical date. ## The genuine tradeoff: transparent backend versus explicit convention Here is where reasonable platform teams differ, and a principal-level answer should present both sides rather than declare one. **A custom XCom backend** that offloads payloads to object storage keeps `return df` working, protects the metadata database without asking anyone to rewrite DAGs, and is a single central change. Against it: it applies to every XCom in the deployment, so all teams inherit the round-trip latency; the DAG source no longer tells a reader where data went; one shared bucket blurs tenant isolation and creates a retention bill nobody owns, since Airflow's cleanup deletes rows and never the objects. **Explicit convention** — tasks write to their own storage and push a URI — keeps ownership, cost and lifecycle with the team that produced the data, and keeps the DAG honest about what it is doing. Against it: it is a change every team must make, and it is only as strong as your enforcement. My default is the convention, with the backend held in reserve for a specific, measured problem — for example, a compliance requirement that pipeline values not reside in the Airflow database. ## Enforcement beats documentation A wiki page does not stop a deadline. Airflow's cluster policies — `dag_policy` and `task_policy` defined in `airflow_local_settings.py` — run at parse time against every DAG and task in the deployment, and can reject or mutate what they see. That is where a platform expresses rules mechanically: reject tasks that opt into pushing when the team is on a deny list, force `do_xcom_push=False` on operator types known to emit noise, require a tag or an owner. Pair it with a CI lint on DAG repositories and a dashboard of the `xcom` table, and you have prevention, detection and a conversation trigger rather than a rule nobody reads. ## What I would actually write down One page: allowed content with examples; a size ceiling with a number; a hard ban on secrets with the reason; the retention window and who runs the clean; the convention for object paths derived from the data interval; and the escalation path for a team that thinks they are the exception. Then the cluster policy that enforces as much of it as parse time can see, and a metric that catches the rest.
- Why is pushing a short-lived API token through an XCom a security problem?XCom values are stored unencrypted in the metadata database, unlike Connections and Variables, which are Fernet-encrypted. The token therefore sits in the clear in the database, appears in backups, and is rendered in the UI to anyone who can view that task instance. Each task should fetch the credential from a Connection or the secrets backend when it needs it.
- How would you enforce an XCom policy across dozens of teams rather than documenting it?Cluster policies. A `task_policy` or `dag_policy` in `airflow_local_settings.py` runs at parse time against every task in the deployment and can reject or mutate what it sees — forcing `do_xcom_push=False` on noisy operator types, requiring owners or tags. Back it with CI lint on the DAG repos and an alert on `xcom` table growth to catch what parse time cannot.
- When would you accept a deployment-wide offloading XCom backend despite its downsides?When there is a specific measured problem it uniquely solves — a compliance boundary requiring pipeline values not live in the Airflow database, or metadata-database growth you cannot fix by asking many teams to rewrite DAGs on your timeline. Go in knowing you have bought latency for every XCom, a shared bucket, and a retention bill Airflow's own cleanup will never pay.
saying these in an interview costs you the question
- Treats XCom as a safe place for credentials or tokens
- Assumes XCom values are encrypted like Connections
- Relies on a wiki page instead of parse-time enforcement
- Never schedules db clean, letting the xcom table grow
- Adopts an offloading backend without owning object retention