skip to content

Apache Airflow

The de-facto orchestrator: DAGs written in Python, operators and sensors for the work, a scheduler driving intervals, and pluggable executors. Interviewers focus on Airflow because its scheduling semantics and XCom limits shape how pipelines get designed.

on this pageshow

explore

questions

page 2 of 2

How do you configure a custom XCom backend in Airflow, and what must the class implement?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Point the core xcom_backend setting at your class, which subclasses BaseXCom and overrides serialize_value and deserialize_value. Serialize writes the payload to external storage and returns a small pointer; deserialize resolves the pointer back into the value.

open as a page

How would you decide between one large Airflow DAG with TaskGroups and several smaller linked DAGs?

level: principalimportance: should knowfreq 36%

basics

~20 s

Split on schedule and ownership, not on size. Work that shares one data interval and one failure story belongs in one DAG, grouped with TaskGroups for readability; work on a different cadence or owned by another team belongs in its own DAG, linked explicitly.

open as a page

How would you choose between CeleryExecutor and KubernetesExecutor for a team's Airflow platform?

level: principalimportance: should knowfreq 40%

basics

~10 s

Characterise the workload first. Many short tasks and stable dependencies favour CeleryExecutor's warm workers; heterogeneous resource needs, per-team images and strong isolation favour KubernetesExecutor's pod-per-task, which costs seconds of startup on every task.

open as a page

What policy would you set for XCom use across a shared multi-team Airflow platform?

level: principalimportance: should knowfreq 33%

basics

~20 s

Treat XCom as control-plane metadata only: small JSON facts and object URIs, never payloads and never secrets, since values are unencrypted and visible in the UI. Add scheduled retention, and enforce the rules with cluster policies rather than documentation alone.

open as a page

In Airflow 2, how can several schedulers run at once without duplicating task instances?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Airflow 2 supports multiple active schedulers with no leader election. They coordinate through the metadata database, taking row-level locks with SELECT ... FOR UPDATE so only one scheduler can claim and queue a given task instance.

open as a page

When is a shared library of custom Airflow operators worth building for a platform team?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

When the same integration plus its policy — auth, tagging, error mapping, a quality gate — is being retyped across many DAGs by teams who should not have to know it. Otherwise prefer provider operators and thin Python tasks over hooks: a custom operator library is a versioned product you then own.

open as a page

In Airflow, how would you choose between ExternalTaskSensor, TriggerDagRunOperator and Dataset scheduling for cross-DAG dependencies?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Match the mechanism to who owns the dependency. Datasets suit a declared data contract between separately owned DAGs; TriggerDagRunOperator suits a producer that deliberately fans out to a downstream it owns; ExternalTaskSensor suits waiting on a DAG you cannot modify, and is the most fragile.

open as a page

showing 31–37 of 37