How would you decide between one large Airflow DAG with TaskGroups and several smaller linked DAGs?
answer
- grouping is cosmetic, splitting is structural
- ask whether they should share a run
- schedule and ownership are the seams
- the rerun boundary is the real cost
- cross-DAG links are weaker links
basics
~20 sSplit on schedule and ownership, not on size. Work that shares one data interval and one failure story belongs in one DAG, grouped with TaskGroups for readability; work on a different cadence or owned by another team belongs in its own DAG, linked explicitly.
solid answer
~50 sA TaskGroup is presentation and namespacing only: the tasks stay in the same DAG, sharing the run, the schedule, the concurrency limits and the clear-and-rerun boundary. So the real question is whether the work should share a *run*. Keep it in one DAG when everything is driven by the same interval, succeeds or fails as a unit, and one team owns it — then the rerun boundary is exactly what an operator wants. Split when cadences differ, when ownership differs (separate deploys, separate on-call, separate pause switch), when one slow section blocks unrelated work under `max_active_tasks`, or when the graph is so wide that clearing it re-runs far more than the failure justifies. The cost of splitting is that the dependency becomes explicit and weaker: `TriggerDagRunOperator`, `ExternalTaskSensor` with its interval-alignment traps, or dataset-driven scheduling — each with its own failure modes and a harder end-to-end view.
code
python · 9 linesfrom airflow.utils.task_group import TaskGroup
with DAG("daily_sales", schedule="@daily", start_date=START, catchup=False) as dag:
with TaskGroup("ingest") as ingest:
pull_orders = PythonOperator(task_id="orders", python_callable=pull_orders_fn)
pull_refunds = PythonOperator(task_id="refunds", python_callable=pull_refunds_fn)
# task ids become ingest.orders and ingest.refunds; same run, same limits
ingest >> build_martsgo deeper
Know that TaskGroups only organise the graph visually and that a separate DAG is what gives separate scheduling; you are not expected to make the call yourself.
Explain the concrete shared properties — one run, one interval, one set of concurrency limits, one clear boundary — that a group keeps and a split breaks.
Argue the operational trade-off with incidents in mind: rerun granularity, contention, and the specific failure modes of trigger operators and cross-DAG sensors.
Own the decomposition policy: where the seams are, which linking mechanism the platform standardises on, how ownership and paging map onto DAGs, and how end-to-end latency stays visible once one pipeline spans several runs.
## First, what a TaskGroup actually gives you `TaskGroup` (and the `@task_group` decorator) nests tasks under a collapsible node in the UI and prefixes their task ids with the group name. That is the whole of it. The grouped tasks are ordinary members of the same DAG: same DAG run, same data interval, same `max_active_tasks`, same `max_active_runs`, cleared together when you clear the DAG, backfilled together, paused together. A TaskGroup is a *readability* tool, not an isolation boundary. (It replaced the old SubDAG mechanism, which looked like isolation but brought its own scheduling and concurrency pathologies and is removed in Airflow 3.) So the decision is not "one DAG or many" in the abstract. It is: *should these tasks share a run?* ## Reasons to keep it as one DAG **One interval, one atom.** If the whole pipeline processes the same day of data and only means something end to end, one DAG makes the unit of scheduling equal to the unit of meaning. Operators reason about one run, one status, one rerun. **Dependencies stay real.** Inside a DAG, an edge is enforced by the scheduler with no timing assumptions. Between DAGs it becomes a sensor waiting on the right interval or a trigger fired by a task — a weaker, more failure-prone link. **Backfill is coherent.** Re-running a historical interval re-runs the whole chain in the right order. Across DAGs, a backfill has to be coordinated by hand, DAG by DAG, in dependency order. **One end-to-end view.** Duration, SLA and lineage are all naturally scoped to the run. ## Reasons to split **Different cadences.** An hourly ingestion and a daily aggregate cannot share a schedule. Forcing them together means either running the aggregate too often or holding the ingestion back. **Different owners.** A DAG is the practical unit of ownership: pausing, deploying, alert routing and permissions are all easier per DAG. If two teams need to pause their halves independently, they need two DAGs. **Blast radius and rerun granularity.** In one DAG, clearing an upstream task re-runs everything downstream of it. If a 300-task DAG regularly needs a rerun of only its last third, the shared run costs money and time on every incident. **Resource contention.** DAG-level `max_active_tasks` and `max_active_runs` apply to the whole DAG, so a long-running section can starve unrelated tasks in the same DAG; separate DAGs (plus pools for the shared external system) isolate them. **Parse and UI cost.** Very large generated DAGs get slow to parse and unreadable in the grid view; the graph stops being a communication tool. ## What splitting costs, mechanism by mechanism `TriggerDagRunOperator` pushes: the upstream DAG fires the downstream one, optionally waiting for completion. It is explicit and easy to follow, but the downstream DAG now has two ways to start (its own schedule and the trigger), and the coupling lives in the upstream repo. `ExternalTaskSensor` pulls: the downstream DAG waits for a specific task in another DAG *for a matching interval*. The classic outage is interval misalignment — an hourly DAG waiting on a daily one without an execution-date offset waits forever, holding a worker slot unless it runs in reschedule mode or is deferrable. Timeouts and clear semantics need deliberate thought. Dataset-driven scheduling inverts it: the producing task declares what it updates, and the consuming DAG is scheduled by that update rather than by a clock. It expresses the real dependency — data, not time — and removes the interval-alignment trap, at the cost of a newer model and a less obvious relationship to intervals and backfills. Whatever the mechanism, you lose the single end-to-end run: measuring true latency from raw landing to published mart now means stitching runs together, and lineage or observability tooling has to do that stitching for you. ## How to answer as a lead Give the rule, then the exceptions. *Cut at the seams where schedule or ownership changes; inside a seam, use TaskGroups for structure and never split just because the graph looks big.* Then name the checks you would run before splitting: how often do we rerun only part of this; do the halves have different SLAs; would splitting turn one strong dependency into a fragile sensor; who gets paged for each half. And name the guardrails after splitting: a naming convention, one documented triggering mechanism used consistently, pools for shared external systems, and lineage-level monitoring so nobody has to reconstruct the end-to-end path by reading DAG code.
- What exactly does a TaskGroup change about how its tasks are scheduled?Nothing. It nests them under one collapsible UI node and prefixes their task ids; they remain ordinary tasks in the same DAG, sharing the run, the data interval, the DAG-level concurrency limits and the clear-and-rerun boundary. If you need independent scheduling, pausing or rerun granularity, a group cannot give it to you — that needs a separate DAG.
- What is the most common failure mode when a downstream DAG waits on an upstream one with ExternalTaskSensor?Interval misalignment: the sensor looks for the upstream task instance at a matching logical date, so two DAGs on different schedules never match and the sensor waits until it times out. Fix it with an explicit execution-date offset or delta, set a timeout so it fails loudly instead of hanging, and use reschedule mode or a deferrable sensor so it does not hold a worker slot while waiting.
- When would dataset-driven scheduling be a better link than a sensor?When the real dependency is "the table has been refreshed" rather than "the clock says 6am". The producer declares what it updates and the consumer is triggered by that update, which removes interval-alignment guesswork and lets the consumer run as soon as the data is genuinely ready. The trade-off is a model that maps less directly onto historical backfills.
- How do you keep resource contention from forcing a split you do not want?Use pools to cap concurrency against the shared scarce resource — a warehouse, an API, a cluster — and set per-task concurrency limits so one heavy section cannot consume every slot; priority weights help order the queue. Splitting for isolation is right only when the halves also differ in schedule or ownership; otherwise you trade a solved problem for a fragile cross-DAG link.
saying these in an interview costs you the question
- Treats TaskGroups as isolated sub-DAGs with their own schedule
- Splits purely because the DAG has many tasks
- Ignores that clearing an upstream task re-runs everything downstream
- Chains DAGs with sensors without considering interval alignment
- Assumes cross-DAG dependencies are as reliable as in-DAG edges