skip to content

In Airflow, what does setting catchup=False on a DAG do?

level: juniorimportance: must knowfreq 78%

answer

  1. it concerns intervals that already passed
  2. matters most when you unpause a DAG
  3. one flag decides: all gaps, or one
  4. False creates only the latest completed interval

basics

~20 s

With catchup=False, Airflow schedules only the most recent completed data interval when a DAG becomes active, instead of creating one run for every interval back to start_date. Older intervals are skipped and must be backfilled deliberately.

solid answer

~40 s

`catchup` controls whether the scheduler fills in the DAG runs it "owes" for intervals that already elapsed. Airflow divides time from `start_date` onward into data intervals — for `schedule="@daily"`, one per day — and owes exactly one DAG run per closed interval. With `catchup=True`, unpausing a DAG whose `start_date` is a year old makes the scheduler create a run for **every** elapsed interval, oldest first, throttled by `max_active_runs`. With `catchup=False` it creates only the latest completed interval and moves forward from there; the earlier intervals are never scheduled automatically. Set it explicitly on every DAG rather than relying on the default, which comes from the `core.catchup_by_default` config and differs between Airflow 2 and Airflow 3. `catchup` only affects scheduler-created runs — manual triggers, API triggers and clearing an existing run are unaffected.

code

python · 12 lines
python
import pendulum
from airflow import DAG
from airflow.operators.empty import EmptyOperator

with DAG(
    dag_id="marketing_emails",
    start_date=pendulum.datetime(2024, 1, 1, tz="UTC"),
    schedule="@daily",
    catchup=False,      # do not send a year of back-dated emails
    max_active_runs=1,
) as dag:
    EmptyOperator(task_id="send")

go deeper

for a junior

Be ready to say in one sentence what happens when you unpause a DAG whose start_date is months old, and which flag decides whether you get one run or hundreds.

for a middle

Explain the mechanics: intervals are owed per schedule, the scheduler creates them oldest first, and max_active_runs is what throttles the resulting queue. Know that the default comes from core.catchup_by_default.

for a senior

Show the judgment call — catchup=True only for runs that are idempotent functions of their interval, catchup=False for anything with external side effects — and describe how you would drain a legitimate backlog without melting the cluster.

for a principal

Own it as a platform convention: an explicit catchup value in every DAG, a house rule on start_date hygiene, and shared pools so one team's replay cannot starve everyone else's scheduled work.

## What `catchup` actually controls Airflow's scheduler does not think in terms of "run now". It thinks in terms of **data intervals**. Given `start_date=pendulum.datetime(2024, 1, 1)` and `schedule="@daily"`, the calendar from that date onward is cut into intervals `[Jan 1, Jan 2)`, `[Jan 2, Jan 3)`, and so on. Each closed interval is owed exactly one DAG run, and Airflow only creates that run once the interval has ended. `catchup` answers one question: when the scheduler notices that several owed intervals have already elapsed with no run, does it create all of them, or only the latest one? - `catchup=True` — the scheduler walks forward from `start_date` (or from the last existing run) and creates a DAG run for every elapsed interval, oldest first. - `catchup=False` — the scheduler jumps to the most recent completed interval, creates that single run, and proceeds normally from there. Earlier intervals are permanently skipped as far as automatic scheduling is concerned. ## The three moments it bites **Deploying a new DAG with a historical `start_date`.** People routinely set `start_date` to "when the data begins" — two years ago. With `catchup=True` that is two years of daily runs the instant the DAG is unpaused. **Unpausing a DAG that sat paused.** Pausing does not stop the clock; it stops run creation. Unpausing after two weeks with `catchup=True` produces fourteen queued runs. **A scheduler outage.** If the scheduler was down overnight, on restart it discovers the missed intervals and, with `catchup=True`, creates them. ## What `catchup=False` does *not* do It is not "backfills off". You can still replay any window explicitly — in Airflow 2 with `airflow dags backfill --start-date ... --end-date ... <dag_id>`; Airflow 3 moved backfills to be scheduler-managed and triggerable from the UI and API. Clearing an existing run also re-executes it. And `catchup=False` does not touch manual triggers: `airflow dags trigger` and the REST API create a run regardless. It also does not shrink an already-created queue. If a hundred runs exist because `catchup=True` was in effect, flipping the flag to `False` does not delete them — you must mark them or delete them yourself. Finally, flipping `False` back to `True` later does **not** recover the skipped middle. The scheduler derives the next interval from the *latest existing run*, so it fills forward from there, not backward into the gap. ## Choosing a value The deciding question is whether a historical run is *useful or harmful*. `catchup=True` is right when each run is an idempotent function of its interval — it reads a partition of source data bounded by `data_interval_start`/`data_interval_end` and overwrites the matching output partition. Replaying interval 2024-03-04 in June produces exactly the same table it would have produced in March, so filling gaps automatically is a feature. `catchup=False` is right when a run has real-world side effects or is inherently "as of now": sending customer emails or notifications, calling a payment or third-party API, syncing "whatever is currently in the bucket", refreshing a dashboard cache. Backfilling those means re-sending two weeks of messages or paying for two weeks of API calls. Because the safe default depends entirely on the DAG's semantics, the practical rule is to write `catchup=` explicitly in every DAG definition, so a reader never has to know which Airflow version's configuration default applies. ## Surviving a legitimate catchup When `catchup=True` is correct but the backlog is large, the controls are: - `max_active_runs` on the DAG — caps how many of its DAG runs execute concurrently. Without it, a backlog can saturate the whole cluster. - `max_active_tasks_per_dag` — caps concurrent task instances for the DAG. - Pools with `pool_slots` — bound pressure on a shared fragile resource such as a source database. - `depends_on_past=True` — forces each task instance to wait for its own previous run to succeed, which serializes the replay in interval order (useful for cumulative state, expensive for throughput). ```python with DAG( dag_id="daily_orders", start_date=pendulum.datetime(2024, 1, 1, tz="UTC"), schedule="@daily", catchup=True, max_active_runs=2, ) as dag: ... ``` ## The classic mistakes Assuming `catchup=False` means "run immediately with today's date" — it does not; the single run it creates still carries the *previous* completed interval's dates. Assuming `catchup=False` protects you from a manual backfill someone launches. And leaving a year-old `start_date` on a DAG whose tasks are not idempotent, then discovering it at 3 a.m. when the queue fills.

  • With catchup=False, how do you run the intervals the scheduler skipped?
    Trigger them explicitly. In Airflow 2 that is `airflow dags backfill --start-date ... --end-date ... <dag_id>`, optionally with `--reset-dagruns` to re-run runs that already exist; Airflow 3 makes backfills scheduler-managed and startable from the UI or API. Clearing existing task instances re-executes those runs. Flipping catchup back to True does not help — the scheduler fills forward from the newest run, not backward into the gap.
  • Why is catchup=True dangerous on a DAG that sends notification emails?
    Because every elapsed interval becomes a real run, and each run sends real email. Unpausing a DAG that was paused for two weeks would deliver fourteen days of back-dated messages to live recipients. Side-effecting DAGs should use catchup=False; catchup=True belongs on DAGs whose runs are idempotent partition overwrites.
  • Does catchup=False change what data a run processes?
    No. It only changes which runs get created. The run it does create still carries the previous completed interval's `data_interval_start` and `data_interval_end`, so a daily DAG unpaused at midday still processes yesterday's window, not today's partial day.

saying these in an interview costs you the question

  • Says catchup=False disables backfills entirely, including manual ones
  • Thinks catchup governs retries rather than missed data intervals
  • Believes catchup=False makes the run process today's in-progress data
  • Assumes catchup runs execute all at once with no concurrency limit
  • Confuses catchup with depends_on_past

context