In Airflow, why does a DAG run's data_interval_start sit behind the time the run actually starts?
answer
- the window, not the clock
- completeness is the reason for the delay
- the run is named for the interval it covers
- it fires only after the interval closes
- logical_date equals data_interval_start for cron
basics
~20 sAirflow names a run for the data window it covers, not the moment it fires. A daily DAG's run whose data_interval_start is May 1 only starts once that day has closed, at May 2 midnight, so the data is complete.
solid answer
~40 sSince Airflow 2.2 every scheduled DAG run carries a half-open **data interval** `[data_interval_start, data_interval_end)`, and the scheduler creates the run only *after* the interval has ended. For `schedule="@daily"`, the run covering 2024-05-01 has `data_interval_start=2024-05-01T00:00` and `data_interval_end=2024-05-02T00:00`, and it begins executing at 2024-05-02T00:00. The reason is data completeness: a task that summarises May 1 cannot run until May 1 is over. `logical_date` (the old `execution_date`) equals `data_interval_start` for cron and timedelta schedules, and `{{ ds }}` is its date part — which is why logs "look a day behind". Tasks should filter source data with `{{ data_interval_start }}` and `{{ data_interval_end }}` rather than calling `datetime.now()`, so a re-run of that interval reprocesses the same window.
code
python · 21 linesimport pendulum
from airflow import DAG
from airflow.operators.python import PythonOperator
def load(start, end):
print(f"select * from orders where created_at >= {start} and created_at < {end}")
with DAG(
dag_id="daily_orders",
start_date=pendulum.datetime(2024, 5, 1, tz="UTC"),
schedule="@daily",
catchup=False,
) as dag:
PythonOperator(
task_id="load",
python_callable=load,
op_kwargs={
"start": "{{ data_interval_start }}",
"end": "{{ data_interval_end }}",
},
)go deeper
Recall that a daily DAG's run for a given date starts the following midnight, and that {{ ds }} names the day being processed rather than the day it ran.
Explain the half-open interval, why the run waits for the interval to close, and which template fields to put in a WHERE clause. Be able to say what logical_date equals for a cron schedule.
Diagnose interval bugs in production: tasks reading now(), filters open on one side, late-arriving data that makes the boundary wrong, and DAGs whose real requirement is a trigger time rather than a window.
Own the convention across a platform — interval-templated filters as a review rule, a house policy on schedule offsets versus late data, and a clear position on which DAGs are window-shaped versus event-shaped.
## The interval model Airflow does not schedule *moments*; it schedules *windows of data*. A DAG's `schedule` and `start_date` cut the timeline into consecutive, non-overlapping data intervals. Each interval is half-open: `[data_interval_start, data_interval_end)` — the start instant is included, the end instant is not, so intervals abut without double-counting a row that lands exactly at midnight. The scheduler creates the DAG run for an interval **once that interval has closed**. So for `schedule="@daily"`: ```text interval [2024-05-01 00:00, 2024-05-02 00:00) logical_date 2024-05-01 00:00 ({{ ds }} = "2024-05-01") run starts 2024-05-02 00:00 (after the interval ends) ``` This single design decision explains the behaviour that confuses almost everyone on first contact: a DAG scheduled `@daily` appears to be "one day behind", an hourly DAG "one hour behind", and a `@monthly` DAG for January does not run until 1 February. ## Why wait for the interval to close Because the run's *job* is to process the interval. A task that aggregates yesterday's orders, or copies the `dt=2024-05-01` partition, cannot produce a complete answer until 2024-05-01 is over. Firing at the interval's *start* would guarantee empty or partial output. Firing at the end and naming the run for the window it covers means the run identity, the data it reads and the partition it writes all agree — which is what makes replaying an interval meaningful. ## The fields you actually template - `{{ data_interval_start }}` / `{{ data_interval_end }}` — the window bounds, as timezone-aware datetimes. These are what your `WHERE` clause should use. - `{{ logical_date }}` — the run's nominal timestamp. For cron and `timedelta` schedules it equals `data_interval_start`. - `{{ ds }}` — `logical_date` rendered as `YYYY-MM-DD`; `{{ ts }}` is the full ISO timestamp. - `{{ prev_data_interval_start_success }}` — the start of the last *successful* run's interval, useful for gap-tolerant incremental loads. A correct extraction filter is bounded on both sides and half-open, mirroring the interval: ```sql select * from orders where created_at >= '{{ data_interval_start }}' and created_at < '{{ data_interval_end }}' ``` The anti-pattern is calling `datetime.now()` (or `CURRENT_DATE`) inside the task. That makes the run's output depend on *when it executed* rather than *what it covers*, so a retry an hour later, or a replay next month, produces a different result — and the whole replay story collapses. ## `execution_date`, `logical_date`, and the version story Before Airflow 2.2 the run carried a single field, `execution_date`, which held the *start* of the period being processed while the run executed at the *end* of it. The name told people it was "when this executes", which it never was. Airflow 2.2 introduced the explicit data-interval model, renamed the field to `logical_date`, and exposed `data_interval_start`/`data_interval_end` as first-class values. Airflow 3.0 removed `execution_date` and its derived macros entirely, so code that still references it fails outright rather than lying quietly. ## Manual runs and non-interval schedules A manually triggered run still needs an interval. The timetable's `infer_manual_data_interval` hook supplies one — for a standard cron schedule, the most recent complete interval preceding the trigger time. So triggering a daily DAG by hand at 14:00 gives you yesterday's window, not today's. That is usually what you want, and is a frequent surprise when it is not. If you genuinely want "fire at 09:00 and treat 09:00 as the timestamp" — no interval offset — Airflow ships `CronTriggerTimetable` (in `airflow.timetables.trigger`), which triggers *at* the cron time rather than at the end of a data interval. Reach for it when the DAG's work is event-shaped rather than window-shaped, for example polling an external system that has no notion of your window. Asset- and dataset-triggered runs likewise have no meaningful window; do not build partition logic on the interval of such a run. ## Timezones `start_date` should be timezone-aware (Airflow uses `pendulum`). Cron schedules are evaluated in the DAG's timezone, so a DAG in a DST-observing zone keeps firing at local 02:00 even as the UTC offset shifts, while `timedelta` schedules are pure elapsed time and drift relative to local clock time across a DST boundary. Storing and comparing everything in UTC downstream avoids a whole class of duplicate-or-missing-hour bugs. ## Debugging checklist When someone reports "the DAG ran but the table is empty", check in order: is the task filtering on `data_interval_*` or on `now()`; is the filter half-open on both sides; and is the source data actually complete by the time the interval closes — if late-arriving rows land two hours after midnight, the interval boundary is right and the *schedule offset* is what needs to change.
- How do you make a daily Airflow DAG fire at 02:00 rather than midnight, without changing which day it processes?Use a cron schedule with the offset, e.g. `schedule="0 2 * * *"`. The intervals then run 02:00-to-02:00, so the run covers 02:00 yesterday to 02:00 today. If you need midnight-to-midnight windows but a 02:00 start, keep `@daily` semantics and shift the filter explicitly, or write a custom Timetable — the offset must live somewhere, and putting it in the schedule is the readable choice.
- Why is calling datetime.now() inside an Airflow task a scheduling bug rather than a style issue?It ties the output to execution time instead of the run's interval. A retry two hours later, a replay next week, and the original run all read different windows, so the task is no longer idempotent, backfills produce different data than the original runs, and gaps or overlaps appear invisibly. Templating `data_interval_start`/`data_interval_end` makes the run reproducible.
- What data interval does a manually triggered Airflow DAG run get?One inferred by the DAG's timetable via `infer_manual_data_interval`. For a standard cron or timedelta schedule that is the most recent complete interval before the trigger time — so a hand-triggered daily DAG processes yesterday, not today. Asset-triggered runs have no meaningful window, so partition logic should not depend on their interval.
saying these in an interview costs you the question
- Says logical_date is when the run executes
- Filters source data with now() or CURRENT_DATE inside the task
- Thinks the interval is inclusive at both ends, double-counting boundary rows
- Claims Airflow runs late because the scheduler is slow
- Uses execution_date in new code without knowing it was removed in Airflow 3