In a scheduled data pipeline, why does the midnight run process yesterday's data rather than today's?
answer
- two different clocks are involved
- a run is named for a window
- the window has to close first
- midnight ends yesterday, it does not start today
basics
~20 sA scheduled run is named for the data interval it covers, not for the clock time it starts. A daily interval only closes at midnight, so the run firing then is responsible for the day that just ended.
solid answer
~40 sEvery scheduled run carries two different times: the **data interval** it is responsible for — a range like `[2026-03-11 00:00, 2026-03-12 00:00)` — and the **wall-clock moment** it actually started. A scheduler cannot process a window until the window is complete, so it waits for the interval to end and then launches the run that covers it. That is why a daily pipeline firing just after midnight on 12 March owns 11 March's data. The interval, not `now()`, is what the task should use to select its input rows and to name its output partition. That design is what makes runs deterministic: re-running the same interval next week selects exactly the same source rows and overwrites exactly the same partition, which is what makes backfills and retries safe.
code
text · 5 linesschedule: daily
data interval: [2026-03-11 00:00, 2026-03-12 00:00)
run starts at: 2026-03-12 00:00:05
run is labelled: 2026-03-11
output written: /warehouse/events/dt=2026-03-11/go deeper
Be able to state that a scheduled run covers a window of time and fires after that window ends, so the first run of a day handles the previous day.
Explain both timestamps precisely and show how the interval bounds feed the task's own filter, and why a now-based filter breaks on retries and reruns.
Show how interval-derived bounds give deterministic partition selection, and discuss the latency floor and the late-arriving-data gap the model creates in production.
Own the framing conversation: whether a stakeholder's 'daily' means cadence or freshness, and when the batch interval model should be replaced rather than tuned.
## The two clocks Every scheduled pipeline run has at least two timestamps attached to it, and most confusion about orchestrators comes from conflating them. - The **data interval** (also called the logical period, the run's window, or its partition): the slice of time whose data this run is responsible for. It is a *range*, not a point — for example `[2026-03-11 00:00, 2026-03-12 00:00)` for a daily schedule. - The **actual start time**: the wall-clock moment the scheduler queued and launched the run. It depends on load, worker availability, and whether an earlier run was still going. Schedulers usually give the run a single display timestamp derived from the interval (commonly its start), which is why a run that executed at 00:05 on 12 March shows up labelled 11 March. It is not a bug or an off-by-one; it is the run being named after the window it owns. ## Why the scheduler waits for the interval to end A daily window cannot be summarised while it is still filling. If a run for 12 March started at 00:00 on 12 March, it would read an empty table. So the rule is: **fire when the interval closes, and process the interval that just closed.** A pipeline with a daily schedule therefore has an inherent latency floor of one full interval plus its own runtime — data from 09:00 on 11 March is not in the warehouse until some time after midnight. The same rule applies at every granularity. An hourly schedule whose interval is `[09:00, 10:00)` fires at 10:00. A monthly schedule for March fires at the start of April. ## What the interval buys you The interval is the input to the task's own filtering logic. A task that derives its bounds from the run's interval is *deterministic*: run it now, run it in six months, run it three times in parallel, and it selects the same source rows and writes the same output partition. ```sql -- deterministic: bounds come from the run's interval SELECT * FROM events WHERE event_ts >= '2026-03-11 00:00:00' AND event_ts < '2026-03-12 00:00:00'; -- non-deterministic: depends on when the query happened to run SELECT * FROM events WHERE event_ts >= CURRENT_DATE - INTERVAL '1 day'; ``` The second query is the classic defect. It works on the happy path and lies on every other path: a retry at 00:58 the next night picks up a different set of rows, a backfill of last quarter reads the last 24 hours instead of the target day, and a run that was delayed by four hours silently shifts its window. Deterministic bounds are the precondition for everything else an orchestrator promises — safe retries, reproducible reprocessing, and the ability to run several historical windows concurrently without them fighting over the same output. ## Where people get caught **"The run timestamp is when it ran."** It usually is not; it is the interval anchor. Log lines and output paths built from it will look a day (or an hour) behind the wall clock, correctly. **"I'll just add a day to fix the label."** Shifting the label to make dashboards look right breaks backfills, because the shift is applied in one place and not the other. Fix the mental model rather than the arithmetic. **"Freshness equals schedule frequency."** With a daily schedule, the *oldest* record in the latest output can be nearly 48 hours old at the moment before the next run lands. If a stakeholder asks for "daily data", ask whether they mean a daily cadence or a freshness target — those are different requirements and only one of them is satisfied by a daily schedule. **Late-arriving rows.** Because the window is closed by the clock and not by a completeness signal, anything that lands after the interval closes is invisible to that run forever unless you deliberately reprocess. A pipeline that must tolerate late data needs an explicit reprocessing window, not a longer schedule. ## How to talk about it A strong answer names the two times, states the wait-for-the-interval-to-end rule, and then connects it to why the model exists at all: because a run identified by a window rather than by a moment can be re-executed at any time and still mean the same thing. That single property is what separates an orchestrated pipeline from a script behind a timer.
- If a stakeholder needs data within minutes of the event, what does the interval model cost them?Latency of at least one interval plus the run's own duration, because the window must close before processing starts. Shortening the interval reduces the floor but multiplies run overhead and small output files. If minutes matter, the answer is a streaming or micro-batch path, not a shorter batch cadence.
- How does naming a run after its interval make a re-run safe?The interval determines both the source filter and the output partition, so re-executing the same interval reads the same rows and overwrites the same target. The result is idempotent by construction: a retry, a manual rerun and a historical backfill all converge on identical output instead of appending duplicates.
- What happens to rows that arrive after their interval has closed?Nothing sees them. The run for that window already completed against the data present at the time, and no later run covers that range. Handling late data requires a deliberate reprocessing window — periodically re-running the last N intervals — or a merge keyed on the record's own event time.
A monthly phone bill is issued on the 1st but covers the month that just ended. The issue date and the billing period are two different things, and nobody expects a bill issued on 1 April to cover April.
saying these in an interview costs you the question
- Says the run's timestamp is simply when the job started
- Filters source rows with now() instead of the interval bounds
- Assumes the midnight run contains that same day's data
- Claims freshness equals the schedule frequency
- Thinks late-arriving rows get picked up automatically