What happens to a pipeline scheduled in a local timezone when that zone's daylight-saving change occurs?
answer
- the local clock is not continuous
- one hour vanishes, another repeats
- a local day is not always 24 hours
- hardcoded 24-hour arithmetic breaks twice a year
basics
~20 sOne local wall-clock hour disappears in spring and one repeats in autumn, so a schedule inside those hours may be skipped or fire twice, and that day's nominal daily interval is 23 or 25 hours instead of 24.
solid answer
~50 sA schedule defined in local time is a wall-clock pattern, and twice a year the wall clock is not continuous. On the spring change, an hour of local time never occurs — a schedule falling inside it has no matching instant, and different schedulers resolve that differently (skip the run, or fire once at the shifted time). On the autumn change, an hour repeats, so the same local time occurs twice and the run may fire twice unless the scheduler deduplicates. Just as important, the **intervals stop being uniform**: a local day is 23 or 25 hours on those dates, and an hourly schedule produces 23 or 25 windows. Anything that assumes "24 partitions per day" or computes a window by adding 24 hours will be wrong. The robust choices are to define schedules in UTC where the business allows it, to keep tasks idempotent so a duplicate firing is harmless, and to avoid scheduling anything in the affected local hours.
code
text · 9 linesspring change, local schedule at 02:30
01:59:59 -> 03:00:00 (02:00-02:59 never occurs)
02:30 has no matching instant -> run skipped or shifted, by scheduler policy
that local day is 23 hours long
autumn change, local schedule at 01:30
01:59:59 -> 01:00:00 (01:00-01:59 occurs twice)
01:30 matches two instants -> possible duplicate firing
that local day is 25 hours longgo deeper
Know that schedules carry a timezone and that daylight-saving changes can move, skip or repeat a run; do not assume a schedule string means UTC.
Explain both changeover directions and why a local day can be 23 or 25 hours, and spot the hardcoded-24 assumption in window arithmetic.
Show the production angle: a skipped run raises no failure alert, so coverage checks over expected intervals and idempotent partition writes are what actually protect you.
Decide and enforce the organisational convention — UTC by default, local only where the window's business meaning demands it — and make sure source, schedule and partition labels share one timezone story.
## Why a local schedule is not a timeline A cron-style schedule such as "03:15 every day" is a pattern matched against a *wall clock*. UTC's wall clock is continuous — every second occurs exactly once — but a local clock in a zone that observes daylight saving is not. Twice a year it jumps. **Spring forward.** The clock advances, and a block of local time never happens. In the affected hour there is simply no instant with that local reading, so a schedule inside it has nothing to match. Schedulers resolve this by policy: some skip the run entirely for that date, some fire it once at the shifted instant. Which one your tool does is a documentation question, not something to guess at in an interview — the point to make is that it is *policy*, not physics. **Fall back.** The clock rewinds, and a block of local time happens twice. Now a single schedule expression matches two distinct instants. Depending on the implementation you may see a duplicate run, or one run with the ambiguity resolved to the first or second occurrence. ## The bigger problem: intervals stop being uniform The skipped or duplicated firing is the headline, but the more damaging effect is on interval arithmetic. If the schedule is anchored in local time: - a daily interval on the spring date is **23 hours** long; - a daily interval on the autumn date is **25 hours** long; - an hourly schedule yields **23 or 25 intervals** for that local day. Any code that hardcodes 24 — computing the window end as "start plus 24 hours", asserting exactly 24 partitions per day, dividing a daily total by 24 to get an hourly average, or comparing today's row count to yesterday's — is wrong on those two dates. The bugs are seasonal, appear on exactly two days a year, and are usually attributed to the source system before anyone suspects the clock. A related trap: the *source data* has its own timezone convention, which may not match the schedule's. If events are stamped in UTC and the window bounds are computed in local time, the boundary shifts by an hour twice a year, and rows get assigned to the neighbouring day. Reconciliation reports show one day short and the next day long, on exactly the changeover date. ## Choosing a timezone deliberately **Schedule in UTC.** No skipped hour, no repeated hour, every interval exactly the same length, and window arithmetic is plain addition. The cost is that the pipeline drifts one hour relative to local business hours twice a year — a report that landed at 06:00 local now lands at 05:00 or 07:00. For most internal data pipelines this is entirely acceptable and is the default worth arguing for. **Schedule in local time.** Correct when the *business meaning* of the window is local — a retailer's trading day, a payroll cutoff, a regulatory daily close. Then the 23-hour and 25-hour days are not a bug, they are the truth about that business day, and the pipeline must handle them rather than normalise them away. Keep the schedule out of the changeover hours, make tasks idempotent so a repeated firing overwrites rather than appends, and never compute window ends by adding a fixed number of hours — derive them from a timezone-aware calendar. **Do not mix.** The worst outcome is a schedule in one zone, source timestamps in another, and partition labels in a third. Pick one convention for window boundaries, state it in the pipeline's documentation, and use it everywhere. ## Operational hygiene Mark the two changeover dates in the team calendar and check the pipelines that matter on those mornings — this is one of the few classes of bug where the reproduction window is known a year in advance. Alert on *interval coverage* rather than run success: a check that every expected window has an output partition catches the skipped spring run, which otherwise fails silently because there is no failed run to alert on. And make every task idempotent, which turns the autumn duplicate firing from a data-corruption incident into a harmless rewrite. ## How to answer Name both directions (skipped hour, repeated hour), say that the resolution is scheduler policy rather than something universal, then pivot to the point most candidates miss: interval lengths become non-uniform, so any hardcoded 24 is a latent seasonal bug. Close with the recommendation — UTC unless the window's business meaning is genuinely local, plus idempotent tasks and coverage-based alerting.
- Why can a skipped spring run be harder to notice than a duplicated autumn run?Because nothing fails. There is no run, so no failure alert, no error log, and no partial output — just a missing partition that downstream queries silently read as zero rows. Detection requires a coverage check that asserts every expected interval produced output, not a check on run outcomes.
- When is scheduling in local time the right choice despite the complications?When the window's business meaning is local: a store trading day, a payroll or regulatory cutoff, a market close. Normalising those to UTC produces windows that are technically uniform but describe the wrong period. Accept the 23- and 25-hour days as correct, and make the code handle variable day length.
- What single defensive property makes the autumn duplicate firing harmless?Idempotence. If the task overwrites a partition keyed by its interval rather than appending, running the same window twice converges on identical output. The duplicate becomes a wasted run instead of a data-quality incident, which is why partition-overwrite semantics matter more than trying to prevent the second firing.
saying these in an interview costs you the question
- Assumes every scheduler handles the missing hour the same way
- Computes the window end by adding a fixed 24 hours
- Says daylight saving only shifts run times, not interval lengths
- Mixes local-time schedules with UTC source timestamps
- Relies on failure alerts to notice a skipped run