skip to content

How do you schedule a nightly Python job across DST so it neither repeats nor skips?

level: seniorimportance: should knowfreq 40%

answer

  1. First ask what the schedule promises
  2. Absolute cadence versus local wall clock
  3. Recompute the local time each cycle
  4. Never add a day to a local value
  5. Idempotency key on the logical run date

basics

~20 s

Decide whether the job is anchored to an absolute cadence or to a local wall clock. Absolute means a stored UTC instant plus a fixed interval; wall-clock means recomputing the local time each night and resolving gaps and repeats explicitly.

solid answer

~50 s

Pick the intent before touching an API. If the job just has to run every 24 hours, anchor it to a UTC instant and add `timedelta(hours=24)`; DST becomes irrelevant. If it must track the local business day, recompute each cycle: `datetime.combine(tomorrow, time(1, 0), tzinfo=zone)`, then `astimezone(timezone.utc)` for the runner to wait on. Never derive the next run by adding a day to the previous *local* value — that is wall-clock arithmetic, it resets `fold` to 0, and it can land on an hour that does not exist. Detect a gap with the UTC round-trip test and log the shift; detect duplicates by keying each run on its logical business date rather than a timestamp, so a second start on the repeated hour is rejected by the store. Finally, measure durations on UTC instants: a six-hour window in wall-clock terms is five or seven real hours across a transition.

code

python · 13 lines
python
from datetime import date, datetime, time, timezone
from zoneinfo import ZoneInfo


def next_run(day, wall, zone):
    local = datetime.combine(day, wall, tzinfo=zone)
    instant = local.astimezone(timezone.utc)
    return instant, local != instant.astimezone(zone)


ny = ZoneInfo("America/New_York")
print(next_run(date(2025, 3, 9), time(2, 30), ny))
print(next_run(date(2025, 11, 2), time(1, 0), ny))

go deeper

for a junior

Know that a job pinned to a local time can run twice or not at all on the two DST nights, and that converting to UTC before waiting is part of the fix.

for a middle

Explain the difference between absolute and wall-clock anchoring, and show recomputing the next local run from the calendar date and zone instead of adding a day to the last one.

for a senior

Demonstrate the operational judgement: gap and ambiguity policies chosen and logged, idempotency keyed on the logical run date, durations measured on instants, and tests pinned to the real transition dates.

for a principal

Own the cross-service contract — which service holds the schedule's intent, what the API returns, and how a tz data update is rolled out — so no downstream consumer has to re-derive the policy itself.

## Anchoring a recurring job A ticket-triage bot does a 6-hour sweep every night, starting at 01:00 local time in `America/New_York`. Twice a year it misbehaves: on the November fall-back night the 01:00 start happens twice, and on the March spring-forward night a 02:00-anchored variant of the same job never fires at all. The bug is not in `datetime`; it is in never having decided what "every night at 01:00" *means*. There are exactly two intents, and the first job of the design is to pick one. **Absolute cadence.** The job must run every 24 hours. Then the anchor is a UTC instant: store the last (or next) run as an aware UTC datetime and add `timedelta(hours=24)` to it. DST never touches the schedule, and the job drifts against the local clock by an hour for half the year — which is fine for a sweep nobody watches. **Wall-clock anchor.** The job must run when the local business day is quiet, so it must track the local clock. Then the anchor is a local time-of-day plus an IANA zone key, and the next run is *recomputed* each cycle: take tomorrow's local date, combine it with 01:00 and the zone, resolve ambiguity and gaps explicitly, and convert the result to a UTC instant that the runner waits on. Never compute the next run by adding `timedelta(days=1)` to the previous *local* value: that is wall-clock arithmetic, it resets `fold` to 0, and it happily lands on an imaginary time. ### Resolving the two anomalies For the repeated hour, the wall-clock recomputation is what saves you. `datetime.combine(day, time(1, 0), tzinfo=zone)` produces one datetime with `fold=0`; converting it to UTC yields exactly one instant, the first pass, and the second pass simply is not on the schedule. The job runs once. Trouble only appears when something *else* also decides when to run — which is where the classic encoding mismatch bites: the schedule is written in config as a naive local string while the runner records `last_run` as a UTC epoch, and the comparison between the two is done after formatting the epoch back into local time. On the fall-back night that formatting is ambiguous, `last_run` renders as 01:00 for a second time, the "has it run yet today?" test says no, and the sweep starts again. Fix it by comparing instants to instants and never round-tripping through a local string, and make the run key a logical business date rather than a timestamp, so a duplicate start is rejected by the store rather than by arithmetic. For the gap, decide and encode the policy. With `fold=0`, `astimezone(timezone.utc)` already shifts an imaginary 02:30 forward to the instant that renders as 03:30 local — usually the right answer for a batch sweep — but detect it (`local != local.astimezone(timezone.utc).astimezone(zone)`) and log that the shift happened rather than letting it be invisible. The cheapest fix of all is to move the anchor: no wall-clock schedule should be pinned inside 01:00–03:00 in a DST zone if it does not have to be. ### Duration is the third trap The sweep is "6 hours". Six hours of what? `start + timedelta(hours=6)` is six hours of *wall clock* inside the zone: on the fall-back night 01:00 plus six wall hours is 07:00 local, which is seven real hours later, so the run overruns its window and can collide with the next stage of the pipeline. On the spring-forward night the same arithmetic gives five real hours and the sweep is cut short. Timeouts, budgets and "did this take too long?" checks must be computed on UTC instants; only the *display* belongs in local time. ### What to demonstrate under questioning Say which intent the job has before touching an API. Show the recompute-from-local-date loop rather than local-plus-a-day. Name the two anomalies and the policy for each. Point at idempotency — a unique run key on the logical date — as the thing that makes a double start harmless even if the scheduling logic is wrong. And say how you would test it: pin the two transition dates for the zone as fixtures and assert the computed UTC instants, because these paths execute twice a year and will never be exercised by chance in CI.

  • Why is adding timedelta(days=1) to the previous local run time the wrong way to schedule?
    Arithmetic on an aware datetime is wall-clock arithmetic inside the zone: it adds to the naive fields, keeps the tzinfo, and resets fold to 0. Across a transition the result is 23 or 25 real hours away, it can land inside the repeated hour with the fold silently chosen for you, and it can land on an hour that never occurs. Recomputing from the local calendar date plus the zone avoids all three.
  • How would you test scheduling code that only misbehaves twice a year?
    Pin the transition dates for the target zones as explicit fixtures and assert the computed UTC instants for the night before, the transition night and the night after — both directions. Include an anchor inside the gap and one inside the repeated hour. Also pin the tz database version the tests run against, so a routine data update fails loudly rather than shifting assertions under you.
  • A duplicate run still slipped through on the fall-back night. Where do you look first?
    At any place a UTC value is formatted into local time and then compared or stored as a string. That round trip is ambiguous during the repeated hour, so a last-run marker can render as the same local time twice and a 'has it run today?' test answers no. Compare instants to instants, and back it with a unique key on the logical run date so a second start cannot commit.

A wall-clock schedule is a promise to the room's clock, an absolute schedule a promise to a stopwatch. Twice a year the two disagree, and code that never said which it meant picks the wrong one.

saying these in an interview costs you the question

  • Adds a day to the previous local datetime each cycle
  • Assumes UTC storage alone fixes a wall-clock schedule
  • Compares last-run markers after formatting them as local strings
  • Treats a wall-clock interval as elapsed real time
  • Relies on the job being naturally idempotent without a run key
  • Believes the scheduler layer resolves DST ambiguity for you

context