A nightly job is expected to run at 02:00 local time. Consider what should happen when the machine was powered off across that time, when the system clock is corrected backwards by an hour, and when a daylight-saving transition means the local hour 02:00 occurs twice or not at all. How would you define the job's behavior?
answer
- monotonic clock for delays, wall clock for calendar times
- policy per job: replay each / coalesce to one / skip and re-anchor
- key each run by OCCURRENCE, persist last-completed key
- clocks back ⇒ ambiguous hour, fire once; clocks forward ⇒ nonexistent hour, fire at transition
- store zone + local time; multi-node needs a shared completion record
basics
~20 sDecide per job whether a missed occurrence must be replayed, coalesced into one run, or skipped; make runs idempotent and keyed by the occurrence they represent. Use a monotonic clock for relative delays and an explicit time zone for calendar times, and treat a doubled or missing local hour as a fire-once decision.
solid answer
~60 sThree distinct problems. **Downtime**: occurrences elapse with nothing running. The policy must be explicit per job — *replay each* (only when occurrences represent distinct work, e.g. per-day billing, and then rate-limited), *coalesce to one* (the usual answer for state-converging work), or *skip and re-anchor* (when the run is worthless if late). Encode the choice, do not inherit a library default. **Clock correction**: relative delays must be measured on a monotonic clock, so a backward step cannot postpone them by the size of the correction and a forward step cannot fire everything at once. Calendar schedules must use wall-clock time, but keyed to an occurrence identity so a repeated instant does not re-fire. **Daylight saving**: with an ambiguous local hour, 02:00 exists twice — fire once, on the first. With a skipped hour, it never exists — fire at the transition instant or the next valid time; do not silently miss a day. Storing the schedule as a time zone plus local time (not a fixed offset) is what makes either decision expressible. Underneath all of it: make each run idempotent and identified by its occurrence key, so replay is safe and duplicates are cheap.
code
text · 20 linesschedule = { zone: "Europe/Berlin", localTime: "02:00", policy: COALESCE }
reconcile():
last = store.lastCompletedOccurrence(job) # e.g. 2026-08-04
missed = occurrencesBetween(last, wallNow(), schedule)
switch schedule.policy:
REPLAY: for occ in missed: runRateLimited(job, occ); store.complete(occ)
COALESCE: if missed nonempty:
run(job, occurrence = missed.last)
store.completeAll(missed) # one run, all marked done
SKIP: store.completeAll(missed); metrics.skipped += missed.size
run(job, occ):
if store.isComplete(occ): return # duplicate-safe (clock went back)
doWork(forOccurrence = occ) # idempotent, keyed by occ
# DST cases
ambiguous local 02:00 (clock back): two instants map to occ 2026-10-25 -> run once
nonexistent local 02:00 (clock fwd): no instant maps -> run at transition instantgo deeper
Know the basics: after downtime the job may need to run once to catch up or not at all, and clock changes can make a nightly time happen twice or not at all — so the job should be safe to run twice.
Distinguish monotonic time for delays from wall-clock time for calendar schedules, and name the three missed-run policies: replay each, coalesce to one, skip and re-anchor.
Design it: occurrence keys with a persisted last-completed marker, idempotent runs keyed by occurrence, rate-limited replay, and explicit handling of the ambiguous and nonexistent local hours.
Own the contract across the fleet: which job classes get which policy, shared occurrence claiming across nodes with clock skew, reconciliation on startup, and how missed or duplicated runs are surfaced to operators and to downstream consumers.
## Why this is a design question, not a configuration question "Run at 02:00" is under-specified. It does not say what 02:00 means when the clock is unreliable, what happens when the machine was not there, or whether the run is about *an instant* or about *a unit of work*. Answering it well means deciding what the job's occurrences mean. ## Two clocks, two jobs - A **monotonic clock** only moves forward at a steady rate and is unaffected by time synchronization or human adjustment. It has no meaning as a date. It is the correct basis for *relative* delays: "in 30 seconds" must mean 30 seconds of elapsed time. - A **wall clock** names a civil instant and can jump forward or backward. It is the correct basis for *absolute* schedules: "at 02:00 on Tuesday." Mixing them is a classic defect. If a relative delay is computed as `wallNow + 30 s` and the clock is corrected 60 s backwards, the task fires 90 s later than intended; if corrected forwards, everything pending fires at once. Conversely, a calendar schedule computed purely from elapsed monotonic time drifts off the calendar within weeks and cannot express "the first of the month" at all. ## Missed runs: define the policy per job When occurrences elapse with nothing running — process down, host powered off, virtual machine suspended, deployment window — you must choose: **Replay each missed occurrence.** Right only when each occurrence corresponds to distinct, non-fungible work: aggregate the 3rd, then the 4th, then the 5th. Two conditions apply. Runs must be idempotent and keyed by occurrence ("aggregate for 2026-08-04"), not by "now," so replay produces correct results. And replay must be rate-limited, because a system recovering from downtime is the worst moment for a burst. **Coalesce to a single run.** Right for state-converging work: refresh a cache, rebuild an index, sweep expired rows. Running five times achieves nothing that one run does not. This is the most common correct answer. **Skip and re-anchor.** Right when a late run is worthless or actively wrong: a market-open snapshot at 11:00 is not a snapshot; a nightly report delivered mid-morning may be worse than none. Skip, record the miss, alert. The defect to avoid is having no policy — inheriting whatever the scheduler does by default and discovering it during an incident. ## Clock corrections A backward correction creates an interval of wall-clock time that occurs twice. A naive scheduler comparing "is it 02:00 yet?" against wall time will fire the job a second time. A forward correction skips an interval, and a naive scheduler may never see the target instant at all. The fix is to schedule against an **occurrence identity** rather than an instant: compute the next occurrence, record that occurrence's key (its date and nominal time), and refuse to run an occurrence whose key has already completed. This makes both duplicates and misses detectable and decidable rather than emergent. Persisting the last-completed occurrence key also makes the schedule survive restarts, which is what turns "missed runs" from an accident into something you can reason about at startup. ## Daylight-saving transitions Two cases, both real, both worth naming explicitly: - **Ambiguous local time** (clocks go back): 02:00 occurs twice in the same local day. Firing twice is almost always wrong — a duplicate nightly report, a double billing run. Default to firing on the **first** occurrence and marking that day's occurrence complete. - **Nonexistent local time** (clocks go forward): 02:00 never occurs; the clock jumps from 01:59:59 to 03:00:00. The naive scheduler simply misses that day. Choose deliberately: run at the transition instant, or at the next valid local time. Silently skipping a day is a data-integrity bug in a job that produces daily artifacts. Both require storing the schedule as **time zone + local time** (e.g. "Europe/Berlin, 02:00"), not as a fixed UTC offset — an offset cannot express a transition, and a job pinned to UTC will simply drift an hour away from the local time the business cares about. Where the exact local hour does not matter, the simplest robust answer is to schedule in UTC and accept the local shift; state that as a decision, not an oversight. ## Multi-node schedules Once more than one instance of a service runs, every node has its own clock and its own view of missed runs. Two consequences: the occurrence-completion record must be **shared** (a lease or a completed-occurrences table), not in-process; and node clock skew means "02:00" differs slightly per node, so the leader for an occurrence must be decided by the shared record, not by whoever thinks it is time first. ## The design in one paragraph Schedule calendar work as time zone plus local time, expand it to concrete occurrences with stable keys, persist the last completed occurrence, and on startup or clock change reconcile: for each occurrence between the last completed key and now, apply the job's declared missed-run policy — replay (rate-limited), coalesce, or skip — recording each decision. Measure relative delays on a monotonic clock. Make each run idempotent under its occurrence key, so duplicates are harmless and replay is correct. Every one of those sentences is a decision someone must make; a scheduler that makes them implicitly is a scheduler that will surprise you exactly once a year.
- A nightly aggregation job was down for three days. Should it run three times or once?It depends on whether each occurrence represents distinct work. An aggregation that produces one artifact per day — a daily total, a daily report — must replay all three, keyed by the day each represents rather than by the current time, and rate-limited so the recovering system is not flooded. A job that converges state, such as rebuilding an index or refreshing a cache, should run once, because three identical runs reach the same end state at triple the cost.
- Why is it wrong to compute a relative delay by adding to the wall-clock time?Because wall-clock time can jump. If the clock is corrected backwards by a minute, a task due in 30 seconds now appears due in 90; if corrected forwards, a batch of pending tasks becomes due simultaneously and fires as a burst. Relative delays mean elapsed time, which is what a monotonic clock measures, so it is unaffected by synchronization and manual adjustment. Wall-clock time remains the right basis for absolute calendar schedules.
- How does this change when the same scheduled job runs on several nodes?Each node has its own clock and its own idea of what has run, so in-process state cannot arbitrate. The record of completed occurrences must live in shared storage, and a node must claim an occurrence there — via a lease or a conditional insert on the occurrence key — before executing it. Clock skew between nodes means several may believe it is time at slightly different moments, so the shared claim, not local time, decides who runs it.
A newspaper delivery. If nobody delivered for three days, do you deliver all three back issues (replay), just today's (coalesce), or none because the news is stale (skip)? The right answer depends entirely on whether each issue is distinct content or just the current state of the world — and the delivery contract has to say which.
saying these in an interview costs you the question
- Computing relative delays from wall-clock time, so time synchronization shifts or bursts the schedule.
- Storing a schedule as a fixed UTC offset instead of a time zone, so daylight-saving transitions cannot be expressed.
- Assuming a nightly job fires exactly once on the day clocks go back — the local hour occurs twice.
- Replaying every missed occurrence unconditionally, flooding a system that just recovered.
- Keying work by "now" instead of by the occurrence it represents, and relying on in-process state to prevent duplicates across nodes.