After a cluster-wide cron was down for six hours, how should each job decide whether to catch up every missed tick, fire once, or skip?
answer
- silence is not success
- last recorded tick versus schedule
- does each tick own distinct work
- value decays with lateness
- throttled, ordered, capped replay
basics
~20 sChoose by what each run means: per-period work such as billing catches up every missed tick, state-sweeping work such as cleanup fires once, and work worthless when late skips. Detect the gap from the last recorded tick, not from now.
solid answer
~50 sA scheduler that only looks forward from *now* silently drops whatever fell inside an outage, so on restart or leader takeover it should compare each job's **last recorded tick** with its schedule to list what was missed. Then it applies that job's **missed-fire policy**. *Catch up all* fits jobs where each tick carries distinct work: a per-day billing run must happen for each missed day, in order and throttled. *Fire once* (coalesce) fits jobs that process current state regardless of time: a minutely cleanup that missed 360 ticks needs one run, not 360. *Skip* fits jobs where a late run is worse than none: an 08:00 digest sent at 14:00. Catch-up must still pass the overlap guard and be idempotent per tick, automatic replay should be capped with an operator alert beyond the cap, and every skipped tick should be recorded so the gap stays visible.
code
pseudocode · 13 lineson scheduler start or leadership takeover:
for job in jobs:
last = last_recorded_tick(job)
missed = ticks_between(job.cron, after = last, until = now)
if missed is empty: continue
if job.policy == CATCH_UP_ALL:
if size(missed) > job.max_catch_up: alert_operator(job, missed)
else: for tick in missed, oldest first: enqueue_claim(job, tick)
else if job.policy == FIRE_ONCE:
record_skipped(job, all_but_last(missed))
enqueue_claim(job, last_of(missed))
else:
record_skipped(job, missed)go deeper
Recall that a scheduler that was down misses ticks, and that there are three basic reactions: run each missed tick, run once, or skip.
Explain how a scheduler finds missed ticks from its run history on start-up or takeover, and match each of the three policies to a kind of job.
Show you can replay safely: ordered, throttled catch-up behind the overlap guard, idempotent per tick, with skipped ticks recorded and an operator alert beyond a cap.
Make the missed-fire policy a required, reviewed field of every job definition, because the right answer depends on business meaning a platform team cannot infer.
## Why missed fires happen in a fleet A **missed fire** (or misfire) is a scheduled tick that passed without a run. In cluster-wide cron it happens whenever nobody was in a position to fire: - the whole fleet was down for a deploy, an incident or a maintenance window; - the shared claim store was unreachable, so no node could win a claim; - leadership was changing hands and the tick fell into the gap; - every node was saturated and the scheduler loop was starved. Nothing fails when a tick is missed, because nothing ran. That makes the default behaviour of many schedulers dangerous: on start-up they compute *the next fire time after now* and carry on, so an outage's ticks simply vanish. ## Detecting what was missed Detection needs durable state that outlives any node: 1. Keep a **run history** with one row per job per scheduled tick, recording status such as running, succeeded or skipped. 2. On scheduler start-up or leadership takeover, read each job's **last recorded tick**. 3. Enumerate the cron expression's instants between that tick and now; those are the missed ticks. 4. Hand the list to the job's policy instead of discarding it. ## The three policies | Policy | What runs | Fits | Risk if misapplied | |---|---|---|---| | Catch up all | One run per missed tick, oldest first | Each tick owns distinct work, such as a billing day or an hourly report period | A long outage replays a flood of runs | | Fire once | One run standing in for every missed tick | Each run processes current state, such as a cleanup sweep or cache refresh | Per-period work silently loses periods | | Skip | Nothing until the next scheduled tick | The run's value decays with lateness, such as a morning digest or a time-boxed promotion | Needed work is lost | Many schedulers also apply a **misfire threshold**: a tick older than it is treated as skipped rather than replayed. The threshold should come from the job's usefulness window, not from a platform default nobody chose. ## Matching the policy to the job Using the archetype of nightly billing, an hourly report and a minutely cleanup after six hours of downtime: - **Nightly billing** that missed its 02:00 tick must catch up; each run bills one day, and coalescing would leave a day unbilled. - **Hourly reports** missed six ticks; if each report covers its own hour, catch up all six in order, but if it is a "current snapshot" report, fire once. - **Minutely cleanup** missed 360 ticks (6 x 60); one run sweeps everything that accumulated, so fire once. - **A 09:00 reminder** whose moment has passed should usually skip, because a reminder delivered at 14:00 annoys more than it helps. The decision depends on business meaning, which is why it belongs in each job's definition rather than in a single platform-wide setting. ```pseudocode on scheduler start or leadership takeover: for job in jobs: last = last_recorded_tick(job) // from run history, UTC missed = ticks_between(job.cron, after = last, until = now) if missed is empty: continue if job.policy == CATCH_UP_ALL: if size(missed) > job.max_catch_up: alert_operator(job, missed) // too far behind to replay blindly else: for tick in missed, oldest first: enqueue_claim(job, tick) else if job.policy == FIRE_ONCE: record_skipped(job, all_but_last(missed)) enqueue_claim(job, last_of(missed)) // one run stands in for all else: // SKIP record_skipped(job, missed) ``` ## Catching up safely Replaying is itself a load event, so it needs guard rails: - **Order**: run missed ticks oldest first when later periods depend on earlier ones. - **Throttle**: one or a few at a time, so the database and downstream providers are not hit by a burst. - **Overlap guard**: catch-up runs pass through the same "is this job already running" check as normal ticks. - **Idempotency**: each replayed run is keyed by its own scheduled tick, so a replay that races with a late original does no double work. - **Cap**: beyond a set number of missed ticks, stop and page someone; thirty days of unbilled customers deserves a human decision. ## Recording the decision Every skipped or coalesced tick should still get a run-history row marked as such. That keeps gap detection honest, since a skipped tick is a decision rather than a hole, and it lets anyone answer later why a given period has no output.
- Why does a scheduler that only looks forward from now cause silent gaps?On start-up it computes the next fire time after the current moment, so any tick that passed while no scheduler or leader was alive is never considered, and because nothing ran, nothing failed. Persisting the last recorded tick per job and comparing on start-up turns that invisible gap into an explicit policy decision.
- When catching up 30 missed daily billing runs, what goes wrong if they all fire at once?Thirty runs hit the database and payment provider together, contend on the same rows and can breach downstream rate limits. Replay them oldest first, one or a few at a time behind the overlap guard, and hold the new nightly tick until the backlog clears so the periods stay in order.
- How do you choose a misfire threshold for a job?Set it from the job's usefulness window: how late can a run be and still be worth doing? For a digest that may be an hour. For billing it is effectively unbounded, so rather than a threshold you cap automatic catch-up and page an operator beyond the cap.
saying these in an interview costs you the question
- The scheduler will automatically run whatever it missed.
- Catching up every missed tick is always the safe choice.
- A minutely cleanup should replay each of its missed ticks.
- Skipped ticks need no record because nothing actually ran.
- One global missed-fire setting is enough for every job.