Last night's reconciliation run is still going when the recurring time trigger fires again tonight — what are the platform's options, and how do you choose?
answer
- a clock has no feedback from a duration
- skip, allow, replace
- what would two writers do?
- a skipped period is lost unless parameterised
- lease plus idempotence beats the setting
basics
~20 sThree options: skip tonight's run, start it anyway alongside the old one, or cancel the old one and start fresh. Which is right depends on whether two runs can safely touch the same data and whether a skipped period is ever processed later.
solid answer
~50 sA recurring schedule fires on a clock, and a clock knows nothing about how long the previous run took — so a run that overruns its interval collides with the next trigger. Platforms expose an overlap rule with roughly three settings. **Skip** keeps exactly one run alive and is right when concurrent runs would corrupt shared state, but it silently drops a period unless the run is parameterised by the period it processes. **Allow** starts both, which is only safe when runs are genuinely independent and partitioned. **Replace** cancels the running one and starts the new trigger, which suits a run whose output is a full refresh and is dangerous for one that is halfway through applying changes. The durable fix is usually not the setting: make the work idempotent, take a lease so the job protects itself regardless of what the schedule does, and alert on a run whose duration approaches its interval.
code
pseudocode · 14 lineson trigger(schedule, scheduledTime):
if now() - scheduledTime > schedule.catchUpWindow:
record(MISSED, scheduledTime) # too late to be worth running
return
if isActive(previousRun(schedule)):
if schedule.onOverlap == "skip":
record(SKIPPED, scheduledTime)
return
if schedule.onOverlap == "replace":
cancel(previousRun(schedule)) # then fall through and start
# "allow" falls through with the previous run still working
startRun(schedule, forPeriod = scheduledTime)go deeper
A recurring schedule fires on a clock and does not wait for the previous run. Know that a platform gives you a rule for the collision: skip the new one, run both, or cancel the old one.
Explain each rule's hazard: allow risks two writers on the same data, skip silently drops a period unless the run takes that period as a parameter, and replace can leave a half-applied cancellation behind.
Reason from the data rather than the setting. Say what two concurrent runs would do to shared state, whether a skipped period is ever recovered, and add the lease, the idempotent design and the duration-against-interval alert that stop this being a nightly gamble.
Set the standard: every scheduled batch takes its period as an input, protects itself with a lease, declares a catch-up window, and is watched on duration against interval. That turns a class of quiet nightly data loss into a property the platform can be audited for.
## Why a schedule collides with itself A recurring schedule is a factory: at each trigger time it creates a run-to-completion job. The trigger is driven by a clock and has no feedback from the duration of the previous run. So the moment a run's duration exceeds the interval between triggers — a night when the data volume doubled, a slow dependency, a retried attempt — the next trigger arrives while the previous run is still working. That is not an exotic edge. A nightly batch that normally takes four hours is one bad night away from it, and the first time it happens is usually in production with real data on the other side. ## The three overlap rules | Rule | What the platform does | Right when | The hazard it carries | |---|---|---|---| | **Skip** | leaves the running one alone, records the trigger as skipped | concurrent runs would corrupt shared state | a period is silently unprocessed unless runs are period-parameterised | | **Allow** | starts the new run beside the old one | runs are independent, partitioned, or read-only | two writers on the same rows; doubled load on the same dependency | | **Replace** | cancels the running one, then starts the new trigger | the output is a full refresh and a stale half-run is worthless | a cancelled run may have applied half its effect | Note what none of them do: none of them makes the work finish faster, and none of them fixes a batch that no longer fits in its window. They only decide who wins a collision. ## Choosing, in the order the decision actually goes 1. **Ask what two concurrent runs would do to shared state.** If the answer is "double-apply" or "deadlock", `allow` is off the table before anything else is discussed. 2. **Ask whether a skipped period is lost.** If each run is parameterised by *which period it processes* — reconcile the day ending at this trigger time — then a later run can cover the gap, and `skip` is cheap. If the run just processes "whatever looks stale right now", the skipped period is genuinely never handled and `skip` quietly loses a day of reconciliation. 3. **Ask what a cancellation leaves behind.** `replace` is safe for a run that computes and then publishes atomically, and unsafe for one that streams changes out as it goes, because cancelling it leaves the world in a partial state that the next run must be able to tolerate. 4. **Only then set the rule** — and record why, because the reasoning is not recoverable from the setting alone. ## The sibling case: the trigger that was missed entirely The same schedule has a second failure mode. If the part of the platform that fires triggers was unavailable across a trigger time, that run never started. When it comes back, two behaviours are possible and platforms differ on the default: fire the missed trigger late, or decide it is too old to be useful and record it as missed. Both are defensible, which is why a catch-up window usually exists to express the judgment — start it if we are within this much of the scheduled time, otherwise skip and say so. The danger with unbounded catch-up is a burst: several missed triggers all firing at once on recovery, hitting the same dependency together — at the worst moment, immediately after an outage. The danger with no catch-up at all is a period that is never processed. The way out of both is the same as above: make the run take the period it is processing as an input, so a missed period can be run deliberately rather than depending on the schedule to remember. ## The fix that outlives the setting The overlap rule is the platform's answer. A robust batch does not depend on it: - **Take a lease.** The run acquires an exclusive lease on the work at start and refuses to proceed if one is held. Now a second run cannot damage anything even if the schedule allows it, and the protection survives being started by hand. - **Be idempotent.** Re-processing a period must land in the same state as processing it once. This is what makes both retries and catch-up runs safe. - **Take the period as a parameter.** `reconcile the window ending at T` rather than `reconcile whatever is stale`. Skipped and late runs become recoverable instead of lost. - **Alert on duration against interval.** A run using most of its interval is the warning that the collision is coming. That is the alert that actually prevents the incident; the overlap rule only decides how it fails. ## What an interviewer is listening for The named settings are the easy half. The signal is whether you reason from the data — what two writers would do, what a cancellation leaves, whether a skipped period is recoverable — and whether you reach for the lease and the period parameter rather than treating the schedule's setting as the whole of the answer.
- Why does an unbounded catch-up window turn an outage into a second incident?Because every trigger missed during the outage fires on recovery, all at once, against dependencies that have just come back and are already cold. A six-hour outage of a fifteen-minute schedule queues two dozen runs that start together. A catch-up window bounds that by declaring how late a run may still be worth starting, and anything older is recorded as missed for a human to decide about.
- If the job takes a lease anyway, does the overlap rule still matter?Yes, but for cost rather than correctness. With a lease, a second run cannot damage anything — it finds the lease held and stops. The overlap rule still decides whether that second run is created at all, and creating runs that immediately exit produces noise in the run history and wasted scheduling. Set it to skip and keep the lease as the guarantee that does not depend on the setting.
- What single alert would have warned you before the collision happened?Run duration measured against the trigger interval. A nightly batch that has crept from four hours to seven is heading for a collision, and the trend is visible weeks ahead. Alerting on the collision itself tells you at the moment it is already too late; alerting on the ratio gives you the window to shorten the work or lengthen the interval.
- Why is `replace` unsafe for a run that streams its changes out as it goes?Because cancelling it stops the run partway through applying effects, so the world is left in a state no complete run would ever produce — some records updated, some not, and nothing recording where it stopped. It is safe for the opposite shape: a run that computes fully and then publishes in one step, where a cancelled run has published nothing at all.
saying these in an interview costs you the question
- Assumes the platform waits for the previous run to finish by default
- Says overlapping runs are fine because each has its own instance
- Picks replace without asking what a cancelled run leaves behind
- Thinks skipping a trigger is free even when nothing processes that period later
- Wants unbounded catch-up so no run is ever lost, ignoring the recovery burst
- Treats the schedule's setting as sufficient protection for shared data