skip to content

In cluster-wide cron with per-tick claims, how do you stop an overrunning minutely cleanup from overlapping the next tick's run on another node?

level: seniorimportance: nice to knowfreq 30%

answer

  1. different tick, different key
  2. guard keyed by job alone
  3. lease, not a plain flag
  4. skip, queue one, replace
  5. duration to interval ratio

basics

~20 s

Per-tick claims never block a different tick, so add a job-level guard: a shared 'running' lease with an owner and expiry, renewed while the run lives. The next tick's winner checks it and skips, queues or replaces the old run.

solid answer

~50 s

The claim on `(cleanup, 02:01)` has nothing to do with the claim on `(cleanup, 02:00)`, so if the 02:00 run is still going on node A when 02:01 arrives, node B wins the new tick and both run, contending on the same rows. The fix is a second, **job-level guard**: a shared record saying 'cleanup is running, owner A, lease until T', renewed by the run while it is alive and released on completion. The next tick's winner checks that guard and applies the job's **overlap policy**: *skip* the tick and record it (typical for sweeps), *queue* one pending run to start when the current one ends, or *replace* by signalling the old run to stop, which is safe only if its late writes are fenced. The guard must be a lease rather than a plain flag, or one crashed owner blocks the job forever. Chronic overrun is also a signal to fix the job itself.

go deeper

for a junior

Recall that a job can still be running when its next scheduled time arrives, and that something must decide whether the next run waits, skips or starts anyway.

for a middle

Explain why a per-tick key cannot block a different tick, and how a job-level running lease with an owner and an expiry closes that gap.

for a senior

Pick the overlap policy per job, make replace safe with fencing, record skipped ticks, and alert when run duration approaches the interval.

for a principal

Treat chronic overrun as a capacity signal and decide whether to lengthen the interval, bound work per run, or move the job to a continuously running worker.

## Why per-tick claims do not prevent overlap In cluster-wide cron with per-tick claims, each node races for a record keyed by `(job, scheduled_tick)`, and the unique key guarantees one winner **per tick**. It says nothing about **different ticks**. Consider a minutely cleanup: 1. At 02:00 node A wins `(cleanup, 02:00)` and starts sweeping. 2. The sweep is slow today and is still running at 02:01. 3. At 02:01 node B wins `(cleanup, 02:01)`, a different key with no conflict, and starts a second sweep. Now two sweeps delete and update the same rows, lock each other, or one deletes data the other is still reading. On a single machine an in-process scheduler can see its own previous run; across nodes, B has no idea A is busy unless something shared tells it. The same gap appears during a catch-up after an outage, when replayed ticks arrive while a normal run is still going, and after a deploy, when a freshly started node has no memory of what the old nodes were doing. The arithmetic of an unguarded overrun is unforgiving: if each run takes three minutes and a new one starts every minute, about **three copies** run at once in steady state, each slowing the others, which lengthens runs further. ## A job-level running guard The fix is a second piece of shared state, keyed by the **job alone**: 1. Before doing any work, the tick's winner tries to acquire the guard `running(job)`, setting owner and a lease expiry. 2. If the guard is free or expired, it takes it and runs, renewing the lease periodically while alive. 3. If the guard is held and unexpired, it applies the job's overlap policy. 4. On completion, the owner releases the guard, conditioned on still being the owner. The per-tick claim answers "who runs this tick?" and the job-level guard answers "may anyone run right now?". Some designs fold both into one conditional write, but the two questions stay distinct. ## Overlap policies | Policy | Behaviour when previous run is live | Fits | Caveat | |---|---|---|---| | Skip | Drop this tick and record it as skipped | Sweeps and refreshes where the next run covers the gap | With chronic overrun, only about one tick in three runs for a 3-minute job | | Queue one | Remember one pending run and start it when the current one ends | Work that must follow every burst but can coalesce | Keep at most one pending, or the queue grows without bound | | Replace | Ask the old run to stop and start the new one | Runs where only the freshest result matters | Needs fencing so the old run's late writes are rejected | | Allow | Start anyway | Truly independent, partitioned work | Usually the accidental default, not a choice | ## Leases, renewal and fencing - A **plain flag** set by a node that then crashes is never cleared; every later tick sees "running" and skips, and the job stops forever with no error. The guard must carry an **expiry**. - The owner **renews** the lease while it is alive; the renewal interval should be well below the lease length so one slow renewal does not free the guard. - A lease can still lapse while its owner is stalled. If a new run then takes the guard, the old run's writes must be rejected using a **fencing token** that increases with each acquisition. - What happens to an interrupted run, whether it is resumed, retried or abandoned, is a recovery question separate from the overlap guard. ## Fixing the cause An overlap guard contains the damage; it does not explain why a minutely job takes three minutes. Operate it like any capacity signal: - **Measure** the ratio of run duration to interval and alert well before it reaches 1. - **Bound work per run**, for example deleting at most a fixed batch of rows, so a run always finishes inside its interval and backlog drains over several ticks. - **Lengthen the interval** if the work does not need minute-level freshness. - **Change the shape**: a job that is always busy may be better as a continuously running worker than as a cron schedule. - **Record skipped ticks** so a skip policy never hides a job that has effectively slowed to a fraction of its schedule.

  • Why must the job-level running guard be a lease rather than a plain flag?
    A plain 'running' flag set by a node that then crashes is never cleared, so every later tick sees the job as running and skips, and the job stops forever without an error. A lease with an expiry, renewed while the run is alive, releases itself after the owner disappears. How the interrupted run is then recovered is a separate concern.
  • A minutely cleanup now takes three minutes. What happens with no overlap guard, and with a skip policy?
    Without a guard, a new copy starts every minute while each lasts three, so about three copies run at once in steady state and slow each other further. With skip, ticks that arrive while a run is live are dropped, so only about one tick in three runs, and the job is effectively every three minutes, which a duration-to-interval alert should surface.

saying these in an interview costs you the question

  • The per-tick lock already stops the next tick from overlapping.
  • A plain running flag is fine because the job always clears it.
  • Replacing the old run is safe without fencing its writes.
  • Overlap cannot happen when every tick has its own unique key.
  • Skipped ticks from overlap need neither a record nor an alert.