skip to content

For cluster-wide cron, how does firing schedules only on an elected leader compare with every node racing for a per-tick lock?

level: middleimportance: must knowfreq 58%

answer

  1. one decider versus many racers
  2. what happens during handover
  3. unique key on job and tick
  4. claim attempts scale with node count
  5. stale holder, increasing token

basics

~20 s

An elected leader runs all schedules from one node: simple, but ticks can be missed or doubled around failover. A per-tick lock lets any live node win each tick with no handover, at the cost of every node attempting a claim every tick.

solid answer

~50 s

With **leader election**, the nodes agree on one leader that alone evaluates schedules while the rest stay idle. It is cheap per tick and keeps schedule state in one place, but correctness hinges on the handover: a tick that falls between the old leader dying and the new one taking over is missed unless the new leader checks what should have fired, and a leader that stalls past its lease can fire alongside its successor. With a **per-tick lock**, every node wakes at each tick and tries an atomic claim keyed by `(job, scheduled_tick)`, a unique insert or conditional write, and only the winner runs. There is no handover, so any surviving node keeps firing, but each tick costs one claim attempt per node and the shared store becomes the critical dependency. Both rely on leases, and long runs need a **fencing token** so a stale holder's writes are rejected.

go deeper

for a junior

Know the two names, an elected leader that runs schedules and a lock claimed per tick, and that both aim for exactly one run per scheduled time.

for a middle

Walk through each mechanism: who evaluates the schedule, how the atomic claim on job and tick works, and what happens to a tick while leadership changes hands.

for a senior

Bring up the leases both designs depend on, how a stalled holder can outlive its lease, and how fencing tokens and a takeover check for missed ticks close those gaps.

for a principal

Choose by job count, existing infrastructure and blast radius: a leader concentrates load and risk in one place, while per-tick claims spread cost across every node and the shared store.

## The problem both approaches solve A fleet of identical nodes needs a recurring job, such as a nightly billing run or a minutely cleanup, to fire **once per scheduled tick** no matter how many nodes exist. Leaving a scheduler active on every node fires the job N times; hard-coding one node fires it zero times when that node disappears. Two mechanisms inside the application solve this by moving the fire decision into shared, consistent state: **leader election** and a **per-tick lock**. ## Option 1: leader-driven scheduling The fleet uses a coordination service or a consensus-backed store to agree on one **leader**. How that agreement is reached is its own topic; what matters here is how scheduling uses it. 1. Every node runs the scheduler code, but only the node currently holding leadership evaluates cron expressions. 2. Leadership is held through a **lease**, a grant that expires unless renewed. 3. When the leader dies, its lease expires and another node acquires it. 4. The new leader starts its scheduler loop and fires future ticks. The strengths are clear: one node does the work, per-tick cost is negligible, and a leader can keep rich in-memory state such as a priority queue of thousands of next fire times. The weak point is **step 3**. If the lease takes 30 seconds to lapse and the old leader died at 01:59:50, the 02:00 tick passes while nobody leads. A new leader that only computes "next fire time after now" never runs it. And if the old leader was merely paused, not dead, it may resume and fire the same tick its successor already fired. ## Option 2: a lock per tick Here there is no leader. Every node computes the same schedule and, at each tick, races for a claim that only one can win. ```pseudocode every node, loop: tick = next_fire_time(job.cron, after = last_seen_tick) // computed in UTC sleep_until(tick) claim = try_claim(job.name, tick, lease = job.max_runtime) // atomic: unique (job, tick) if claim.won: run(job, tick, claim.fencing_token) // writes carry the token last_seen_tick = tick ``` The claim is typically an insert into a table whose primary key is `(job_name, scheduled_at)`, or a conditional set-if-absent in a key-value store. The first attempt succeeds; the rest see a conflict and skip. Because every node attempts every tick, losing any node, even the one that won last time, changes nothing: the next tick is won by whoever is alive. Note that the loop advances from `last_seen_tick`, not from the current time, so a tick delayed by a slow wake-up is still attempted rather than silently skipped. ## Side-by-side | Aspect | Elected leader | Lock per tick | |---|---|---| | Who evaluates schedules | One node | Every node | | Cost per tick | One evaluation | One claim attempt per node | | Losing a machine | Needs a leadership handover | No handover; others keep claiming | | Missed-tick risk | Gap during failover | Only if every node or the store is down | | Duplicate risk | Stale leader after a pause | Claim expiry while holder is paused | | Critical dependency | Coordination service | Shared claim store | | Good fit | Many jobs, rich schedule state | Few jobs, an existing shared database | The cost row is worth quantifying. A 200-node fleet with one minutely job makes 200 claim attempts a minute, about 288,000 a day, for 1,440 runs. That is trivial for most databases, but it grows with nodes times jobs. ## Leases and fencing tokens on a cron run Both designs depend on **leases**, because a claim or a leadership that never expires would be stuck forever after its holder dies. Leases create one hazard: a holder that stalls, for example in a long pause or a network partition, can outlive its lease without knowing it. - A **fencing token** is a number that increases with every grant of the lease or claim. - The holder attaches the token to every write it makes on behalf of the run. - The target of those writes remembers the highest token it has seen and rejects anything lower. - A stale holder that wakes up late therefore has its writes refused, while the newer holder proceeds. For a per-tick lock, this matters only if expired claims can be re-taken, which is how a design recovers a tick whose runner vanished. For a leader design it matters on every handover. ## Choosing - Prefer a **leader** when there are many schedules, the fleet already runs a coordination service, and you want one place that holds the whole schedule in memory. - Prefer a **per-tick lock** when there are a handful of jobs, a shared relational database already exists, and you want to avoid reasoning about handover gaps. - Many production schedulers **combine** them: a leader (or a small set of scheduler nodes) decides what is due, and a per-tick claim guards against the leader's own duplicates. - In either design, a new leader or restarted node should compare the last recorded tick with the schedule, so a failover gap becomes an explicit missed-fire decision rather than a silent loss.

  • With leader-driven cron, how can a tick be missed during failover, and how do you close the gap?
    If the leader dies at 01:59:50 and its lease takes 30 seconds to lapse, the 02:00 tick passes with no leader, and a successor that only looks forward from now never runs it. Close the gap by having every new leader compare each job's last recorded tick with its schedule on takeover and apply that job's missed-fire policy.
  • Why does a per-tick lock need a fencing token if the tick key is already unique?
    The unique key stops two nodes from claiming the tick at the same time. But if an expired claim can be re-taken, the original holder may wake from a stall and keep writing after the successor took over. A token that increases with each claim, checked by the store receiving the writes, rejects the stale holder's late writes.
  • What does a per-tick lock cost on a 200-node fleet running one minutely job?
    Every node attempts a claim every minute: 200 attempts a minute, about 288,000 a day, for 1,440 actual runs. That is usually trivial, but it scales with nodes times jobs, so fleets with thousands of schedules usually let only a leader or a few scheduler nodes make claims.

saying these in an interview costs you the question

  • Leader election guarantees no tick is ever missed or doubled.
  • A unique per-tick key alone stops a stale holder from writing.
  • A per-tick lock still needs a leader to decide who races.
  • Per-tick claims are free, so they scale to any fleet and job count.
  • The lock can be held without expiry because the job always finishes.