skip to content

A recurring job is configured to run every 30 seconds, but under load individual runs start taking 90 seconds. Describe what a typical scheduler does in that situation and how you would make the job behave sensibly.

level: seniorimportance: should knowfreq 45%

answer

  1. schedulers serialize periodic tasks — no overlap by default
  2. E > P ⇒ back-to-back runs, zero idle, cadence silently lost
  3. positive feedback: job hammers the resource that slowed it
  4. fix: measure duration, timeout each run, gap-based interval
  5. explicit missed-run policy + shared lease for multi-node

basics

~20 s

Schedulers serialize a periodic task rather than overlapping it, so missed occurrences pile up and each run starts immediately after the last — effectively continuous execution with no idle gap. Fix by measuring duration, skipping missed occurrences, bounding each run, and switching to a gap-based schedule.

solid answer

~60 s

Standard periodic schedulers do **not** run two instances of the same recurring task at once; they serialize. So when execution exceeds the period, the next occurrence is already overdue when the current one ends and starts immediately. The effective period becomes the execution time, the idle gap drops to zero, and the job is now hammering whatever resource made it slow in the first place — a positive feedback loop. If the scheduler instead *does* allow overlap (some do, and application-level schedulers vary), you get concurrent instances of a job that was almost certainly written assuming exclusivity: double processing, lock contention, and unbounded thread growth as instances accumulate. Fixes, layered: (1) measure and alert when duration approaches the period; (2) give each run a timeout so one hung attempt cannot suspend the schedule; (3) switch to an interval measured from the previous run's *end*, so slowness stretches the period instead of collapsing the gap; (4) make the job idempotent and guard it with a "skip if already running" flag; (5) if the work simply does not fit, chunk it or shard it.

code

text · 14 lines
text
P=30s, E=90s, interval measured from previous START:
t=0   [========= run 1 (90s) =========]
t=30  due (blocked)  t=60 due (blocked)  t=90 due
t=90  [========= run 2 (90s) =========]   <- no gap, ever
effective period = 90s, idle time = 0

Same job, interval measured from previous END, plus guards:
periodicJob():
  if !lease.tryAcquire(ttl=2*P): metrics.skipped++; return   # overlap guard
  try:
    withTimeout(maxRun): doBoundedBatch()                    # bound the run
  finally:
    lease.release()
=> period becomes E + P; target always gets P of quiet

go deeper

for a junior

Say that the scheduler will not run two copies at once, so the next run starts right after the previous one finishes and the job effectively runs all the time instead of every 30 seconds.

for a middle

Add why the cadence is lost silently, the difference between skipping and catching up on missed occurrences, and that switching to an interval measured from the previous finish restores a guaranteed gap.

for a senior

Discuss the feedback loop against the slow dependency, per-run timeouts, duration monitoring as the leading indicator, and an explicit skip-if-running guard with idempotent job design.

for a principal

Treat it as capacity design for background work: does the work fit its window at all, should it be chunked or sharded, what is the missed-run policy per job class, and how are leases and idempotence guaranteed across a multi-node deployment.

## The situation Period P = 30 s, execution E = 90 s. The schedule asks for something arithmetically impossible: three runs' worth of due times pass during every single run. ## What schedulers actually do **Serialization is the norm.** Nearly every general-purpose scheduler treats a periodic task as a single logical job and will not start occurrence *n+1* while occurrence *n* is running. What you get is: - occurrence 2 becomes due at t=30 while run 1 is still going, - occurrence 3 due at t=60, occurrence 4 at t=90 — all overdue, - run 1 finishes at t=90, and the next run starts *immediately*. The steady state is back-to-back execution at the execution time: a run every 90 s with zero idle. Two harms follow. First, the configured cadence is silently gone; anything downstream expecting a 30-second refresh is now three times stale, with no error anywhere. Second — the dangerous one — **the job's pressure on the system goes up exactly when the system is struggling.** Normally the job worked for 30 s out of every 90 s at most; now it works 100% of the time against the same slow database or API, which makes it slower still. That is a positive feedback loop, and it is how a harmless maintenance job becomes the proximate cause of an outage. A second question is what happens to the *missed* occurrences. Some schedulers coalesce them (one immediate run, then re-anchor); others attempt to catch up by replaying each. Catch-up replay is worse: it fires a burst of runs at a system that just fell behind. **Overlapping is possible and worse.** Some application-level and distributed schedulers will start a new instance regardless. Then you have concurrent copies of a job usually written assuming it is alone: rows processed twice, lock contention or deadlock between instances, memory growing with the number of live instances, and — because each instance is slow — an ever-growing set of them until resources run out. ## Making it behave ### 1. Measure the run, not just the schedule Export duration percentiles per job and alert when p99 duration exceeds, say, half the period. The transition from "comfortable" to "continuous execution" is gradual in reality and invisible without this metric. Duration approaching the period is the single best leading indicator. ### 2. Bound each run Give every run a deadline or timeout. Without one, a run that blocks forever suspends the schedule permanently — indistinguishable, from the outside, from a cancelled schedule. With one, a pathological attempt is abandoned and the cadence resumes. ### 3. Prefer a gap-based interval when duration is variable An interval measured from the previous run's *end* makes the period stretch to E + P automatically. The target always gets P of quiet, no matter how slow the job becomes. This is self-limiting by construction and is the right default for anything polling or scanning a shared resource. You give up cadence guarantees, which is the correct trade when the cadence was already unachievable. ### 4. Decide the missed-run policy explicitly For most maintenance work, **skip** is right: if three refresh windows passed, run once now and re-anchor to the next due time — there is no value in doing the same refresh three times. **Catch-up** is only right when each occurrence corresponds to distinct work that must not be lost (per-hour billing aggregation, for example), and even then it should be rate-limited so recovery does not become a thundering burst. Make this a stated policy per job, not an accident of the library's default. ### 5. Guard against overlap explicitly Do not rely on the scheduler's serialization if you might ever move to a different scheduler or run more than one instance of the service. Use a guard the job itself owns: a flag or lease checked at entry that causes an immediate no-op return (and increments a "skipped because still running" counter) if a previous run is in flight. In a multi-node deployment this must be a shared lease with an expiry, not an in-process flag — otherwise every node runs its own copy. ### 6. Make the job idempotent If running twice concurrently or replaying a missed occurrence is harmless, every other problem here becomes survivable. Idempotence is the property that turns a scheduling correctness problem into a scheduling efficiency problem. ### 7. Fix the size mismatch All of the above manages symptoms. If a 30-second job genuinely takes 90 seconds, the work does not fit the window. Chunk it (process a bounded batch per run and continue next time), shard it across instances by key range, or lengthen the period to match reality. Choosing a period without measuring the work is the original defect. ## The summary an interviewer wants Overrun does not produce an error — it produces silent loss of cadence, and, with fixed-rate semantics, a job that consumes resources continuously at the worst possible time. The engineering answer is: measure duration, bound each run, prefer gap-based intervals for variable work, define the missed-run policy deliberately, guard overlap with a lease, and ultimately make the work fit the window.

  • Why is continuous execution after overrun more dangerous than simply missing the cadence?
    Because it is a positive feedback loop. The job was previously idle most of the time; once runs exceed the period it consumes the shared resource — database, API, disk — without pause, which makes that resource slower, which makes runs longer, which removes any remaining gap. A background job that was harmless at 30-second intervals can become the dominant load on a struggling dependency precisely during an incident.
  • If missed occurrences pile up, should the scheduler catch up on all of them?
    Usually no. For idempotent maintenance work such as a cache refresh or a cleanup sweep, running once and re-anchoring to the next due time achieves the same end state, while replaying every missed occurrence adds a burst of load to a system that just recovered. Catch-up is justified only when each occurrence maps to distinct work that would otherwise be lost, such as per-interval aggregation, and even then it should be rate-limited.
  • Your service now runs on three nodes and each node runs the same scheduled job. How do you prevent three concurrent executions?
    An in-process guard is not enough, because each node has its own memory. You need a shared lease or lock in a store all nodes can see, acquired at job entry with an expiry longer than the expected run time so a crashed holder does not block forever, and released on completion. Nodes that fail to acquire it return immediately and increment a skipped counter. Making the job idempotent remains the durable safety net, since leases can expire mid-run.

A cleaner scheduled to sweep a hallway every 30 minutes, where sweeping now takes 90. Nobody is cloned to sweep in parallel; the cleaner simply never stops, permanently in the way of everyone using the corridor — and the hallway is not cleaner for it.

saying these in an interview costs you the question

  • Assuming the scheduler starts a second concurrent instance when a run overruns — the usual behavior is serialization.
  • Believing overrun produces an error or warning; it is silent and only visible as lost cadence.
  • Missing the feedback loop where a slowed job consumes its dependency continuously.
  • Relying on the scheduler's serialization for overlap safety in a multi-node deployment, where each node runs its own copy.
  • Blindly replaying every missed occurrence after a stall, adding a burst to a recovering system.

context