skip to content

A recurring job stops running entirely, with no error in the logs and the process still healthy. What common property of periodic scheduling explains this, and how do you defend against it?

level: middleimportance: must knowfreq 58%

answer

  1. escaping exception cancels the whole recurrence, not just one run
  2. failure stored in an unread result handle — nothing logs it
  3. catch EVERYTHING inside the task body, return normally
  4. catch block itself must not throw
  5. alert on stale last-success timestamp, not process health

basics

~20 s

Most schedulers cancel a repeating task permanently if one run throws an exception that escapes it, and the failure is only recorded in a result handle nobody inspects. Defend by catching everything inside the task body and monitoring last-success time, not just process health.

solid answer

~60 s

Repeating schedules are typically specified as: run, and if the run completes normally, schedule the next occurrence. An escaping exception is therefore not a failed *run* — it is the **end of the schedule**. The task is removed and never runs again. It is silent because the exception is usually captured into the handle representing the scheduled task's outcome, and nothing ever reads that handle. Nothing is thrown on any live thread, so no default error handler fires and no stack trace reaches the log. Defenses, in order: 1. **Wrap the entire task body** so nothing escapes: catch every error type, log it with context, and return normally. This is the only fix that actually keeps the schedule alive. 2. **Log inside the catch** — including errors you consider impossible. 3. **Monitor liveness, not process health**: export a last-successful-run timestamp or a run counter, and alert when it goes stale. Process-up says nothing about schedule-alive. 4. **Bound each run with a timeout**; a hung run suspends the schedule just as effectively as a dead one.

code

text · 17 lines
text
# scheduler's internal loop
run(task)
if completed normally: reschedule(task, nextDue)
else: store exception in handle; STOP           # <- nobody reads the handle

# safe task body
periodicTask():
  try:
    doTheWork()
    metrics.lastSuccess = now()                 # liveness signal
  catch anything as e:
    try: log.error("job failed", e) ; metrics.failures++
    catch: pass                                 # catch block must not throw
  # returns normally -> next occurrence is scheduled

# alert rule
alert if now() - metrics.lastSuccess > 3 * period

go deeper

for a junior

Know the rule: if a repeating task throws and the exception escapes, the scheduler stops repeating it, and nothing logs the failure. Wrap the task body in a catch-all that logs and returns normally.

for a middle

Explain why it is silent — the exception is stored in a result handle nobody reads — and give the two-part defense: catch everything inside the body, and monitor a last-success timestamp.

for a senior

Add the cases the wrapper misses (hung runs with no timeout, cancelled tasks, shut-down pools, shared workers starved by one blocked job) and design the liveness alert threshold relative to the period.

for a principal

Treat it as an observability contract for all background work: every recurring job exports liveness and consecutive-failure counters, alerts on absence of success, and has a defined policy for transient versus permanent failure escalation.

## The rule that surprises everyone A periodic schedule is a loop the scheduler runs on your behalf, and its continuation condition is *the previous run finished normally*. Expressed as pseudocode: ``` run task if it completed normally: schedule the next occurrence else: record the failure and stop rescheduling ``` So a single escaping exception — on run number one or run number one million — terminates the entire recurrence. Not "skip this run." Not "retry next period." Terminate. The rationale is defensible: the scheduler has no idea whether the failure is transient or a permanent misconfiguration, and blindly repeating a task that throws instantly could produce an infinite error loop. So it stops and reports. The problem is *how* it reports. ## Why it is silent When you schedule a repeating task you typically get back a handle representing the task's eventual outcome. The escaping exception is stored **inside that handle**. It is not rethrown on any running thread, so: - No uncaught-exception handler fires, because the exception was caught by the scheduler. - No stack trace is printed, because nothing printed it. - The process stays perfectly healthy — this is one task among many. - Health checks pass, memory looks fine, CPU looks fine. The only way to see the failure is to inspect the handle, which requires a thread to ask for the result — and nobody does, because the task "runs forever" and there is no result to collect. The failure is therefore perfectly, structurally invisible. Teams commonly discover it days later from a downstream symptom: a cache never refreshed, a queue never drained, a report never generated, a certificate never renewed. ## The defenses ### 1. Never let anything escape the task body The only structural fix. Wrap the whole body in a catch that handles *every* throwable category the platform has, not just the ones you expect: ``` periodicTask(): try: doTheWork() catch anything as e: log.error("periodic task failed", e, context) metrics.failures++ # always return normally -> schedule survives ``` Two details people get wrong. First, catching only the "expected" error type leaves everything else escaping — and the errors that kill schedules in production are exactly the unexpected ones (a null dereference, a class-loading failure, an out-of-range index in code you did not think could fail). Second, the catch block itself must not throw; a logger that fails, or a metric emitter that throws, will escape and kill the schedule from inside the safety net. ### 2. Log with enough context to act on A schedule that survives failures but logs nothing is only a slower kind of invisible: the work is not happening and everything looks fine. Log the exception with the task identity, the attempt number, and any input that would let someone reproduce it. ### 3. Monitor the schedule, not the process This is the defense that catches everything, including the cases your wrapper misses (the task hung, the pool was shut down, the scheduler thread died, someone cancelled it). Emit either: - a **last-successful-run timestamp** per job, and alert when `now − last_success > k × period`, or - a **monotonically increasing success counter**, and alert when it stops increasing. The important property is that this is *liveness* monitoring: it detects absence of work. Almost all error monitoring detects presence of errors, which is exactly the wrong shape for this failure — there are no errors, only silence. Choose `k` to tolerate a couple of missed runs so ordinary jitter does not page anyone. ### 4. Bound each run A task that blocks forever — a network read with no timeout, a lock never released, a deadlock — never completes, so the next occurrence is never scheduled. The schedule is just as dead, and no exception was ever thrown, so the wrapper does not help. Give each run its own timeout or deadline so a hung attempt is abandoned and the recurrence continues. On a scheduler with a small worker count, one hung task can also starve *other* schedules that share those workers. ### 5. Distinguish transient from permanent Swallowing every error forever is not always right either. If a job fails identically for hours, something needs to escalate. A good pattern is to keep the schedule alive but track consecutive failures, log at increasing severity, and alert after a threshold — you get both liveness and honesty. ## The mental model to carry away There are three distinct states a recurring job can be in, and only the first is healthy: 1. **Running and succeeding** — the counter advances. 2. **Running and failing** — errors appear in logs; visible, fixable. 3. **Not running at all** — total silence, everything green. State 3 is the one this question is about, and the only way to detect it is to monitor for the absence of successful runs.

  • Your wrapper catches the platform's ordinary exception type, yet the schedule still died. What could have happened?
    Either an error outside that type escaped — many platforms distinguish ordinary exceptions from severe errors such as out-of-memory or linkage failures, and a catch of the narrower type will not stop those — or the catch block itself threw, for example because the logger or metrics call failed. It is also possible the task never returned at all: a blocking call with no timeout suspends the recurrence without any exception. Catch the broadest category, keep the handler trivially safe, and bound each run with a timeout.
  • What single monitor would you add to detect this class of failure across all your scheduled jobs?
    A per-job last-successful-run timestamp exported as a metric, with an alert when the age exceeds a small multiple of the job's period. It is liveness monitoring — it detects the absence of work rather than the presence of errors — so it catches every cause at once: swallowed exceptions, hung runs, cancelled tasks, a shut-down pool, or a scheduler thread that died. Error-rate alerts cannot detect it, because a dead schedule produces no errors.
  • If catching everything keeps the schedule alive, is it right to swallow failures indefinitely?
    No — that trades a silent dead schedule for a silently useless one. Keep the recurrence alive, but count consecutive failures, log each with context, escalate severity as the streak grows, and alert past a threshold. That way transient errors self-heal while a persistent misconfiguration still reaches a human.

A wind-up clock where the person rewinding it is told to stop forever the first time a spring slips. The clock does not sound an alarm; it just goes quiet, and everyone assumes the time on its face is current until a meeting is missed.

saying these in an interview costs you the question

  • Believing a failed run simply skips and the schedule continues on its own.
  • Assuming an uncaught-exception handler or the default error log will surface it — the scheduler already caught it.
  • Catching only the narrow expected exception type, letting severe errors escape and kill the schedule.
  • Relying on process health checks or CPU/memory dashboards to notice a dead schedule.
  • Forgetting that a hung run with no timeout stops the recurrence just as effectively as an exception.

context