skip to content

In a media-render job platform, a job keeps crashing its worker before any failure is reported; how do you stop it cycling forever?

level: seniorimportance: nice to knowfreq 30%

answer

  1. a crash never reports its failure
  2. where is the attempt counter bumped?
  3. limit enforced when claiming
  4. exponential delay, randomised
  5. quarantine rather than delete

basics

~20 s

Count an attempt when the job is claimed, not when a failure is reported, and check the limit at claim time. A job that uses up its attempts by crashing workers moves to a terminal quarantined state for human review instead of being reclaimed again.

solid answer

~50 s

A **poison job** is one whose input reliably breaks the worker, for example a malformed file that makes the renderer exhaust memory and get killed. Because the process dies, no error handler runs, so a retry limit that counts *reported failures* never triggers: the lease expires, the job is reclaimed, and it kills the next worker. The fix is to increment `attempt` in the claim itself and refuse to claim once `attempt >= max_attempts`, moving the job to `dead` with its run history instead. Between reclaims, apply exponential backoff with jitter so a crash loop is slow and spread out rather than instant. Exempt graceful releases, such as a worker drained for a deploy, so innocent jobs don't burn attempts. Alert on correlated worker deaths that share a job id, and once quarantined, hand the job to a failure queue for a human to inspect and replay.

go deeper

for a junior

Remember that a job can kill its worker so the error is never recorded, and that attempt limits still need to stop it.

for a middle

Explain why the attempt counter must be incremented at claim time, and how exponential backoff with jitter spaces out reclaims.

for a senior

Show how you tell poison jobs from infrastructure churn: release outcomes, correlated worker deaths, isolation of suspects, and quarantine with an alert.

for a principal

Decide attempt budgets and isolation policy per job type, weighing the fleet damage of a crash loop against wrongly killing jobs during unstable periods.

## What makes a poison job different In a background-job platform that renders media, most failures are **reported** failures: the worker catches an error, records it, and the job moves to `retry_wait`. A **poison job** is worse. Something in its input, such as a corrupt file, an absurd resolution or a pathological edit list, makes the worker process itself die: it runs out of memory and is killed, or it triggers a crash in a native library. No error handler runs, so nothing is recorded. From the platform's view, the job just goes silent. Its lease expires, the job is reclaimed, and the next worker dies the same way. As an illustration, if the crash happens a minute in and the lease is 60 seconds, the job is reclaimed roughly every two minutes, about 30 worker deaths an hour, each taking down any other jobs that worker was running. ## Count attempts at claim time The core fix is **where** the attempt counter is incremented: - **At failure time** (in the error handler): a crash never reaches the handler, so the counter never moves and the loop is infinite. - **At claim time** (inside the conditional update that grants the lease): every execution is counted, whether it ends in success, a reported error or a crash. The claim then refuses to hand out a job that has already used its budget: ```pseudocode on reclaim(job): // lease expired, no outcome recorded recordRun(job, job.attempt, LEASE_EXPIRED) if job.attempt >= job.maxAttempts: setState(job, DEAD, reason = "crashed workers on every attempt") alert(job) else: delay = random(0, min(3600, 30 * 2^(job.attempt - 1))) setState(job, RETRY_WAIT, notBefore = now + delay) ``` ## Backoff with jitter between reclaims Even below the limit, reclaiming instantly makes a crash loop as fast as possible. **Exponential backoff** spaces attempts out: with an illustrative base of 30 seconds and a one-hour cap, the delay after the nth failure is `min(3600, 30 x 2^(n-1))`, so 30, 60, 120 and 240 seconds after failures one to four, 450 seconds (7.5 minutes) in total before a fifth attempt. **Jitter** randomises each delay, for example `random(0, delay)`. When one bad upstream change poisons thousands of jobs at once, jitter stops them from all retrying in the same second and knocking over the fleet in synchronised waves. ## Detecting and quarantining 1. **Enforce the limit at claim time**, as above, so the loop is bounded no matter how the worker died. 2. **Correlate worker deaths.** If the last jobs held by several crashed workers share one job id, that job is the likely culprit, even before its limit is reached. 3. **Isolate suspects.** A job with repeated unexplained deaths can be routed to a dedicated, resource-limited worker that runs only one job at a time, so it cannot take innocent jobs down with it. 4. **Quarantine, do not delete.** Move it to the terminal `dead` state with its full run history, alert, and hand it to the platform's failure queue so a human can inspect the input, fix the cause and replay it deliberately. ## Not punishing innocent jobs Claim-time counting charges an attempt for every execution, including ones that failed for reasons unrelated to the job. The run outcome decides what counts: | How the attempt ended | Consume an attempt? | |---|---| | Reported error | yes | | Lease expired with no outcome | yes, the job is a crash suspect | | Graceful release during a deploy or scale-down | no, record it as released | | Whole-pool outage affecting many jobs | usually no; operators may reset attempts after the incident | Without the release exemption, a job that happens to be running during several deploys could be declared dead having done nothing wrong. It also helps to set `max_attempts` with some headroom for ordinary infrastructure churn. ## What to surface - A per-job-type count of `lease_expired` outcomes, which rises sharply when a poison input appears. - Worker exit reasons, such as memory kills versus clean exits, joined with the job ids the worker held. - Jobs in `dead`, with their run history, as an actionable list rather than a silent pile. - A replay action that resets the attempt counter explicitly and records who replayed the job and why.

  • Why add jitter to the backoff delay between attempts?
    When a shared cause, such as a bad upstream release, breaks many jobs at once, plain exponential backoff makes them all retry at the same moments, producing synchronised load spikes that can cause fresh failures. Randomising each delay spreads the retries across the interval, so the fleet sees a smooth trickle instead of waves.
  • How do you avoid charging an attempt when a worker is drained for a deploy?
    On graceful shutdown the worker releases its leases explicitly and records the run outcome as `released`. The platform then either decrements the attempt or counts only runs whose outcome is an error or an expired lease. Only silent disappearances and reported errors should consume the job's budget.
  • What should happen after a poison job is quarantined?
    It stays in a terminal state with its full run history and an alert fires. The job is handed to the platform's failure queue, where someone inspects the input, fixes the renderer or rejects the input, and replays the job deliberately. The replay resets the attempt counter on purpose and is recorded in the run history.

saying these in an interview costs you the question

  • Increment the attempt counter in the worker's error handler.
  • A job that never reports a failure has not failed yet.
  • Reclaim immediately after a lease expires; the crash was probably a fluke.
  • Delete poison jobs automatically to protect the fleet.
  • Every worker death should consume one of the job's attempts.