skip to content

questions

5

In a background-job platform, what is a job state machine, and why does it keep a separate run record for each attempt?

level: juniorimportance: must knowfreq 58%

answer

  1. finite states, legal moves only
  2. terminal states are absorbing
  3. transition as a conditional update
  4. hot row versus append-only history
  5. where does the attempt counter live?

basics

~20 s

A job state machine is the fixed set of job states (queued, running, retry-wait, succeeded, dead) and the only transitions allowed between them. Per-attempt run records keep the job row small while preserving who ran each attempt and why it failed.

solid answer

~50 s

A job state machine is the explicit list of states a job can be in, typically `queued`, `running` (leased to one worker), `retry_wait`, `succeeded`, `cancelled` and a terminal `dead`, plus the only transitions allowed between them. Every transition is a guarded write, for example 'move to `succeeded` only if the job is `running` under attempt 3', so two actors cannot both apply conflicting moves. The job row holds only the *current* state and an attempt counter. Each attempt gets its own append-only **run record** with worker id, start and end times, outcome and error. That split keeps the row the dispatcher queries small and fast, while the run history answers what actually happened: how many times the job ran, on which workers, and why each attempt failed. Retry limits, poison-job detection and support debugging all depend on that history.

go deeper

for a junior

Be able to list the usual job states, say which ones are terminal, and explain why each attempt gets its own record instead of overwriting the job row.

for a middle

Explain that every transition is a conditional update naming the expected state and attempt, and walk through how that resolves a late worker racing a reclaim or a cancel.

for a senior

Show how the run history feeds retry limits, poison-job detection and support investigations, and how you keep the dispatcher's hot row narrow as history grows.

for a principal

Discuss retention and archiving of run history at high job volume, and which states and outcomes the organisation standardises across job types so tooling works for all of them.

## What a job state machine is A **background-job platform** runs work outside the request path: a tenant asks for a large data export or a media render, the API records a **job**, and a pool of **workers** executes it later. Many actors touch the same job concurrently: the API that creates or cancels it, the dispatcher that hands it out, the worker that runs it, a reaper that recovers abandoned work, and operators who replay failures. A **job state machine** is the agreement that keeps those actors consistent. It is a finite set of **states** plus the list of **legal transitions** between them. Anything not on the list is rejected. Without it, each component invents its own meaning of 'running' or 'failed', and races turn into jobs that are both finished and still retrying. ## Typical states and transitions | State | Meaning | Can move to | |---|---|---| | `queued` | waiting for a worker | `running`, `cancelled` | | `running` | leased to exactly one worker until the lease expires | `succeeded`, `retry_wait`, `dead`, `running` again under a new attempt if the lease lapses | | `retry_wait` | an attempt failed; waiting until `not_before` | `queued` | | `succeeded` | terminal: result stored | nothing automatic | | `dead` | terminal: attempts exhausted or a permanent error | `queued` only through an explicit operator replay | | `cancelled` | terminal: the tenant withdrew it | nothing automatic | Two properties matter most: - **Terminal states are absorbing.** No automatic path leaves `succeeded`, `dead` or `cancelled`. A late report from an old worker must not reopen them. - **Recovery is a transition, not a side channel.** Reclaiming a job whose worker vanished is modelled as a legal move that bumps the attempt number, so it is visible and countable. ## Transitions are guarded writes A transition is implemented as a **conditional update** (compare-and-set): the write names the state and attempt it expects to find, and changes nothing if they no longer match. ```sql UPDATE jobs SET state = 'succeeded', finished_at = CURRENT_TIMESTAMP WHERE job_id = :job_id AND state = 'running' AND attempt = :my_attempt; -- 0 rows changed: another actor already moved the job, so this worker stops ``` The guard resolves the everyday races: 1. A worker finishes late while the reaper has already handed the job to a new attempt. 2. A tenant cancels while the job is running. 3. Two dispatcher instances try to claim the same queued job. In each case exactly one conditional write wins and the loser learns it from the affected-row count. ## Job record versus run history The **job record** is the hot row the dispatcher scans, so it stays narrow: - job id, tenant id, job type and parameters - current `state`, current `attempt`, `max_attempts` - lease owner and lease expiry while running - `not_before` for delayed retries The **run history** is append-only, one row per attempt: - attempt number and worker id - claimed, started and finished timestamps - outcome: succeeded, failed, lease expired, released, cancelled - error class and message ```json { "jobId": "exp-8812", "state": "retry_wait", "attempt": 2, "maxAttempts": 5, "runs": [ {"attempt": 1, "worker": "w-14", "outcome": "lease_expired"}, {"attempt": 2, "worker": "w-03", "outcome": "failed", "error": "upstream timeout"} ] } ``` ## What the run history buys you - **Retry limits** compare the attempt counter with `max_attempts`; the history explains each attempt that was used. - **Poison-job detection** spots a job whose attempts all ended in `lease_expired` because it keeps killing workers. - **Debugging and support** can answer 'why did my export take four hours?' with queue wait versus run time per attempt. - **Metrics and billing** can separate time spent waiting from time spent executing. - **Audit** keeps a trail after the job row is compacted or archived. ## Common mistakes - Tracking a job with a single `done` boolean, which cannot express retrying, cancelled or reclaimed work. - Overwriting `last_error` on each attempt, which destroys the evidence of the earlier failures. - Letting any component write any state unconditionally, so a stale worker can flip a cancelled job back to succeeded. - Deriving the attempt count by counting run rows instead of incrementing it inside the claim itself, which can race when two claims happen close together.

  • Should the attempt counter live on the job row, or be derived by counting run records?
    On the job row, incremented inside the same conditional update that claims the job. Then claiming and numbering the attempt happen in one step, so two near-simultaneous claims cannot both believe they are attempt 3. Run records are then written with that number. Counting run rows afterwards is fine for reporting but races if used to decide the next attempt number or whether the limit is reached.
  • How should cancellation work for a job that is already running?
    The API cannot safely stop a remote worker mid-side-effect, so it sets a cancel-requested flag or moves the job to a `cancelling` state. The worker sees the flag on its next heartbeat or checkpoint, stops at a safe point, and the final move to `cancelled` is a conditional write. If the worker finishes first, its guarded `succeeded` write and the cancel race cleanly: whichever lands first wins.

A parcel tracking system shows one current status on the parcel, but keeps a separate scan log of every depot and delivery attempt; the status alone could not tell you why the parcel was late.

saying these in an interview costs you the question

  • A single done flag is enough to track a background job.
  • Any component may set any state as long as it knows the job id.
  • Overwriting the last-error field on each retry keeps enough history.
  • A late worker report may reopen a job that already reached a terminal state.
  • Run history is optional logging that retry logic never needs.
open as a page

In a job scheduler that hands out time-limited worker leases, how does it detect a worker that died mid-job and reclaim that job?

level: middleimportance: must knowfreq 70%

basics

~20 s

Each claim records a lease owner and expiry, and the worker extends the expiry with periodic heartbeats. When heartbeats stop, the lease lapses, and a reaper or the claim query itself treats the job as claimable and starts a new attempt.

open as a page

In a background-job platform, a worker stalls past its lease and the job is reassigned; how do you stop the stalled worker committing a stale result?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Give every claim a monotonically increasing attempt number, a fencing token, and make every commit conditional on it still being current. The stalled worker's write matches zero rows, so it aborts; external side effects must check the token too or be idempotent.

open as a page

In a multi-tenant job platform, one tenant enqueues two million exports at once; how would you keep that backlog from starving every other tenant?

level: principalimportance: should knowfreq 45%

basics

~20 s

Stop dispatching from one global FIFO. Queue work per tenant, choose the next tenant by weighted fair share of worker time, cap each tenant's concurrent jobs, and let those caps flex when others are idle so capacity is not wasted.

open as a page

In a media-render job platform, a job keeps crashing its worker before any failure is reported; how do you stop it cycling forever?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Count an attempt when the job is claimed, not when a failure is reported, and check the limit at claim time. A job that uses up its attempts by crashing workers moves to a terminal quarantined state for human review instead of being reclaimed again.

open as a page