skip to content

A long-running job holds per-key running totals. Why does deploying new code against it need a plan a stateless service does not?

level: juniorimportance: must knowfreq 60%

answer

  1. a service keeps nothing; this job remembers
  2. the new build inherits bytes
  3. two matches: which step, which format
  4. failure is refusal, or a silent empty start

basics

~20 s

New code has to read back the retained set the previous build wrote — everything the job holds between records. A stateless redeploy discards nothing; here, entries the new build cannot match or decode mean a refused start, or history silently lost.

solid answer

~50 s

A stateless service can be replaced process by process because nothing of value lives inside it. A long-running job's **retained set** — the running accumulators, records buffered awaiting a group, deduplication entries, join sides and pending per-key callbacks it holds between records — was written by the build you are replacing, so the new build has to find it and decode it. Two independent matches have to succeed: the runtime must decide which stored entries belong to which stateful step of the new program, and the new code must be able to read the stored values in the format the old code wrote them. Runtimes differ in what happens when either fails — some refuse to start and name the step, some start that step with an empty retained set and resume from nothing. A batch model that keeps nothing between runs has no such problem: you just run the new code.

go deeper

for a junior

Recall that a job holding running totals, deduplication entries or join sides carries data the next build must read back, and that a stateless service carries none. Naming that difference is the whole of the junior answer.

for a middle

Explain the two matches that must succeed — which stateful step the stored entries belong to, and whether the new code can decode them — and say that a failure may be a refused start on one runtime and a silent empty start on another.

for a senior

Show that you verify a history-dependent number after the deployment rather than trusting a green job, and that you know the fallback's cost grows with the age of the retained set and depends on the source still holding that period.

for a principal

The call you own is the standing convention: what every long-lived job in the organisation must decide before its first deployment about step identity, stored formats and expiry, so that no team discovers in week three that its only route is to start again.

## What the retained set is, and why it is not a cache A long-running job's **retained set** is everything it is still holding between records: running accumulators, records buffered awaiting their group, deduplication entries, both sides of a join, and pending per-key callbacks. It is the job's real storage system, and it is not a cache. Drop a cache and the next answer is slower; drop the retained set and the next answer is **wrong** — a running total restarts at zero, a deduplication entry no longer suppresses the duplicate that arrives an hour later, one half of a join never finds its partner. That is the entire reason replacing the code of such a job is a different operation from replacing a stateless service. A stateless service keeps nothing worth keeping in the process, so a new build can be started, given traffic and the old build killed; nothing is read back because nothing was ever written. A stateful job's new build inherits bytes written by the build it replaces. ## The two matches a redeploy has to make | Match | What is compared | What failure looks like | |---|---|---| | **Identity** | each stateful step of the new program against the label the stored entries carry | entries belong to no step the new program has, or a step starts and finds none of its own | | **Format** | the new code's idea of a stored value against the bytes the previous build actually wrote | a decode error, or — worse — a successful decode into a value that is quietly not what was stored | Both have to succeed, and they fail independently. Code that changed nothing about what is stored can still lose its entries because the step it belongs to is no longer identified the same way; a step that is identified perfectly can still be handed bytes it cannot read. ## What actually happens on a failed match — this varies This is where engineers who have only operated one runtime get caught, because the behaviours are genuinely different: - some runtimes **refuse to start**, naming the step whose entries could not be matched; - some **start that step with an empty retained set**, which is the dangerous one: the job is green, the dashboards are live, and the numbers are quietly missing everything accumulated before the deployment; - some start and **fail later**, at the first read of an entry the new code cannot decode, which can be hours after the deployment looked successful; - some require an explicit acknowledgement before they will proceed with entries they could not place. The operational rule that follows is the same on all of them: after a redeploy, check a number that depends on accumulated history, not just that the job is running. ## Where the problem does not exist at all 1. **The two-phase disk-to-disk batch model** keeps nothing between runs — every run re-derives its result from the input — so a new build is simply the next run's code. Long-lived state questions do not apply to it. 2. **Finite jobs generally**, where the whole answer is computed from an input that is read again each time. 3. **Stateless steps** inside an otherwise stateful continuous job: a filter, a projection, or an enrichment that looks a value up outside the job holds nothing between records, so it can be changed freely. Saying which of these you are in is the first move, not a detail: a great deal of anxiety about deployments is spent on jobs that hold nothing. ## Why the plan is a first-deployment decision The identity a stored entry carries is written **with the entry**, by whatever rule was in force at the time. That makes two things irreversible in practice: - an explicit name pinned to a stateful step after the fact labels the step, not the entries already written under the old label, so the new build looks for a name that nothing on storage carries; - the older the job, the more expensive the fallback, because rebuilding the retained set from the source is only possible for as long as the source still holds the period the entries cover, and it costs a catch-up run over all of it. So the compatibility plan — how a stateful step will be identified, and what happens to a stored value when its shape changes — belongs in the first deployment of a job that is expected to live, alongside the expiry rule that keeps the retained set bounded. It is cheap to decide on day one and can be impossible to decide in week three.

  • The job had been running for six weeks before the redeploy. Does the age of the retained set change anything?
    It changes the fallback, not the mechanism. The older the job, the more history the entries represent, so rebuilding them from the source costs a catch-up run over that whole period and is possible only while the source still holds it. Age also makes a silent empty start harder to notice, because the wrong numbers look plausible for a while.
  • Does a nightly job that re-reads its whole input and rewrites its output have this problem?
    No. It keeps nothing between runs, so the new build simply runs. The distinction is not batch against continuous, it is whether anything is carried from one run or one record to the next: a finite job that persists nothing is free to change, and a continuous job of only stateless steps very nearly is.
  • Is it enough that the new build compiles against the same value class the old one used?
    No. What is on storage is bytes, not a class. The code that turns a stored value into bytes and back can change — a different encoder, or a different configuration of one — while the class in the source looks identical, and then the format match fails even though the code matches.

Replacing a stateless service is swapping a vending machine for a newer one. Replacing a stateful job is swapping the machine while keeping the old one's coin box — the new machine has to open that box and count what is in it, and if the lock does not fit you either stop or start the day's takings at zero.

saying these in an interview costs you the question

  • Calls a stateful job's redeploy just a rolling restart
  • Treats the retained set as a cache that can be dropped without changing results
  • Assumes every runtime refuses to start rather than resuming with empty state
  • Assumes rebuild-from-source is always available as a fallback
  • Says the deployment tooling handles it, naming nothing that reads the stored values
  • Checks only that the job is running after the deployment, never a history-dependent number