skip to content

When a job is redeployed, what matches each stateful step to the entries the previous build wrote for it?

level: middleimportance: must knowfreq 55%

answer

  1. matched by a label, not by source code
  2. pinned name, or derived from position
  3. an unrelated edit can shift the derived one
  4. the label is written with the entries

basics

~20 s

An identity label written with the entries. Either the author pinned a stable step identifier — a name fixed to that stateful step — or the runtime derived one from the step's position in the job's step graph, which later edits can shift.

solid answer

~50 s

Stored entries carry a label saying which stateful step wrote them, and the new build is matched against that label. There are two ways the label is produced, and which one you get depends on the runtime. Some let the author pin a **stable step identifier** — a name fixed to the step so that a redeploy can match the new step to the state the old one wrote. Where no name is pinned, or where the model offers none, identity is **derived** — typically from the step's position among the stateful steps of the job's step graph, the graph a program becomes before it runs. Derived identity is the trap: inserting, removing or reordering a step can shift it, so an edit that has nothing to do with state loses the state. And the label is written with the entries, so pinning a name afterwards labels the step, not the millions of entries already on storage.

go deeper

for a junior

Know that stored entries carry a label saying which step wrote them, and that a redeploy matches the new program's steps against those labels rather than against the old code.

for a middle

Explain both ways the label is produced — pinned by the author, or derived from the step's position in the job's step graph — and give one edit that shifts a derived identity without touching the stateful step.

for a senior

Demonstrate that you pin identifiers in the first deployment and know why pinning them in the tenth is itself a state-losing change, plus the two routes out when it has already happened.

for a principal

Own the convention across teams: naming rules that survive a rewrite, a rule that every long-lived job pins on day one, and a register of the stateful steps whose entries could never be rebuilt from source.

## The label, not the code, is what is matched When a job with long-lived state is redeployed, nothing compares your old source to your new source. What exists on storage is a set of entries, each carrying a label that says which stateful step of the program wrote it. Starting the new build is a lookup: for each stateful step in the new program, find the entries filed under that step's identity. So the only question that matters is **how that identity is produced**, and the answer is not the same across this product class. ## Two ways identity is produced | Source of identity | How it is set | What breaks it | |---|---|---| | **Pinned** — a stable step identifier | the author fixes a name to the stateful step, in code, before the job first runs | changing the name later; nothing else | | **Derived** | the runtime computes it, typically from the step's position among the stateful steps of the job's step graph, sometimes also from the step's shape | inserting, removing or reordering steps; changing the stateful part of the program's shape | A runtime whose model is **record-at-a-time with key-bound state** — state a step may read or write only under the grouping key of the record it is handling — usually offers the pinned form and falls back to the derived one when the author did not use it. A runtime that runs continuous work as **repeated small finite jobs** — cutting an endless input into short chunks and running a complete job over each — commonly identifies its stateful steps by their position inside the plan compiled for each cycle, and may give the author nothing to pin at all; there the rule is stated the other way round, as *do not change the shape of the stateful part of the query*. And the **two-phase disk-to-disk batch model** keeps nothing between runs, so it has no labels and no problem. A candidate who states one of these as the universal mechanism has told you which runtime they have used. ## Why derived identity is a trap Derived identity couples state survival to edits that have nothing to do with state: - adding a step in front of a stateful one can move it in the ordering the identity is computed from; - deleting a step that was never stateful can do the same; - reordering two stateful steps can swap their labels, which is worse than losing them — each finds entries, and they are the wrong ones; - changing what the stateful step itself is, even to something equivalent, can change a shape the identity is computed from. Which of these actually shifts an identity on a given runtime is not reliably predictable from the source, and it is not a property you want to be learning during a production deployment. ## The part that is genuinely irreversible The label is written **with the entries**. That has one consequence worth stating slowly, because it is the reason this subject appears in senior interviews: 1. entries written under a derived identity carry that derived identity; 2. pinning a name to the step in a later release changes what the **new** step asks for; 3. the entries on storage are still filed under the old derived label, so the new step asks for a name nothing carries and finds nothing; 4. the pinning that was supposed to protect you is therefore the change that loses the state. The routes out of step 4 are the same ones as for any unreadable state, and all of them cost: **offline state rewriting** — reading the stored snapshot as an ordinary dataset outside the running job, relabelling the entries and writing a new snapshot the next build can read, available on some runtimes and not others — or rebuilding the retained set from the source, for as long as the source still holds the period. ## The practice this produces - Pin an identifier on every stateful step **in the first deployment** of any job expected to live, and pin one on steps that are not yet stateful but might become so. - Choose names that describe the step's job, not its position or its current implementation, so a rewrite of the step keeps the name. - Where the runtime offers nothing to pin, the equivalent discipline is to treat the shape of the stateful part of the program as a published interface: change what happens around it freely, change it itself only with a migration plan. - Record, next to the job, which stateful steps hold entries that cannot be recomputed, because those are the ones where a lost label is not recoverable at any price.

  • The job has run for a year with no identifiers pinned. Can you add them now and be safe from then on?
    Adding them changes what the new steps ask for, while the entries on storage stay filed under the old derived labels — so the release that pins the names is itself the one that loses the state, unless you relabel the entries by rewriting the stored snapshot offline, or accept a rebuild from source. Safe from then on, yes; free, no.
  • Two stateful steps swap places in an edit. Why is that worse than one of them losing its entries?
    Losing entries is visible: a total restarts, a duplicate slips through, and a check of a history-dependent number catches it. Swapped identities mean both steps find a full, plausible retained set belonging to the other step, so the job runs happily and produces confidently wrong answers with nothing to alert on.
  • Does pinning an identifier protect the state against any change to that step?
    No. It settles only the identity match. The stored values must still be readable by the new code, so a changed value shape or a changed record encoder — the code that turns a stored value into bytes and back — can break the redeploy with the identity matching perfectly.

saying these in an interview costs you the question

  • Believes the runtime matches state by comparing the old and new source
  • States a pinned author-set identifier as the mechanism every engine uses
  • Thinks only edits to the stateful step itself can affect it
  • Assumes an identifier pinned later applies to entries already written
  • Names the step after its position or its current implementation
  • Treats a matched identity as proof the stored values are readable