skip to content

Which changes to a stateful step's stored values can still be read back after a redeploy, and which cannot?

level: middleimportance: should knowfreq 50%

answer

  1. sort the edit, not the release
  2. additive usually survives
  3. renamed, re-keyed, re-encoded does not
  4. the encoding decides what may be added
  5. a rename can fail with no error at all

basics

~20 s

Additive changes usually survive: a field added with a default supplied for entries that lack it, and a widened type. Renames, a changed grouping key, a swapped record encoder and removed or reordered stateful steps do not, because the stored bytes no longer answer what the new code asks.

solid answer

~50 s

Sort every edit into what the stored bytes can still answer. Usually readable: adding a field to a stored value when the reader supplies a default for entries written without it, and widening a type the encoding can widen. Usually not readable: renaming a field where the encoder identifies fields by name, changing the grouping key or its type — every entry is filed under a key the new code will never ask for — swapping the **record encoder**, the code that turns a stored value into bytes and back, since the class in the source can look identical while the format changes, and removing or reordering stateful steps, which breaks identity before format is even reached. Two cautions: which additive changes are legal is a property of the encoding, not of the job, and a general-purpose language serializer may reject *any* change; and the failure can be a refused start, a silent empty entry or a default read where a real value was expected.

code

json · 5 lines
json
{
  "writtenByOldBuild": { "count": 17, "lastSeenMillis": 1758300000000 },
  "readByNewBuild_fieldAdded": { "count": 17, "lastSeenMillis": 1758300000000, "channel": "unknown" },
  "readByNewBuild_fieldRenamed": { "eventCount": 0, "lastSeenMillis": 1758300000000, "channel": "unknown" }
}

go deeper

for a junior

Recall the direction of travel: adding a field with a default is usually fine, and renaming a field or changing the grouping key usually is not. Knowing the two sorts exist is enough at this level.

for a middle

Explain why additive edits work — old entries still answer what the reader asks, and the missing field takes a default — and name the edits that change what a value is or where it is filed, including a swapped record encoder.

for a senior

Show the failure modes: that a rename on a name-identified encoding produces defaults with no error at all, and that verification must run against entries written by the previous build, never against a freshly started job.

for a principal

Set the standard: the stored format of every long-lived stateful step is a published contract with a named owner and a review rule, so no team discovers the encoder was part of the contract during an incident.

## Sort the edit, not the release A release is a bundle of edits and tells you nothing. What decides whether a redeploy reads its state back is each individual edit, sorted by one question: **can the bytes already on storage still answer what the new code asks of them?** There are three sorts, and the third is the one that is missed. | Sort | Examples | Usual outcome | |---|---|---| | Does not touch stored values | changed business rule, new output field derived per record, a step added that holds nothing | reads back; the entries are irrelevant to the edit | | Additive to a stored value | a field added with a default supplied for entries that lack it, a widened numeric type | reads back **where the encoding supports it** | | Changes what a stored value *is*, or where it is filed | rename, re-key, changed collection shape, swapped record encoder, removed or reordered stateful step | does not read back | ## Why the first two usually work An entry written last month contains the fields the old code stored. If the new code asks for those fields plus one more, and can say what the missing one should be for old entries, every old entry is still usable — the accumulated value survives and the new field starts at its default. Widening a type is the same argument: the old value is representable in the new type, so nothing is lost. Both depend entirely on the **record encoder** — the code that turns a stored value into bytes and back. The rules for which additive changes an encoding permits belong to the encoding, not to the job, and they differ sharply: a schema-driven encoding is built for exactly this and has explicit rules for defaults; an encoding that writes only a language object's default serialized form may reject any change to the class at all. The practical consequence is that *what you may add is a property of how you chose to store the value*, which is why that choice belongs in the first deployment. ## Why the third sort does not work - **Rename.** Where the encoder identifies fields **by name**, the reader asks for a name the entry does not carry, so the accumulated value is invisible and the new field starts at its default — with no error. Where the encoder identifies fields by a **numeric tag**, renaming touches only the source and is harmless. You cannot tell which case you are in by reading the class; you have to know the stored format. - **Re-key.** Changing the grouping key, or its type, refiles everything. The entries are stored under keys the new code will never ask for, and they cannot be translated in place because a different key means a different bucket and usually a different owning worker. - **Changed shape.** A single stored value that becomes a collection of values, or a collection whose element type changes, is a different stored thing wearing the same field name. - **Swapped record encoder.** The class in the source is unchanged and the bytes are in a different format. This is the edit people do not see coming, because the diff looks like configuration. - **Removed or reordered stateful steps.** This fails the identity match before format is ever considered: entries are filed under a step that no longer exists, or a step matches entries another step wrote. ## What failure looks like The same three behaviours as any failed redeploy, and they differ by runtime: a **refused start** naming the step; a **silent start with an empty retained set**; or a **deferred failure** at the first read of an entry the new code cannot decode, possibly hours in. Rename is the nastiest of the set, because on a name-identified encoding it produces none of the three — it produces a correct-looking job reading defaults. ## How this is worked in practice 1. For each edit in the release, ask what it does to a value already on storage; if the answer is not obvious, the edit is in the third sort until proven otherwise. 2. Prefer additive edits. A field you stop using can stay stored and ignored far more cheaply than a field you rename. 3. When an edit belongs to the third sort, it needs one of the escape routes — a converter, a transition window, or a rebuild — and those cost time and correctness, so batching several such edits into one deliberate migration beats spreading them over three releases. 4. Verify against real stored entries, not a fresh job. A redeploy tested on a job started five minutes ago has tested nothing: it is the entries written by the *previous* build, in the *previous* format, that carry the risk.

  • A field is removed from a stored value. Is that readable?
    Usually yes, and it is the cheaper direction: old entries carry a field the new reader ignores. But it is one-way — once entries are written without it, going back means the field is absent for every entry in between, so treat a removal as a decision you will not reverse rather than as a harmless tidy-up.
  • Why is swapping the record encoder so easy to miss in review?
    Because the stored value's class is untouched, so the diff shows only a changed setting or a changed default, and nothing in the source says the bytes on storage now mean something else. It is worth treating the encoder of any stateful step as part of its published contract and calling it out explicitly in review.
  • The team wants to add a new stateful step to an existing job. Is that a compatibility problem?
    Not a compatibility one — a new step simply starts with an empty retained set, and the existing steps are unaffected as long as the edit does not shift their identity. It is a correctness problem instead: the new step's answers are wrong until it has seen enough input to be meaningful, and consumers need to be told which period that covers.

saying these in an interview costs you the question

  • Judges a release as a whole instead of sorting each edit
  • Says a rename is harmless because the value's position is unchanged
  • Claims a rename always loses the value, whatever the encoding
  • Thinks the class in the source determines the stored format
  • Treats a re-key as just another field change
  • Tests the redeploy on a freshly started job with no old entries