A job holding months of per-key entries must change its grouping key. What are the options, and what does each cost?
answer
- the entries are filed, not just shaped
- three routes, no free one
- bounded lifetime enables the window
- source retention enables the rebuild
- the plan is a day-one artefact
basics
~20 sThree routes: rewrite the stored snapshot offline into the new key and start from it; run a transition window where the new build writes the new key and still reads the old; or start empty and rebuild from the source. Each trades downtime, complexity or a period of wrong numbers.
solid answer
~50 sA changed grouping key refiles everything, so no entry can be read where the new code looks for it — this is a migration, not a release. Three routes exist. **Rewrite offline**: read the stored snapshot as an ordinary dataset outside the running job, re-key the entries, write a new snapshot the new build starts from; cheapest in correctness, but this facility exists on some runtimes and not others, and the job is down for the length of the rewrite. **Transition window**: the new build writes under the new key while still reading old-key entries, until nothing old is still in range; no downtime and no wrong numbers, but two code paths, more retained bytes while both live, and it only ends if entries have a bounded lifetime. **Rebuild from source**: start empty and refill from replayed input; simple, but it needs the source to still hold the period, it runs at catch-up throughput, and the numbers are wrong or absent meanwhile.
go deeper
Recall that changing the key entries are filed under means the new code cannot find them, and that the ways forward all cost something — rewriting them, running both for a while, or starting again from the input.
Name the three routes and the precondition of each: a snapshot you can read as data, entries with a bounded lifetime, or a source that still holds the period. Say why the entries cannot simply be relabelled in place.
Show judgement: pick a route from those preconditions, put a closing condition on a transition window, and state what consumers are told about the period during which a rebuilt retained set is incomplete.
Own the day-one artefacts that make this survivable across a fleet — expiry defaults, pinned step identifiers, a recorded dependency on source retention — and the standing rule for when a retained set is too valuable or too large to live inside the job at all.
## Why a re-key is the hard case Changing the grouping key a stateful step's entries are filed under is the worst of the incompatible edits, and it is worth being precise about why. It is not that the values cannot be decoded — they can. It is that they are **filed in the wrong place**. Every entry sits under a key the new code will never ask for, and because a key determines which bucket it belongs to and therefore which worker owns it, the entries cannot simply be relabelled where they lie: moving them is a redistribution by key, the step that sends every record to the worker that owns its key. That is a job in itself, which is exactly why the routes below are the shape they are. Before choosing, establish two facts, because they eliminate routes: 1. **Do the entries have a bounded lifetime?** If an expiry rule drops an entry a fixed interval after its last write or last access, the transition window has an end. If entries live forever, it does not. 2. **Can the source still supply the period the entries cover?** If not, rebuilding is off the table and the entries are, in the strict sense, irreplaceable. ## The three routes | Route | What happens | Costs | Ruled out when | |---|---|---|---| | **Offline rewrite** | the stored snapshot is read as an ordinary dataset outside the running job, entries are re-keyed, and a new snapshot is written for the new build to start from | job is stopped for the length of the rewrite; the transformation must be written and tested; the snapshot's own format is a compatibility surface | the runtime exposes no way to read or write a snapshot as data | | **Transition window** | the new build writes under the new key and still reads old-key entries, until no old entry can still matter | two code paths in one job, both maintained; more retained bytes while both live; a date to remove the old path, or it is permanent | entries have no bounded lifetime, or the old key cannot be computed from the record any more | | **Rebuild from source** | the step starts with an empty retained set and is refilled from replayed input | the source must still hold the period; a catch-up run; wrong or absent numbers until it is caught up | the source no longer holds it, or the entries were derived from something that cannot be replayed at all | ## Choosing between them The honest decision procedure is short: - If the runtime supports reading and writing a snapshot as data and the stop is acceptable, **rewrite offline**. It is the only route that preserves every entry exactly and finishes at a known moment. - If a stop is not acceptable and entries expire, run a **transition window** and write the removal date into the work item that opens it. The failure mode of this route is not technical, it is organisational: the second code path is never deleted, and two years later nobody knows whether the old branch still fires. - If neither is available, **rebuild**, and treat the period during which the numbers are rebuilding as a consumer-facing fact: say which outputs are incomplete and until when. The mechanics of pushing a long period of historical input through a job is a separate subject with its own owner; what belongs here is that it is the fallback, and that its availability is decided by how long the source retains data — a limit set by somebody else, often without knowing your job depends on it. ## What to do with the entries you cannot rebuild Some retained sets cannot be reconstructed at any price: a running value derived from records the source has since dropped, an entry set fed by an input that no longer exists, an aggregate over a period nobody archived. For those, the offline rewrite is not the cheap option but the **only** one, and if the runtime does not offer it, the honest answer is that the re-key cannot be done without losing history — at which point the design conversation is whether the retained set should have been inside the job at all. ## The judgement this is really testing An interviewer asking this is checking whether you know that the plan is written **before the first deployment**, not at the moment of the re-key. Concretely: an expiry rule, so a transition window can ever end; a stored format with room to move; a note of which entries can be rebuilt from the source and for how long that stays true; and step identifiers pinned, so the re-key is the only problem you are solving rather than one of three. Every one of those is cheap on day one and expensive or impossible on day six hundred.
- The transition window is running and the old path has been cold for a month. What has to be true before you delete it?That no entry filed under the old key can still affect an answer: the expiry interval for those entries has fully elapsed since the window opened, and late input that would have landed on an old-key entry is no longer accepted. Cold for a month is evidence, not proof — the check is the expiry horizon, and a metric counting old-path reads is what makes the deletion safe rather than hopeful.
- Why is the rebuild route often decided by somebody outside the team?Because its precondition is how long the source keeps the input, and that retention is usually set by whoever owns the source, for their own reasons and without knowing which jobs depend on being able to replay it. If your only migration plan is rebuild, your migration plan is a setting on a system you do not control.
- Could you avoid all of this by running the new build alongside the old one and cutting over?Running two builds and comparing their numbers is a different technique with its own owner, and it does not by itself move a single entry: the new build still starts with no history under the new key, so it has to be filled by one of these three routes anyway. It changes when you cut over, not what the state costs.
It is a library that decides to shelve by author instead of by subject. You can close for a weekend and reshelve everything, run both schemes and let the subject shelves empty as books are returned, or throw the catalogue away and re-enter the books from the purchase records — and the last one works only if the purchase records still go back far enough.
saying these in an interview costs you the question
- Calls a changed grouping key just another field change
- Assumes every runtime can read and write a stored snapshot as data
- Opens a transition window with no date or condition for closing it
- Assumes the source still holds every period the entries cover
- Ignores that entries with no expiry rule make the window endless
- Treats the rebuild period as invisible rather than telling consumers