skip to content

When should data movement be split out of the schema change set rather than shipped inside it, and what does that split cost?

level: principalimportance: should knowfreq 45%

answer

  1. the release waits for whatever is inside it
  2. two conditions, not one
  3. correct against half-migrated data
  4. history stops reproducing the data
  5. name the completion criterion and owner

basics

~20 s

Split data movement out when the run would hold the release open and the new code is correct against incomplete data. The cost: the change history no longer reproduces the data, and an unfinished run tends to become permanent.

solid answer

~50 s

Anything inside the ordered change set runs while the release waits for it, so a data movement measured in hours turns a deploy into an outage window and makes an abort ugly. Splitting the data run out — shipping the structural change with the release and driving the data separately — removes that, but only if the new code is genuinely correct against **half-migrated** data. Three costs are worth naming. The change history stops being a complete description of the database, so replaying it from empty no longer reproduces what production holds. Somebody must own finishing the run, and the change that hardens the result — the constraint or the removal of the fallback path — can only ship once it provably has. And the tolerance code written "for the migration window" outlives the window unless a specific person and a specific follow-up change are named up front.

go deeper

for a junior

Recall that anything inside a schema change set runs while the deploy waits, so a long data change is sometimes shipped separately from the structural one.

for a middle

Explain both conditions for splitting: the run is long enough to matter, and the new code works correctly against rows the run has not reached yet. Know why the second is the one teams skip.

for a senior

Show what you put in place around a split run — a completion criterion expressed as a query, a named owner, and the follow-up change that hardens the result and deletes the fallback.

for a principal

Weigh the durable cost: once data movement lives outside the ordered history, replaying that history no longer reproduces the database, and every rebuild and recovery inherits the gap unless you deliberately close it.

## Why the question exists at all A schema change is usually near-instantaneous; the data movement attached to it is not. When both live in the same ordered change set, the release waits for the slow half. That has three consequences a lead has to weigh: the deploy's duration becomes the data's duration, an abort part-way through leaves an ambiguous state, and every instance waiting to start is waiting on a data run rather than on a structural change. So there is a genuine architectural choice: **is the data movement part of the release, or a separate piece of work the release merely permits?** ## When keeping it inside is right It is not automatically wrong to ship data movement inside the change set — for most changes it is the simplest correct answer. Keep it inside when: - the affected row count is small enough that the run is bounded and predictable; - the new code **cannot** function against un-migrated rows, so the release must not begin serving until the data is complete; - the movement and the structure are one logical change and splitting them creates a window nobody would otherwise have to reason about; - the data is reference or lookup rows, which are tiny and are part of the schema's meaning. The advantage of staying inside is that the history remains a complete description: replay it from empty and you get the same database, structure and data alike. ## When to split it out Split it when both halves of this are true: 1. **The run is long enough to matter** — long enough that holding the release open for it is unacceptable, or that it must be startable, stoppable and restartable on its own schedule. 2. **The release is correct without it.** The new code must handle a row whose new value is absent: derive it at read time, fall back to the old column, or treat absence as a defined state. If it cannot, splitting does not remove the dependency — it converts a slow deploy into a partially broken one. The second condition is the one teams skip. The honest test is to ask what a user sees for a row the run has not reached yet, and to be able to answer without saying "the run will be finished by then". ## What the split costs | Cost | Why it bites | | --- | --- | | The history no longer reproduces the data | A database rebuilt from empty has the structure but not the movement, so environments diverge | | Applied-state tracking is now yours | Nothing records that the run finished in a given environment; you need your own record | | Two code paths live simultaneously | The read-time fallback exists only for the window, but it is real code with real bugs | | Completion becomes a prerequisite | The constraint that hardens the column, and the deletion of the fallback, cannot ship until the run provably finished everywhere | | Ownership drifts | A run with no named owner stalls, and a stalled run is invisible until something depends on it | The deepest of these is the first. Once data movement lives outside the ordered history, "replay the history" and "reproduce the database" stop being the same statement, and every environment rebuild, every recovery drill, and every new developer's local database inherits the gap. Some teams close it by re-expressing the finished movement as a cheap statement appended to the history afterwards, once the volume is known to be trivial on a fresh database — a fresh build has no legacy rows, so the same movement is often instant there. That is worth doing deliberately rather than discovering the divergence later. ## Making the split finish A split data run is a piece of work with a beginning and an end, and the end is the part that gets lost. Name three things when you approve the split: 1. **A completion criterion** — a query whose result is zero, not a feeling that it is probably done. 2. **A named owner** for the run in each environment, because "the team" does not restart anything. 3. **The follow-up change** that hardens the result and deletes the tolerance code, written as a real ticket at the time of the split, not afterwards. Without them the common outcome is an application permanently written to tolerate a state that was supposed to last a week, and a column that can never be made required because nobody can prove it is fully populated. That is the true cost of the split, and it is why the default for anything small should remain: ship it inside the change, where finishing is not optional.

  • What is the honest test of whether the release is correct without the data run?
    Ask what a user sees for a row the run has not reached, and require an answer that does not depend on timing. If the code derives the value, falls back to the old column, or treats absence as a defined state, the release is independent. If the answer is "it will be done by then", the dependency is still there and merely hidden.
  • How do you stop a split run from leaving the change history unable to reproduce the database?
    Close the gap deliberately: once the movement has finished and its volume is understood, append the same transformation to the history as a cheap statement. A freshly built database has no legacy rows, so it is usually instant there, and replay again produces the same data as production.
  • Why is the tolerance code the part that tends to survive?
    Because nothing fails while it exists. The fallback path keeps working, the constraint is simply never added, and the ticket to remove it competes with feature work forever. The only reliable counter is to create that follow-up change when the split is approved and tie it to a completion criterion that someone owns.
  • Does splitting always beat keeping the movement inside the change?
    No. For small, bounded data — reference rows, a few thousand updates — keeping it inside is simpler and stronger: finishing is not optional, the history stays complete, and no tolerance code is needed. Splitting buys a shorter release at the price of a window, so it should be justified by duration, not adopted by habit.

saying these in an interview costs you the question

  • Splits the data run out but ships code that assumes it has already finished.
  • Treats splitting as the default rather than something duration justifies.
  • Forgets that the change history no longer reproduces the data after a split.
  • Has no completion criterion, so nobody can say whether the run finished.
  • Leaves the tolerance code and the missing constraint in place indefinitely.
  • Assumes a run with no named owner will be restarted after it stalls.