skip to content

An ordered stateful rollout is held after replacing only the highest-numbered copies — what state is the group left in?

level: seniorimportance: nice to knowfreq 28%

answer

  1. a hold is deterministic, not partial
  2. above the line new, below it old
  3. the only staged exposure available here
  4. you own the mixed window while held
  5. holding is not reverting

basics

~20 s

A cleanly split group: every copy above the hold position runs the new version with its own store intact, every copy at or below it is untouched on the old one. Nothing is half-replaced, and nothing is reverted.

solid answer

~50 s

Because the replacement is serialised in a fixed order, stopping it is deterministic rather than messy — the position of the hold tells you exactly which members are on which version. That determinism is what makes a deliberate hold useful: it is the only staged exposure available to a workload that cannot run an extra copy, so you can put the new version on one member of a real group, let it take real traffic and real replication for a day, then release the hold. What you own while held is the mixed-version window for as long as you stay there. What a hold is not is a reversion: the replaced copies are already running the new code and their stores may already hold what it wrote, so moving the hold back down does not un-write anything.

code

yaml · 10 lines
yaml
update:
  order: highestFirst      # the reverse of the start-up order
  concurrency: 1           # one copy replaced at a time
  holdBelowOrdinal: 3      # positions 3 and 4 are replaced; 0, 1 and 2 are not
  advanceWhen: caughtUp    # the workload's own catch-up signal, not a port check

# group of five, positions 0..4
#   4 -> new version, caught up
#   3 -> new version, caught up
#   2, 1, 0 -> untouched, previous version

go deeper

for a junior

The key idea is that the stop is clean: whole copies are on one version or the other, and the position of the hold tells you which.

for a middle

Explain why the determinism follows from serialising the steps, and what the mixed-version group can and cannot do while it is held.

for a senior

Use it deliberately: advance one member, hold across a full traffic cycle, and say out loud what the hold is costing you each hour.

for a principal

The trade is evidence against exposure. Every hour held buys observation and accumulates durable data written by a version you have not committed to.

A rollout that stops part-way is normally an incident. On an ordered stateful replacement it is a well-defined state, and it is useful enough that platforms let you ask for it deliberately. ## Why the stopped state is clean Replacement here is strictly serialised: one copy at a time, in a fixed order, with a gate between steps. At any instant exactly one copy is in transition and every other copy is wholly on one version or the other. So a stop — whether you asked for it or a gate never went true — leaves: - every copy **above** the hold position on the new version, with its own name and its own store, caught up and serving; - every copy **at or below** it untouched, still running the previous version, never stopped at all; - **no copy in between.** There is no partially-updated member, because the unit of change is the whole copy. This is the opposite of what people expect from a stalled batched rollout, where the batch boundary is a scheduling detail. Here the boundary is a position in the order, and it is the same position for everyone reading the system. ## Why you would ask for it A data-owning group cannot run an extra copy to test with — the name and the store admit one holder. That removes the usual options for exposing a new version to a slice of reality before committing. The hold puts them back: 1. Advance the rollout by exactly one copy, the newest member, the one nothing else was brought up behind. 2. Leave it there under real load and real replication for as long as you want evidence for — hours, or a day. 3. Release the hold and let the remaining copies follow, or move the hold back up and replace the one member you advanced. That is genuine staged exposure, bought entirely out of the ordering. ## What the hold costs while it lasts - **You own the mixed-version window for the duration.** Two versions of the same group are replicating to each other and reading stores of one format for as long as you hold, which may be far longer than a rollout would have taken. - **The group's capabilities are the intersection.** Anything the new version can do that requires every member to understand it is unavailable until the last copy is replaced. - **Operational surface doubles.** Two sets of logs, two sets of metrics, two behaviours to reason about at three in the morning, and any incident during the hold needs the hold in its first sentence. - **The last step is still untested.** The member replaced last is the one the rest of the group was brought up behind, so the riskiest step in the rollout is precisely the one a hold near the top has not exercised. ## A hold is not a reversion This is the confusion worth naming explicitly. Holding stops *further* replacement; it does nothing to the copies already replaced: - they are running the new version, not the old one; - their stores already contain whatever the new version has written since it started; - moving the hold position back down does not put the old code back on them by itself, and it certainly does not un-write data. Going back to the previous version is a separate operation with its own risks, and the durable data written during the hold is exactly what makes it hard. Deciding how long to hold is therefore also deciding how much new-version data you are willing to have in the stores before you commit. ## Reading the state from outside When you inherit a group in this state, three facts tell you everything: the hold position, the version each side of it runs, and how long it has been there. From the first two you know which members can do what; from the third you know how much data the new version has already written, and therefore how expensive the decision in front of you is. The decision itself is only ever one of three, and naming them keeps the conversation short: - **release the hold** and let the remaining copies follow, accepting the new version everywhere; - **stay held** deliberately, which means committing to operate a mixed group and to keep both versions compatible indefinitely; - **go back**, which is a separate operation whose cost is set by the durable data the new version wrote while you were deciding. Drifting is the fourth option nobody chooses on purpose, and it is the one most held rollouts end up in: the hold stops being a decision anybody is waiting on and becomes the group's shape, still carrying every obligation of the window that was meant to be temporary.

  • Why is the final step of an ordered stateful rollout often the one to worry about?
    Because the fixed order runs the reverse of start-up, so the copy replaced last is the one the rest of the group was brought up behind — frequently the member holding a coordinating role. A hold near the top can look completely successful for a day and still tell you nothing about that step.
  • How do you decide how long to hold?
    By what evidence you are waiting for, and against how much new-version data accumulates while you wait. A hold long enough to cover a full daily traffic cycle usually buys the most; a hold of several days mostly buys more durable data written by a version you have not committed to, plus a longer mixed-version window to operate.

saying these in an interview costs you the question

  • Calls a held rollout a reversion of the copies already replaced
  • Expects the held copies to be somewhere between the two versions
  • Thinks holding costs nothing while the group stays mixed
  • Assumes the last copy is the safest step because it is last
  • Believes releasing a hold requires restarting the whole group