skip to content

What must an ordered stateful rollout wait for before it touches the next copy, and why is a started process not enough?

level: middleimportance: should knowfreq 46%

answer

  1. started is not the same as caught up
  2. one stopped is not one unusable
  3. the workload publishes its own lag
  4. a shallow gate stacks up behind members
  5. no deadline that continues anyway

basics

~20 s

It must wait until the replaced copy has rejoined its group and caught up. A started process answers its port long before its store is replayed, so advancing on that gate stacks up members that are behind.

solid answer

~40 s

The gate has to mean "this member is carrying its share again": store reattached, unflushed work replayed, replication lag back inside an acceptable bound, and the peers counting it as a full member. A process-is-running check, or a shallow ready check on a port, goes true seconds after start-up and minutes before any of that is true. Advance on the shallow gate and the rollout walks the whole group faster than the data can catch up — nominally one copy is missing at a time, in reality three are unusable at once and the group can lose its majority. Only the workload knows its own lag, so the gate has to be a signal the workload publishes, and the platform must not have a timeout that shrugs and continues.

code

pseudocode · 11 lines
pseudocode
for copy in group ordered by ordinal, highest first:
    stop(copy)                       # ask it to stop, wait out the grace period
    detach(copy.store)
    create(copy.name, newVersion)    # same name, same position
    attach(copy.store)               # the predecessor's data, not a fresh store

    wait until copy.caughtUp         # published by the workload:
                                     #   rejoined the group AND lag within bound
                                     # no deadline that continues on expiry

    # only now is the next copy touched

go deeper

for a junior

Learn the distinction first: a process that has started is not the same as a member that has its data back. The rollout has to wait for the second one.

for a middle

Explain what happens between the two — store reattached, unflushed work replayed, group rejoined, lag closed — and why only the workload can report it.

for a senior

Work the failure out loud: a five-member group on a two-second gate loses its majority while the platform still believes one copy was ever stopped.

for a principal

Per-member catch-up time is the multiplier on every rollout the workload will ever have. Treating a long one as a defect rather than a schedule is where the leverage is.

Ordered one-at-a-time replacement is only as safe as the condition that separates one step from the next. That condition is the whole mechanism; everything else is bookkeeping. ## What "one at a time" actually promises The promise is not "only one copy is stopped at a time". It is "only one copy is **unusable** at a time". Those are the same thing only if the gate between steps is true when the copy is genuinely usable again. If the gate is true earlier, the two claims come apart immediately, and the group's safety margin comes apart with them. ## The three gates, and what each one really means | gate | goes true when | what it misses | |---|---|---| | the process is running | the first process has not exited | says nothing about data at all | | a shallow ready check | a port accepts a connection, or a handler returns success | true seconds after start, while the store is still being replayed | | the workload's catch-up signal | the member has rejoined and its lag is inside a bound | nothing relevant — but only the workload can compute it | A liveness-style check answers "should this container be restarted?". A ready check answers "may this copy be sent traffic?". Neither answers "is this member carrying its slice again?", and that third question is the only one an ordered stateful rollout may advance on. ## What the copy is doing while the shallow gate is already green 1. **Reattaching the store** it inherited from its predecessor, and opening the data files on it. 2. **Replaying** whatever was written but not yet folded into the main data structure at the moment it was stopped. 3. **Rejoining** the peer group and being recognised as a member again. 4. **Catching up** on everything its peers accepted while it was away — which grows with how long the step took and how busy the workload is. Step 4 is usually the long one, and it is invisible from outside the process. There is no generic platform signal for it, which is exactly why the gate has to be something the workload publishes about itself. ## How a shallow gate fails, concretely Take a five-member group, each member needing about ten minutes to catch up, and a gate that goes true two seconds after the process starts: - at minute 0 the highest member is replaced and the gate goes green almost at once; - by minute 1 the next is stopped, while the first is still nine minutes behind; - by minute 2 a third is stopped, and now **three** of five members are either absent or unusable. The platform's own view is that it never had more than one copy stopped, and it is right about that and wrong about everything that matters. The group has lost its majority and stops accepting writes, and the rollout keeps marching. ## Designing the gate - **Let the workload own it.** Publish a signal that is true when the member's replication lag is inside a threshold *you* chose, not when a handler returns. - **Choose the threshold from the failure you are avoiding**, usually "could this member serve its slice if another one failed right now?". - **Do not let it time out into a pass.** A gate with a deadline that continues anyway is not a gate; if catch-up is genuinely slower than the deadline, the rollout should sit and wait for a human. - **Do not confuse it with the restart-on-failure check.** If your catch-up signal is wired to the check that restarts containers, a slow-catching-up member gets killed and restarted, which resets its catch-up and guarantees it never finishes. - **Measure the real per-member time** before you plan the rollout: it is the multiplier on the whole thing. The last bullet is the one people are surprised by. Because the replacement reattaches its predecessor's store, catch-up should be short — a replay plus the writes accepted during the step. When a member instead takes tens of minutes, the usual cause is that it is rebuilding its slice from its peers rather than opening the store it inherited, and that is a defect to fix rather than a duration to schedule around.

  • What breaks if the catch-up signal is also wired to the check that restarts a failing container?
    A member that is legitimately slow to catch up gets killed for it. The restart throws away the progress it had made, so it starts catching up from further behind, fails again, and loops. Catch-up is a reason to wait, never a reason to restart; keep the two signals separate and give the restart check a far looser threshold.
  • Why should the step not have a timeout that advances the rollout anyway?
    Because the timeout fires exactly when the assumption behind one-at-a-time is false. If a member cannot catch up inside the expected window, continuing adds a second unusable member to the first, which is the failure the serialisation existed to prevent. Stopping and surfacing it is the correct behaviour.

saying these in an interview costs you the question

  • Says a passing port check means the member is back in the group
  • Treats catch-up as instant because the store was reattached
  • Lets the step time out and advance to the next copy anyway
  • Assumes the platform can measure catch-up without the workload
  • Thinks one-at-a-time alone means only one member is ever unusable