skip to content

Why does a platform replace the copies of a data-owning, individually-named workload one at a time in a fixed order?

level: middleimportance: must knowfreq 58%

answer

  1. not interchangeable copies
  2. one name, one store, one holder
  3. no spare copy to surge with
  4. a majority has to stay alive
  5. replaced in reverse of start-up order

basics

~20 s

Because those copies are not interchangeable. Each owns a slice of durable data and a name the replacement must reclaim, and the group needs a majority alive, so only one may be missing at a time.

solid answer

~50 s

Interchangeable copies get replaced in batches, sized by a surge allowance and an unavailability budget, because any copy serves any request. A data-owning copy holds one slice of the workload's data on a store attached to it and answers to a name its peers dial directly, so removing two removes two slices and can cost the group its majority. The platform therefore serialises: stop one copy, create its replacement under the same name, reattach the same store, wait until it has rejoined and caught up, and only then touch the next. Surge normally is not available here — you cannot run a second holder of a unique name, or a second writer against a store that admits one — so the sequence is strictly stop-then-create and the group genuinely runs one member short for the length of each step.

go deeper

for a junior

Remember the distinction itself: some copies are interchangeable and some own a slice of data and a name. The second kind is replaced one at a time.

for a middle

Explain the mechanics: stop, release the store, create under the same name, reattach the same store, wait, next. Say why an extra copy is not available here.

for a senior

Show that you have watched one: name the majority the group needs, the per-member catch-up time, and what the total duration means for your maintenance window.

for a principal

The trade is redundancy against update latency. Sizing a group so that a planned one-member absence still leaves room for an unplanned one is a design decision, not a rollout setting.

A workload's copies come in two kinds, and the replacement policy follows entirely from which kind you have. ## Interchangeable copies and copies that own something An **interchangeable** copy holds nothing that outlives it. Any request may go to any of them, one is as good as another, and the platform may destroy several at once and create their replacements anywhere there is room. A **data-owning** copy is the opposite. It holds one slice of the workload's durable data on a store attached to it, it answers to a name of its own that peers and sometimes clients address directly, and the name and the store are bound together. Replacing it is not "make another one of these"; it is "hand this exact slice and this exact name to a process running different code". A sharded search index is the standard example: five copies, five shards, each copy writing its shard to its own store and replicating to the next. ## Why the steps are serialised - **Each absence costs a distinct slice.** Two interchangeable copies missing costs you two units of capacity; two data-owning copies missing costs you two *shards*, and the queries that need them cannot be served by anyone else. - **The group usually needs a majority.** Peer groups that replicate among themselves stop accepting writes when fewer than half the members are present. In a five-member group, one absence is survivable and two is a coin toss against an unplanned failure. - **Rejoining is not instant.** A replaced copy has to reattach its store, replay whatever it had not yet flushed, and catch up on everything its peers wrote while it was gone. Copies replaced in parallel are all behind at the same time. - **The blast radius of a bad version is data, not just traffic.** A broken interchangeable copy serves errors; a broken data-owning copy can write into the slice it owns. ## Why there is normally no extra copy to work with Batched replacement of interchangeable copies can use a **surge** allowance: extra copies created *before* the old ones go, so served capacity never dips. Ordered stateful replacement normally cannot: - the name admits **one holder**, so the replacement cannot be created while the copy it replaces still exists; - the store admits **one writer**, so attaching it to two copies at once is either refused or corrupting; - therefore the step is strictly stop, then create, and the group is one member short until the new copy is caught up. The one honest exception is not a replacement at all: some systems let you *add* a member, let it copy the data, and then remove the old one. That costs a full extra copy of the slice and a second reshuffle, which is why it is reserved for growing a group rather than for updating one. ## The order, and why it runs backwards Members of such a group are brought up in ascending order because later members join the ones already running. Platforms that order the replacement run it the other way — highest first — and the reason is the same dependency: 1. the **newest** member, the one nothing else was brought up behind, takes the first risk; 2. each following step moves down toward the member the rest of the group is most sensitive to; 3. that member is replaced **last**, so if you stop early, the part of the group you have not touched is the part you most want intact. ## What one step actually does 1. Ask the copy to stop, and wait out its grace period. 2. Release its store. 3. Create a copy with the **same name and position**, running the new version. 4. Reattach the **same** store — the replacement warm-starts from the data already there; only the container's throwaway writable layer is lost. 5. Wait for the workload's own signal that it has rejoined and caught up. 6. Only then move to the next copy. ## The two policies side by side | | interchangeable copies | data-owning copies, ordered | |---|---|---| | what a copy holds | nothing durable | one slice, on its own store | | addressed as | any member of a pool | its own stable name | | how many change at once | a batch, sized by two budgets | exactly one | | extra copies during the change | allowed, via a surge allowance | normally impossible | | gate before the next step | the new copy reports ready to serve | the new copy reports caught up with its peers | | worst moment | the whole batch is missing | one member missing, for much longer | The practical consequence is that the rollout's wall-clock duration is the number of members multiplied by how long one member takes to come back, and that number, not the platform, is what decides whether the update fits in a maintenance window.

  • Can an ordered stateful replacement ever use an extra copy the way a batched one uses surge?
    Only where the name and the store are not single-holder, and then it is really a different operation: add a new member, let it copy the slice from its peers, then remove the old one. It costs a full extra copy of the data and a second rebalance, so it is used to grow a group rather than to update one.
  • What happens to a copy's data when its container is replaced during the update?
    Nothing recreates it. The store is detached from the old container and reattached to the replacement under the same name, so the new version opens the data its predecessor left. Only the container's own writable layer — anything written outside the attached store — is lost, which is why that layer is the wrong place for anything the next version needs.
  • Does raising the unavailability budget make this rollout faster?
    No. That budget governs how many interchangeable copies may be missing at once, and it has nothing to change here: the constraint is that each name and each store admits one holder, and that the group needs a majority. The lever that actually shortens the rollout is the per-member catch-up time.

saying these in an interview costs you the question

  • Says the replacement gets a fresh, empty store each time
  • Thinks an extra copy can be created first, as with interchangeable replicas
  • Wants to raise the unavailability budget to replace two at once
  • Treats the order as arbitrary rather than the reverse of start-up
  • Calls the copies interchangeable because they run the same image