skip to content

Each copy of a five-member data group needs forty minutes to catch up after replacement — how would you judge whether that rollout is acceptable?

level: principalimportance: should knowfreq 34%

answer

  1. members times per-member catch-up
  2. challenge the number before scheduling around it
  3. a rollout spends redundancy bought for failures
  4. linear growth means an eventual wall
  5. split into groups that roll in parallel

basics

~20 s

Judge it on three things: whether forty minutes is inherent or a defect, whether redundancy still covers an unplanned loss while one member is away, and whether a rollout that long can finish inside any window you can commit to.

solid answer

~50 s

Start by challenging the number. The replacement reattaches its predecessor's store, so catch-up should be a replay plus the writes accepted during the step; forty minutes usually means the member is rebuilding its slice from peers instead, and fixing that is worth more than any scheduling. If it is genuinely inherent, then the rollout is about three and a half hours during which the group runs a member short for each step, so size the group for a planned absence plus an unplanned failure rather than for failures alone. Then check the shape at scale: duration grows linearly with membership, so a hundred-member group is days and one-at-a-time stops being viable — at that point you split the workload into independent groups that roll in parallel. Finally, set policy: a measured worst-case duration per workload, and a catch-up gate the workload publishes.

go deeper

for a junior

The arithmetic is the entry point: total time is the number of members multiplied by how long one member takes to come back.

for a middle

Explain why a reattached store should make catch-up short, and what it means when it does not.

for a senior

Show the operational consequence: for each step the group is a member short, so one unplanned failure during the rollout is the one that hurts.

for a principal

Own the structural call — how much redundancy the rollout is allowed to spend, and the group size at which the workload must be split rather than rolled.

This is a capacity and risk question wearing a rollout's clothes. The arithmetic is trivial; everything interesting is in what you do about it. ## Do the arithmetic first, and state the assumptions Assume five members, forty minutes of catch-up each, and a couple of minutes to stop and start: - one full ordered rollout is about **5 x 42 minutes, a little over three and a half hours**; - for each of those steps the group is running **one member short**, and with a proper catch-up gate it is back to full strength between steps; - the mixed-version window lasts the **whole** rollout, not a moment of it. The same shape at a hundred members is 100 x 42 minutes, about **2.9 days**. That is the number that decides whether one-at-a-time is a strategy or a trap. ## Challenge the forty minutes before you plan around it The replacement inherits its predecessor's store, so an honest catch-up is a replay of unflushed work plus whatever the peers accepted during the step — minutes, usually. When it is tens of minutes, look for a cause rather than a schedule: - the member is **rebuilding its slice from peers** instead of opening the store it inherited, which means the store is not actually being reattached, or the version change invalidates what was on it; - the stop is **not clean**, so there is no recent checkpoint and the replay starts far back; - the catch-up is **rate-limited** to protect live traffic, in which case the number is a deliberate trade and can be tuned per rollout; - the slice is simply **too large**, which is a sharding decision, not a rollout one. Only the last two are inherent. Halving per-member catch-up halves every future rollout of this workload, which is a far better return than negotiating a longer maintenance window. ## Then decide what redundancy you are actually buying A group sized so that it survives exactly one failure has no margin during a rollout: for each step, the planned absence has already spent it, and a single unplanned loss at that moment takes the group below its majority. Three practical positions: | position | what it means | what it costs | |---|---|---| | size for failures only | any rollout step plus one failure is an outage | cheapest, and honest only if outages during rollouts are acceptable | | size for a failure during a rollout | one extra member so a step and a failure overlap safely | one member's worth of hardware and one more slice to keep caught up | | never roll during risk windows | keep the smaller group, forbid rollouts near peaks | free in hardware, expensive in patch latency | The second is what most groups holding real data should choose, and the reason to state it explicitly is that the cost shows up in the capacity budget while the benefit shows up as an outage that did not happen. ## Then decide the policy other teams have to follow 1. **A declared worst-case rollout duration per workload**, measured rather than estimated, published beside the workload. Everything downstream — patch cadence, incident expectations, freeze windows — depends on it. 2. **The advance gate must be the workload's own catch-up signal**, never a port check, and it must not time out into a pass. 3. **A rule for holding across a peak**: either the group may sit part-replaced through peak traffic, in which case the mixed-version obligations are permanent requirements, or it may not, in which case the rollout must be startable early enough to finish. 4. **A size at which the group must be split.** Once a full rollout cannot finish inside the longest window you can commit to, the answer is not a faster rollout; it is several independent groups, each rolling one member at a time in parallel with the others. ## The reasoning to show The trap in this question is treating the three and a half hours as the problem. It usually is not. The problem is that the duration is multiplied by a per-member cost nobody has examined, that it consumes redundancy bought for failures, and that it scales linearly into a wall. A good answer challenges the number, prices the redundancy, and names the size at which the shape has to change.

  • At what point does one-at-a-time stop being viable, and what replaces it?
    When the full rollout can no longer finish inside the longest window you can commit to — duration grows linearly with membership, so this arrives suddenly for large groups. The replacement is not a faster rollout but a different shape: split the workload into several independent groups, each with its own majority, each rolling one member at a time in parallel with the others.
  • What would make you suspect the forty minutes is a defect rather than a fact?
    If the time is roughly proportional to the size of the whole slice rather than to the length of the step. Catch-up after reattaching a store should scale with what happened while the member was away; catch-up that scales with total data means the member is rebuilding from its peers, which usually means the store is not actually being reattached.
  • How does this change the patch cadence you can promise?
    It sets a floor on it. If a full rollout is three and a half hours you can patch this workload weekly without difficulty; if it is three days you cannot honestly promise anything faster than that, and any security commitment shorter than the rollout duration is a commitment the workload's shape cannot keep.

saying these in an interview costs you the question

  • Accepts the per-member catch-up time without asking why
  • Counts a planned absence as free against the group's redundancy
  • Assumes the rollout duration stays flat as the group grows
  • Proposes replacing several members at once to save time
  • Treats the mixed-version window as a moment rather than the whole rollout