skip to content

A nightly script restarts each broker node with a fixed 60-second sleep between steps; afterwards acknowledged records are missing - why?

level: seniorimportance: should knowfreq 54%

answer

  1. the sleep is not a gate
  2. catch-up time varies with traffic
  3. current copies fall step by step
  4. the last current copy goes down

basics

~20 s

The sleep was shorter than the catch-up each returning node owed, so current copies were never restored between steps. The count fell step by step until the loop stopped the node holding the last current copy of a unit.

solid answer

~40 s

Sixty seconds is a guess at the catch-up distance a returning node owes, and on a busy night it is short. Each step therefore leaves some units of ownership with one fewer current copy than before, and the shortage carries into the next step instead of being repaired. Eventually the loop reaches a unit whose only current copy is on the node it is about to stop. From there the platform either refuses writes for that unit until a current copy returns, or restores service from a copy that was behind - and every record the behind copy never received is gone, including records the broker had already acknowledged to the writer. Nothing in the loop errors, because a sleep cannot fail. The defect is the gate, not the restart command.

code

json · 6 lines
json
[
  { "step": 1, "stoppedNode": "n1", "holdsCopyOfPartitionA": false, "currentCopiesOfPartitionAAtStop": 3, "caughtUpAgainWithin60s": true },
  { "step": 2, "stoppedNode": "n2", "holdsCopyOfPartitionA": true, "currentCopiesOfPartitionAAtStop": 3, "caughtUpAgainWithin60s": false },
  { "step": 3, "stoppedNode": "n3", "holdsCopyOfPartitionA": true, "currentCopiesOfPartitionAAtStop": 2, "caughtUpAgainWithin60s": false },
  { "step": 4, "stoppedNode": "n4", "holdsCopyOfPartitionA": true, "currentCopiesOfPartitionAAtStop": 1, "caughtUpAgainWithin60s": false }
]

go deeper

for a junior

Take away the shape: a pause is not a check, and a maintenance loop that never verifies anything can take down the last good copy of some data.

for a middle

Trace the arithmetic out loud - how the count of current copies for one unit falls across the steps - and say what the platform does once it reaches zero.

for a senior

Argue for the fix at the right level: gate on the condition, verify the precondition before the first stop, and abort instead of skipping when it will not come good.

for a principal

Ask why a loop with no gate was permitted to run unattended against production at all, and where else the same pattern is walking a node set.

## How a clean-looking loop loses acknowledged records The script did everything the runbook said: one node at a time, in order, with a pause between. What it did not do is check anything. A sleep asserts that time has passed, and the risk is about whether data is current, so the loop never noticed that it was spending redundancy it was not putting back. Walk one unit of ownership - a partition, or on queue-shaped brokers a queue - through the loop. Say it is held by three nodes, and the cluster has four: 1. **Stop the node that holds no copy of it.** Nothing changes for this unit; three current copies. 2. **Stop the second node.** Its copy leaves the caught-up set. Sixty seconds later the script moves on, but on the night's write rate the copy needs several minutes. Two current copies, one behind. 3. **Stop the third node.** Same story. One current copy, two behind. 4. **Stop the fourth node** - the one holding the only current copy. Zero current copies. At step four the unit has no node that holds everything the writers were told had been stored. ## What happens at zero What the platform does next is one of two things, and both are bad in different ways: - **It refuses writes for that unit** until a node with a current copy is back. The writers see errors or stall; nothing is lost, but the outage is real and it is the operator's doing. - **It restores service from a copy that was behind.** Availability returns immediately, and every record that copy never received disappears - including records the broker had already acknowledged, which is why the loss is invisible to the writers, who were told the writes had landed. Which of those happens is a platform and configuration matter, and a candidate should say so rather than assert one. The important point is that the operator chose neither - the loop chose for them. ## Why nobody noticed - **A sleep cannot fail.** There is no exit code for having waited too little. - **The end state looks healthy.** By morning every node is up, every unit has its copies current again, and the count of missing records is not a metric anyone has. - **The window of exposure is transient.** The shortage existed only while the loop ran, so a dashboard sampled afterwards shows nothing. - **The loss surfaces downstream.** Someone finds a gap in a reconciled total days later, by which point the maintenance is not a suspect. ## What would have prevented it The repair is not a longer sleep. A longer sleep is the same guess with more slack, and it will be wrong the night traffic doubles or a copy has an unrelated shortage to work off. The repair is to change the kind of thing the loop waits on: 1. **Gate on the catch-up condition** - before stopping any node, require that no unit of ownership is short of its expected current copies, taking the restarted node's copies into account. 2. **Check the precondition before the first stop, too.** A cluster that was already short before maintenance began will be short at every step, and the loop must refuse to start rather than compound it. 3. **Abort, do not skip.** If the condition has not come good within a generous bound, the loop stops with the unit named. Continuing is precisely the behaviour that produced the loss. 4. **Make the safety independent of the script.** Where the platform can enforce a floor on current copies, set it, so that an operator command that would take the last current copy down is refused rather than obeyed. ## The part that generalises This failure is not really about restarts. It is about a loop whose gate is the wrong kind of fact, running long enough to exhaust a reserve that nothing was measuring. The same shape appears wherever maintenance walks a set of nodes: the individual step is safe, the sequence is not, and only a condition evaluated between steps can tell the difference. The framing to give in an interview is blunt: **a rolling restart is only a safe operation if something between the steps is checking that the previous step has been paid back.** Otherwise it is a batch restart with extra steps and a longer total outage.

  • Would a much longer sleep have been an acceptable fix?
    No, only a less frequent failure. The needed wait varies with write rate, with how much each copy missed and with the bandwidth catch-up is allowed, so any constant is wrong on some night. It also pays for safety with wall-clock time on every step, making the loop so slow that someone eventually shortens it again.
  • How would you tell after the fact whether records were actually lost?
    Compare what the writers believe they stored against what the stream holds for the affected period - a producer-side count or reconciled total against the stream's own. Cluster-side evidence is indirect but useful: the times units were short of current copies, and any record of service being restored for a unit from a copy that was not current.

saying these in an interview costs you the question

  • Blames the writers rather than the restart loop
  • Assumes acknowledged records cannot be lost by a planned change
  • Says a longer sleep would have made the script correct
  • Thinks the loop was safe because no step reported an error
  • Believes redundancy is restored the moment the process returns