skip to content

You must restart every node of a six-node in-memory tier with no node promoted in its place. How do you choose the order and pacing so the system of record behind it survives?

level: seniorimportance: should knowfreq 44%

answer

  1. name the model before the plan
  2. the two knobs are order and pacing
  3. cheapest share first, measure it
  4. pace on load behind, not process liveness
  5. budget load, fill and settle per node

basics

~20 s

Name what one node's restart removes - a share of the keyspace, or a share of capacity. Then one node at a time, cheapest share first, and no next node until the load behind the tier is back inside its headroom.

solid answer

~40 s

Start by naming the model, because the answer differs. If the keyspace is split across the six, restarting one removes roughly a sixth of the entries and sends that share's reads to the system of record behind the tier. If each node holds the whole keyspace, restarting one removes no entries but concentrates traffic on the remaining five. Then set the two knobs. Order: cheapest share first, so the first cycle tests your estimates where being wrong is cheapest. Pacing: the next node does not go until the load behind the tier is back inside its headroom - not merely until the previous process is up. Budget the unavailable time honestly: a node restoring from a copy or replaying a write log is not serving while it does that.

go deeper

for a junior

Know that restarting nodes one at a time is safer than all at once, and that a node that has just come back is holding nothing yet.

for a middle

Explain why the pause between nodes exists: the restarted node's share of reads is reaching the system of record behind the tier until it is filled again.

for a senior

Show the whole plan - name the model first, order by cheapest share, pace against a measured signal behind the tier, and budget the load, fill and settle phases for each of the six cycles.

for a principal

Weigh the cycle against the release cadence: if a safe rolling restart takes an afternoon, either deployments get rarer, the tier gets cheaper to restart, or the system behind gets permanent headroom.

## First, say what one node's restart actually removes A rolling restart is six restarts in a row, and what each one costs depends on a fact about the tier that has to be stated before anything else: - **The keyspace is split across the nodes.** Restarting one removes roughly its share of the entries. Reads for that share now reach **the system of record behind the tier**, so the spike behind it is roughly the share of the arrival rate that node was serving, multiplied by one over the miss rate. Six restarts means six of those. - **Each node holds the whole keyspace and traffic is spread across them.** Restarting one removes no entries from the tier as a whole. The reads it was serving go to its peers, so what you are managing is concentration inside the tier - remaining capacity and per-node ceiling - not load behind it. The rest of the plan follows from which of those you are in. This question also assumes nobody is promoted in a restarted node's place; arranging for a peer to take over so the restart is invisible is failover, a different subject with its own failure modes. ## The two knobs There are exactly two, and naming them is half the answer: - **Order** - which node goes first. - **Pacing** - how long you wait before the next one. ### Order 1. Rank the nodes by the share of reads they serve, or the share of keyspace they hold. 2. Start with the **smallest share**. The first cycle is where your estimates get tested, and you want them tested where being wrong is cheapest. 3. Measure that cycle end to end: how long the node was unavailable, how long it took before its share of reads stopped reaching the system behind, how high the load behind the tier actually went. 4. Use those measurements to decide whether the plan survives the biggest node, and revise before you get there rather than after. ### Pacing Pace against an **observable behind the tier**, not against a stopwatch and not against process liveness. A node whose process is up but holding nothing is still sending its whole share of reads behind the tier; it is the worst moment of the cycle, not the end of it. The condition for starting the next node is: the previous node is serving, its share of reads is largely being answered from memory again, and the load on the system behind has returned inside its normal headroom. If you cannot see that last quantity, you cannot pace a rolling restart - you can only hope. ## Budget the unavailable window honestly A node is not instantly back. Depending on the tier's **posture**, coming back may include **restoring from a copy** or **replay at start** of a **write log**, and in many designs the process does not accept callers until that finishes. So each cycle is: | phase | what is happening | who feels it | |---|---|---| | stop | the node leaves service | peers, or the system behind | | load | restoring from a copy, or replay at start | nobody is served by this node | | fill | entries accumulate, or a pre-load runs | the system of record behind the tier | | settle | load behind the tier returns to normal | nobody, if you waited | Multiply that by six. The common error is to budget only the stop-and-start and call it a quick cycle, then discover the tier spent the afternoon partially empty. On a store that keeps nothing the load phase vanishes but the fill phase is guaranteed, so pacing matters **more** there, not less. ## The levers that make the cycle cheaper - **Hold the node out of rotation while it fills**, if the routing in front of the tier allows it, and **pre-load a chosen key set** before letting callers reach it. Then the fill phase costs scheduled load instead of arrival-driven load. - **Schedule the whole cycle against the traffic trough.** The multiplier behind the tier is the same at any hour, but the absolute spike scales with the arrival rate. - **Do the arithmetic before the first node**: the share of reads one node serves, times one over the miss rate, against the spare capacity of the system behind. If a single node's share does not fit, no ordering saves you and the answer is pre-loading, a longer window, or not restarting today. - **One node at a time** unless that arithmetic leaves room to spare. Two at once doubles the spike and halves the time you have to notice it. ## What this question is not It is not the miss path. Per-request protection of the system behind - coalescing duplicate work, locking a key while it is rebuilt, short-lived markers for absent entries - is somebody else's design and does not change the order or the pacing. It is also not failover: nothing here is promoted, and every node comes back as itself.

  • What is the wrong signal to pace a rolling restart against?
    Process liveness. A node that has just started is serving nothing from memory, so its whole share of reads is reaching the system of record behind the tier - that is the peak of the cycle, not its end. Pace against the load behind the tier returning inside its headroom, or against the node answering its share from memory again.
  • Does a tier where every node holds the whole keyspace make a rolling restart free?
    No, it moves the cost. Entries are not lost from the tier, so the system behind is largely spared, but the restarted node's traffic concentrates on its peers. The question becomes whether the remaining nodes can carry it at the ceiling that matters for them, and whether the returning node is sent traffic before it holds anything.
  • The arithmetic says one node's share does not fit behind the tier. What now?
    Ordering cannot save you, so change one of the inputs: pre-load the node before it takes traffic so its share never reaches the system behind, move the cycle to a deep trough where the absolute rate is lower, add temporary headroom behind the tier, or postpone. Restarting anyway and watching is how a maintenance window becomes an incident.

saying these in an interview costs you the question

  • Restarts the next node as soon as the previous process is up.
  • Ignores how long a node is unavailable while it loads or replays.
  • Never says whether nodes hold slices or whole copies of the keyspace.
  • Answers with per-request miss-path tricks instead of order and pacing.
  • Restarts several nodes at once to shorten the maintenance window.
  • Assumes a peer will be promoted, which turns this into failover.