skip to content

Your cluster is saturated at peak and you add two nodes to relieve it — why does the incident get worse before it gets better?

level: seniorimportance: should knowfreq 57%

answer

  1. capacity spends before it pays
  2. the fill reads off the saturated nodes
  3. empty nodes receive no traffic
  4. levers ranked by latency of effect
  5. detached storage inverts this

basics

~20 s

Because relief has to be copied onto the new nodes, and that copying reads records off the very nodes that are already at their ceiling. The cluster does extra work for hours and gains nothing until the new copies are current.

solid answer

~50 s

Adding nodes buys capacity on a delay and charges for it up front. The newcomers hold nothing, so they absorb no traffic; getting them work means filling their copies from the existing nodes, over the same disks and links that are already saturated. Until a copy becomes current enough to serve, the cluster is carrying production traffic plus copy traffic on unchanged hardware — measurably worse. Relief then arrives unit by unit, in the order copies become current, typically long after the incident ended. The levers that act inside the incident are the ones that reduce or shed the work arriving, not the ones that add machines. Expanding during an incident is still right when the shortfall will outlast it — but then it is a recovery action with a paced copy phase, not the fix for tonight, and it should be said that way on the call.

go deeper

for a junior

The thing to hold on to is that new machines start empty, so they cannot take any traffic yet, and getting data onto them makes the busy nodes busier for a while.

for a middle

Explain where the extra load comes from — reads and network out of the already-saturated nodes, plus cache displacement — and why clients send nothing to a node that owns nothing.

for a senior

Demonstrate the call itself: rank the levers by how fast they act, keep hunting for a real mitigation, start the expansion only if the shortfall outlasts the night, and pace the fill.

for a principal

The question you own is why the estate had no headroom at peak, and whether a storage design where capacity can be added in seconds is worth its other costs.

## The shape of the mistake A cluster is at its ceiling. Someone adds machines, because machines are what you add. Twenty minutes later the numbers are worse, and the reasonable-sounding conclusion — "the expansion did not help" — is wrong in an important way. The expansion *is* helping; it just spends before it pays, and the spending lands inside the incident. The reason is the subject of this leaf: a node that joins holds nothing. It is a member, not a participant. Every unit of ownership it will eventually serve — every partition (or queue) — has to be assigned to it and then filled, and the filling is the problem. ## Where the extra load comes from During the fill, the cluster is doing everything it was doing before, plus: - **Reads on the busy nodes.** The records the newcomers need are fetched from the nodes that currently hold them — the saturated ones. Those reads land on the same disks that are serving production. - **Network out of the busy nodes.** The copy traffic leaves through the same links already carrying writes and reads. - **Cache displacement.** The copying walks over old records, which pushes the recently written records — the ones readers actually want — out of whatever memory was caching them, so ordinary reads get more expensive too. - **Nothing on the credit side yet.** A copy that is still filling cannot serve, so none of this buys any throughput until it is current. Clients do not spread requests over the new nodes in the meantime, either. Traffic follows ownership: a client sends each request to whichever node owns that unit, so an empty newcomer receives nothing at all until ownership moves. ## What actually helps inside the incident | Lever | Time to effect | What it costs | |---|---|---| | Reduce or shed the work arriving | immediate | the work you shed | | Move demanding readers or writers off the hot units | minutes | coordination with their owners | | Free headroom on the existing nodes | minutes | whatever you gave up to free it | | Add nodes and fill them | hours | more load first, then relief | The ranking is not a rule about machines being useless. It is a statement about *latency of effect*. During an incident you are choosing among levers by how fast they act, and adding capacity is the slowest lever available while also being the only one that raises the ceiling permanently. ## When expanding mid-incident is still the right call Three conditions, and it needs all three: 1. **The shortfall outlasts the incident.** If the cluster is short of capacity in a way that will still be true tomorrow, starting the expansion now is simply starting it at the earliest honest moment. 2. **The copy phase is paced.** The fill has to be allowed to take longer so that it does not deepen the incident. How that pacing is expressed differs by platform and is a decision of its own. 3. **Everyone on the call knows relief is not coming tonight.** The failure mode that hurts is not the expansion; it is an incident commander who stops looking for a real mitigation because machines are on the way. There is one shape where this reasoning inverts. Where records live on shared or remote storage, ownership can move to a newcomer without copying anything, so expansion is a control-plane change of seconds and it *is* a legitimate mitigation inside an incident. Knowing which shape your cluster is before the call is what makes the answer to "should we add nodes?" fast. ## Two adjacent traps - **Expanding to fix something that is not a capacity problem.** If the cause is one demanding tenant, a slow disk on one node, or work concentrated onto a few nodes, more machines dilute nothing; the concentration simply reappears. - **Adding many nodes at once.** Each newcomer must be filled, and filling several simultaneously multiplies the copy traffic the saturated nodes must produce. A smaller expansion completes sooner and hurts less on the way. The sentence to be able to say on a call: *the machines are useful, they are not a mitigation, and here is what we are doing for the next ninety minutes instead.*

  • When is expanding during a live incident still the right call?
    When the shortfall will outlast the incident, so the expansion has to happen anyway and starting now is honest; when the copy phase can be paced so it does not deepen the event; and when everyone understands relief lands later. It is a recovery action taken early, not a mitigation.
  • The cluster is saturated because one tenant is heavy — does adding nodes help there?
    Barely. Extra machines dilute a cluster-wide shortage, not a concentration: traffic follows ownership, so if the demand sits on a few units of ownership it stays on whichever nodes hold those units. You get the copy traffic now and very little relief later.
  • How do you tell mid-incident whether the expansion is progressing or stuck?
    Watch how far behind their leaders the newcomers' copies still are, and whether that distance is falling at a rate that would finish in a sensible time. A distance that is flat or growing means the fill is losing to production traffic, and the expansion will not complete without something else giving way.

saying these in an interview costs you the question

  • Treats adding machines as the first mitigation in a saturation incident
  • Expects new nodes to take a share of traffic on arrival
  • Concludes the expansion failed when the numbers get worse first
  • Adds many nodes at once and multiplies the copy traffic
  • Uses expansion against load concentrated on a few units of ownership