A broker cluster is sized exactly to its measured peak rate - what routinely arrives that the machines have no room for, and what should the design target have been?
answer
- a peak is measured on a good day
- replay, retries, a dead node's share
- they arrive together, not in turn
- one of n nodes is 1/(n-1) each
- target a fraction, not the whole
basics
~20 sThree things reliably arrive on top of measured traffic: a reader replaying from the start, a writer retry surge re-sending records already accepted, and a dead node's share landing on its surviving peers. The design target is a fraction of what the machines can do, not all of it.
solid answer
~50 sA measured peak is measured with every node alive, nothing replaying and nothing retrying, so sizing to it exactly buys a cluster with no answer for the day those conditions stop holding. Three claimants reliably arrive, often together: a reader replaying a retention span from the start, which is read demand the live traffic never showed; a retry surge re-sending records the cluster already accepted, which is ingress arriving twice; and a lost node, whose traffic and whichever partitions it was leading land on the remaining nodes. On a cluster of `n` evenly loaded nodes, losing one adds about `1/(n-1)` to each survivor - a fifth on six nodes, half on three. So the design target is expressed as a fraction of what a node sustains at peak, chosen so that the cluster still has room after the failures it claims to survive. The reserve is not idle money; it is the part of the machine the failure case is paid for with.
go deeper
Recall that a cluster should not be sized to run flat out at its busiest moment, because a machine can be lost and the rest have to carry its work. Naming one thing that arrives unplanned is enough here.
Name the claimants and say why they arrive together: a replay, a retry surge, and a dead node's share. Explain that the design target is a fraction of what the machines sustain, applied to the peak number.
Do the failure arithmetic out loud. Show what losing one of n nodes does to the survivors' utilisation, and describe the feedback loop where slowness causes retries which cause more slowness.
Set the rule and defend its price. Decide whether the estate survives a node or a domain, state the target as a derivation rather than a quoted percentage, and record what the number deliberately does not cover.
## Why a cluster at one hundred percent of its measured peak has no answer Every measurement is taken under conditions that will not always hold. A peak rate is observed with the full node count present, every reader keeping up, and writers sending each record once. Size to that number exactly and the cluster is correct on precisely the days nothing goes wrong - which is not what it is for. The **spare share** is capacity deliberately left unused at peak so that the cluster has somewhere to put the load that arrives when something does go wrong. It is a sizing decision, not an operational one: it is spent by dividing the derived demand by a target utilisation rather than by one. ## The three claimants that reliably arrive These are not the only unplanned loads a cluster sees, but they are the three that show up often enough that a sizing write-up which ignores them is incomplete: 1. **A replay.** Where records are retained and can be re-read, a consumer that resets and reads a retention span from the start generates read demand unrelated to the write rate, served largely from the volume rather than from memory holding recent records. On platforms that delete on acknowledgement there is nothing to replay - which is worth stating, because it changes the reserve those platforms need. 2. **A retry surge.** When writes slow or fail, writers re-send. The re-sent bytes cross the network and reach the accepting node whether or not anything later recognises them as duplicates, so they are real ingress. This is the claimant with a feedback loop attached: slowness causes retries, retries cause slowness. 3. **A lost node's share.** Its clients reconnect elsewhere, its traffic moves, and whichever partitions it was accepting writes for are taken over by other nodes. Nothing was added to the cluster's total work - it is simply spread over fewer machines. The reason they belong in one answer is that they **co-occur**. A node dies, so writes to it fail, so writers retry, so a consumer that fell behind during the disruption replays. The reserve has to cover the combination, not each in turn. ## The arithmetic of a lost node On a cluster of `n` evenly loaded nodes, losing one spreads its share over the remainder, so each survivor takes roughly `1/(n-1)` more than before: | nodes | extra load per survivor after one loss | utilisation that becomes | |---|---|---| | 3 | +50% | 60% becomes 90% | | 4 | +33% | 60% becomes 80% | | 6 | +20% | 60% becomes 72% | | 10 | +11% | 60% becomes 67% | Two things fall straight out of that table. **Small clusters need a much larger reserve than large ones** for the same survivability, because one node is a larger fraction of them. And if the failure unit is a whole failure domain rather than a single node, the sum is over everything in that domain at once, which for a three-domain cluster means losing a third of the machines, not one of them. ## What the design target looks like The target is stated as a utilisation at peak, derived rather than inherited: - start from the demand derived from measurement, already scaled by the peak-to-average ratio; - decide what must be survived while still at peak - one node, one failure domain, or nothing; - add the replay and retry allowance the platform's shape makes possible; - divide the demand by the resulting target fraction to get the node count; - write down what the number excludes, so the exclusion is a decision rather than a discovery. A specific percentage quoted from anywhere is worth much less than the derivation; the same target that is prudent on ten nodes is reckless on three. What does not change is the direction: the reserve exists so that the worst interval and the worst day can be the same day. ## What the spare share is not - It is **not growth headroom.** Coping with more traffic next year is a separate exercise on a separate timescale; the spare share is for load that arrives this afternoon with no notice. - It is **not a per-client allowance.** Limiting what one writer may take is a different mechanism with a different purpose; the reserve is about the machine, not about fairness between tenants. - It is **not free of cost**, and pretending otherwise loses the argument. It is bought deliberately, and the honest framing is what it buys: the difference between a degraded interval and an outage. ## The failure it prevents The shape of the incident is always the same. A cluster runs comfortably at its measured peak for months. A node is lost during the busiest interval. The survivors take its share and cross their line. Writes slow, writers retry, the retries add ingress, a consumer falls behind and later replays over the volumes, and the cluster that was exactly the right size is now the wrong size in every resource at once. Nothing in that chain is exotic; all of it was foreseeable from the sizing sheet.
- Why does a three-node cluster need a bigger reserve than a ten-node one?Because one node is a third of it. Losing one of three adds about half again to each survivor, so a cluster running at 60% of what its machines sustain lands near 90%. Losing one of ten adds about a ninth, landing near 67%. The same survivability claim therefore demands a far lower peak utilisation on small clusters.
- Does a retry surge really add load if duplicates are eventually recognised?Yes. Duplicate handling, where a platform offers it, works after the bytes have arrived: the re-sent record still crossed the network, still consumed request handling and still competed with live traffic. It may avoid storing the record twice, which protects the volume, but it does nothing for the resources the surge actually saturates.
- How do you justify the reserve to someone who calls it idle hardware?By naming what it is bought for and what the alternative costs. State the failure the cluster claims to survive, show the arithmetic of a lost node at peak, and price the degraded interval against the outage. If the organisation would rather accept the outage, that is a legitimate answer - it just has to be an answer, not an oversight.
- Should the reserve be sized for a lost node or a lost failure domain?Whichever the cluster claims to survive, and the two differ enormously. A single-node reserve on a three-domain cluster does not cover a domain loss, which removes a third of the machines at once. Pick the unit deliberately, write it down, and make sure the copies and the coordination role are placed consistently with it.
A fire escape is not sized for how many people use the stairs on an ordinary Tuesday. It is sized for the one day everybody leaves at once - and nobody calls the unused width on Tuesday a waste.
saying these in an interview costs you the question
- Calls reserved machine waste because the cluster never exceeded its measured peak
- Assumes a measured peak already includes a replay or a failure
- Sizes for one lost node on a cluster that must survive a lost failure domain
- Treats a retry surge as free because duplicates may be recognised later
- Uses the same utilisation target on a three-node and a thirty-node cluster
- Confuses the reserve with next year's growth headroom