skip to content

A cluster sized from a daily average of 8,000 records per second sees 40,000 in its busiest ten minutes - which number should have sized it?

level: middleimportance: should knowfreq 55%

answer

  1. averages do not saturate anything
  2. busiest interval over sustained rate
  3. machines take the peak
  4. storage integrates the average
  5. match the interval to what gives way

basics

~20 s

The busiest interval, not the daily mean. A peak-to-average ratio of five means the machines must carry 40,000 records per second, while the volume size still comes from the sustained rate accumulated over the retention span.

solid answer

~40 s

Split the question by resource. Anything that saturates instantaneously - network interfaces, request handling, the write path - is sized by the busiest interval, so here by 40,000 records per second, a peak-to-average ratio of five against the 8,000 mean. Anything that accumulates is sized by the sustained rate instead: stored bytes are the average rate integrated over the retention span, and a five-times burst lasting ten minutes barely moves that total. The trap is the averaging interval itself: a mean taken over a day, or even over a minute, can show plenty of room while writers are already timing out inside it. Average over the interval at which the resource actually gives way, and take the maximum of those intervals, not their mean.

go deeper

for a junior

Recall that traffic is not flat, and that a cluster has to carry its busiest minutes rather than its daily average. Saying which of the two numbers you would size from already answers most of this.

for a middle

Compute the ratio and use it: peak divided by sustained. Explain that machines take the peak while stored bytes take the sustained rate over the retention span, and say why the two differ.

for a senior

Interrogate the measurement itself. Name the averaging interval, take a maximum rather than a mean, separate write and read peaks, and identify when a peak is really a retry surge rather than demand.

for a principal

Decide which peaks the organisation pays to survive and which it accepts degrading through, and write that down. An unstated peak assumption is how a cluster comes to fail on the one day its owners most needed it.

## What the ratio is The **peak-to-average ratio** is the multiplier between the rate a cluster sustains over a representative interval and the rate it carries in its busiest one. In the example it is: ``` sustained rate 8,000 records/s busiest ten minutes 40,000 records/s peak-to-average 40,000 / 8,000 = 5x ``` It is the single number that separates a cluster that runs from a cluster that runs on average. Most real traffic has one: business hours against nights, a batch job that opens at the top of the hour, a mobile client population that wakes on a push, a retry surge after a dependency recovers. ## Why the average cannot size the thing that fails Saturation is not an average phenomenon. An interface that can move 1 GB/s does not move 5 GB for one second and catch up later; the excess queues, and when the queue is full the excess is refused, dropped or delayed. A cluster sized at the mean is, by construction, over its line for as long as the busiest interval lasts. What that looks like from outside **differs by platform**, and saying so is part of the answer: - on some platforms the writes are accepted but acknowledged more slowly, so the visible symptom is climbing write latency and writers holding records in memory; - on others the request is refused outright with an error and the writer must retry, which adds its own traffic to the interval that is already the worst one; - on a rented cluster the ceiling may be enforced as a policy rather than reached as a physical limit, and the surplus is refused or slowed by rule. In every case the second-order effect is the same shape: **writers that cannot hand records over start buffering, and a buffer that fills either blocks the application or discards.** The incident is rarely reported as "the cluster was busy"; it is reported as the application being slow. ## Which numbers take the peak and which take the average This is the distinction most sizing exercises get wrong in one direction or the other - either everything is sized for the peak, which buys volumes nobody needs, or everything is sized for the mean, which buys a cluster that fails daily. | quantity being sized | which measurement | why | |---|---|---| | network interfaces, in and out | busiest interval | saturation is instantaneous; there is no catching up | | request handling and per-record work | busiest interval | the same records arrive in the same second | | traffic between record-serving nodes | busiest interval | it tracks ingress, so it peaks with it | | bytes stored for a retention span | sustained rate | storage integrates the rate over the whole span | | the reserve left free on the machines | applied to the peak number | the reserve exists for the worst interval, not the mean | A worked contrast: at 8,000 records per second of 2 KB, a seven-day retention span with three copies stores roughly `8,000 x 2 KB x 604,800 s x 3`, and ten minutes at five times the rate adds well under one percent to that. The same ten minutes decides the entire network and node count. ## The averaging interval is itself a decision A rate is meaningless without the interval it was averaged over, and the interval has to match what gives way: 1. **Sub-second bursts** are usually absorbed by buffers and rarely size anything on their own. 2. **Seconds to a few minutes** is where interfaces, request handling and the write path actually saturate - this is the interval to take the maximum over. 3. **Hours and days** are the wrong granularity entirely for machine sizing, though they are the right one for stored bytes. The classic misread is a dashboard averaging over a minute that shows comfortable room while writers time out every few seconds inside it. Nothing is wrong with the measurement; it is answering a different question from the one being asked. ## Getting an honest number - Take the **maximum** of the short-interval rates across a representative period, not the mean of them, and not the single highest spike of all time unless that spike must be survived. - Measure the **write rate and the read rate separately**, because their busiest intervals need not coincide - a nightly re-read can peak when writes are at their quietest, which is good news you can only see if you looked. - Record the **record rate and the byte rate** at peak, since a burst of small records and a burst of large ones stress different resources. - Expect the ratio to **change with the traffic mix**, and re-measure after a launch rather than inheriting last year's multiplier. - Treat a peak that is mostly retries as a **different problem** from a peak that is mostly new records; the first can shrink when the underlying cause is fixed.

  • Why does the volume size not follow the peak?
    Because stored bytes are a rate accumulated over the retention span, not a rate that must be met instantly. Ten minutes at five times the rate adds ten minutes of extra bytes to a span measured in days - well under a percent. Networks and request handling have to meet the peak the moment it arrives; storage only has to hold the integral.
  • A dashboard averaged over one minute shows headroom while writers time out. What is wrong?
    The averaging interval is longer than the interval over which the resource saturates. A minute containing ten seconds at four times the line rate averages out to something comfortable. Re-measure at the granularity at which the interface, the request path or the write path actually gives way, and take the maximum rather than the mean.
  • Should the peak used for sizing be the highest rate ever seen?
    Only if that event must be survived. A single unrepeatable spike sizes an expensive cluster for a day that may never return. A defensible approach is the peak of the busiest ordinary interval, stated with the event class it excludes, so the organisation chooses knowingly rather than discovering the exclusion during one.

saying these in an interview costs you the question

  • Sizes machines from a daily or hourly mean rate
  • Quotes a rate without saying what interval it was averaged over
  • Sizes the volume for the peak rate rather than the sustained one
  • Assumes the write peak and the read peak happen together
  • Treats a retry surge as ordinary new traffic when sizing
  • Believes a burst is always absorbed because buffers exist