skip to content

A checkout service is measured at 1,200 requests/second average and 3,000 requests/second at its daily peak hour. Its owning team wants to know how many servers to run if each server can safely sustain 150 requests/second. How should peak traffic and headroom, not just the average, shape that server count?

level: middleimportance: must knowfreq 80%

answer

  1. size to peak, not average
  2. peak-to-average ratio varies by domain
  3. headroom = 20-50% above measured peak
  4. node failure/deploy drains eat into 'capacity'
  5. autoscaling has reaction lag; headroom covers the gap

basics

~20 s

Never size servers off the average - size off the busiest moments plus extra room. Take the peak number, divide by what one server handles, then add a safety margin so a small spike doesn't cause an outage.

solid answer

~40 s

Divide peak QPS by per-server capacity: 3,000 / 150 = 20 servers just to meet the measured peak exactly, with zero slack. That's fragile - any traffic above the historical peak, a slow node, or a deploy taking capacity offline pushes the fleet over 100% utilization. Add headroom, commonly 20-50% above measured peak depending on volatility and blast-radius tolerance: at 30% headroom, target capacity is 3,000 x 1.3 = 3,900 QPS, or 26 servers. The exact headroom percentage is a judgment call balancing cost (idle capacity) against risk (outage during an unmodeled spike), and should be informed by how much peak traffic itself varies day to day, not just a fixed rule of thumb.

go deeper

for a junior

Should recognize, when prompted, that peak matters more than average for server sizing, and be able to divide peak QPS by per-server throughput.

for a middle

Should independently apply a headroom margin on top of measured peak without being asked, and explain in general terms why zero-margin sizing is risky.

for a senior

Should reason about where the headroom percentage should come from (traffic volatility, business cost of an outage), how autoscaling interacts with static headroom, and how node failures/deploys erode nominal capacity.

for a principal

Should connect headroom policy to broader organizational practices - pre-warming for known events, load-shedding and graceful degradation as complements to headroom, and the cost/risk trade-off as an explicit, revisitable business decision rather than a fixed engineering constant.

## Why the average is the wrong number to size on Peak-vs-average reasoning is the step that converts a raw capacity estimate into an actual provisioning decision, and it is one of the most consistently tested judgment calls in system design interviews because it separates candidates who can do arithmetic from candidates who understand what the arithmetic is for. The average request rate answers 'how much work does this system do over a full day, on average,' which is useful for cost modeling and long-run resource utilization targets, but it is close to useless for deciding how many servers must be running right now, because a server fleet has to survive the worst second of the day, not the average one. ## The mechanism The mechanism starts with identifying peak load, ideally pulled from real monitoring data rather than assumed: here, `3,000` requests/second measured at the daily peak hour, versus `1,200` average. The ratio between them, `2.5x` in this case, is the **peak-to-average ratio** discussed in QPS estimation more generally, and it varies by domain - a checkout flow tied to a fixed business day has a more pronounced peak than a service used uniformly around the clock by a globally distributed base. Dividing peak QPS by the throughput one server can sustain (`150` requests/second here, itself an estimate that should come from load testing, not a guess) gives the server count needed to exactly meet peak: `3,000 / 150 = 20` servers. ## Why sizing to the bare peak is fragile But sizing to exactly meet a historically observed peak is fragile for several concrete reasons, which is why **headroom** - deliberately over-provisioning beyond the measured peak - is standard practice, not paranoia. 1. **First, historical peaks are a sample, not a ceiling**: the next day's peak could exceed today's, especially for a growing product, a marketing campaign, or a seasonal event (a checkout service in particular should expect Black Friday-scale peaks well above a typical day). 2. **Second, fleets are not perfectly homogeneous or perfectly available at every instant** - nodes fail, get drained for deploys, or get evicted by the orchestrator for maintenance, so 'capacity' isn't the same as 'capacity currently healthy and serving.' 3. **Third, request cost itself isn't perfectly uniform**; a burst of unusually expensive requests (larger payloads, cache misses forcing database round-trips) can consume more per-server capacity than the average request implies, even at the same QPS. Headroom absorbs all three of these without requiring a human to react in real time. ## Choosing the margin The standard practice is to add a percentage margin on top of measured peak - commonly in the 20-50% range - chosen based on how volatile the traffic pattern is and how costly an outage would be. A checkout service, where an outage directly blocks revenue and often coincides with the business's own highest-value traffic (a flash sale is simultaneously a peak and a moment where failure is maximally expensive), justifies headroom toward the higher end of that range. At 30% headroom: `3,000 x 1.3 = 3,900` target QPS, requiring `3,900 / 150 ~= 26` servers - six more than the bare-peak figure, purely as insurance. This is the concrete trade-off headroom encodes: extra servers cost money and sit partially idle most of the time, but that idle capacity is what prevents cascading failure (overloaded servers slow down, which increases queueing and retry traffic, which further overloads the remaining healthy servers - a classic death spiral) during the exact moments the business cares about most. ## The failure mode, and what headroom buys The failure mode headroom guards against shows up repeatedly in real incidents: a service sized to its historical peak gets hit by a slightly larger event - a viral social post, a competitor's outage driving traffic over, a marketing email landing better than expected - and the fleet saturates, latency climbs, clients start retrying (adding more load onto an already-overloaded system), and the whole thing cascades into a full outage rather than gracefully degrading. - **Autoscaling groups mitigate but don't eliminate this risk**, because scale-up takes time (new instances need to boot, warm caches, and join load balancer rotation), so headroom functions as the buffer that covers exactly that reaction-time gap - production systems commonly combine both: static headroom sized for known daily peaks, plus autoscaling to absorb the unexpected on top of that baseline. - **E-commerce platforms sizing for Black Friday, or ticketing platforms sizing for an on-sale moment**, are the canonical real-world cases where teams explicitly multiply peak estimates by a large headroom factor and pre-warm capacity well before the event, precisely because reactive autoscaling alone is too slow for a demand spike that arrives in seconds rather than minutes.

  • How does autoscaling change the calculus around how much static headroom you need to keep provisioned at all times?
    Autoscaling reduces, but doesn't eliminate, the need for static headroom, because scaling up takes real time - new instances must boot, warm up, and be added to load-balancer rotation, often tens of seconds to a few minutes. Static headroom covers that reaction window; autoscaling then handles sustained or larger-than-expected growth beyond it, so mature systems typically use both together rather than relying on autoscaling alone.
  • What happens if per-server capacity (150 requests/second here) was estimated from a synthetic benchmark rather than real production traffic?
    Synthetic benchmarks often use simplified, uniform request payloads that don't reflect the real mix of cheap and expensive requests in production, so the true sustainable per-server rate under real traffic composition is frequently lower than the benchmark suggests. This is a common source of capacity plans that look correct on paper but saturate earlier than predicted in production.
  • Why might a service intentionally choose a lower headroom percentage than 30%, despite the added risk?
    Cost - idle headroom capacity is paid for continuously but used only during peak, so for a service where an outage is low-stakes (internal tooling, non-critical batch jobs) or traffic is highly predictable with low variance, a team may accept a thinner margin, like 10-15%, to reduce infrastructure spend, explicitly trading resilience for cost efficiency.

It's like staffing a restaurant: if you schedule cooks based on the average number of meals served across the whole day, you'll be understaffed and customers will wait an hour during the Friday dinner rush - you staff for the rush, and you keep one extra cook on call in case the rush is bigger than usual or someone calls in sick.

saying these in an interview costs you the question

  • Sizes the server fleet directly off average QPS
  • Sizes exactly to the measured historical peak with zero headroom margin
  • Can't explain why node failures or deploys reduce effective available capacity below the nominal fleet size
  • Doesn't distinguish autoscaling's reaction-time lag from the need for baseline static headroom
  • Picks a headroom percentage with no reasoning tied to traffic volatility or outage cost

context