skip to content

A full capacity ramp test takes hours, so it cannot run on every commit. How would you catch a capacity regression — a change that quietly cuts maximum sustainable throughput by 20% — before it reaches production?

level: seniorimportance: nice to knowfreq 32%

answer

  1. you cannot ramp on every commit
  2. fixed rate, compare to a baseline
  3. cost per request drifts first
  4. a flaky gate gets overridden

basics

~20 s

Stop measuring the maximum. Run a short test at a fixed request rate well below the limit and compare resource cost per request and latency against a stored baseline. Cost per request drifts detectably long before the ceiling moves, and it needs only minutes of stable load.

solid answer

~50 s

Measuring maximum throughput is the wrong signal for a per-change gate: it takes a long ramp and is the noisiest number in the whole test. Instead I fix the offered rate at, say, 50% of the known limit, run for a few minutes after warm-up, and compare against a stored baseline on cheaper, lower-variance metrics — CPU-seconds per request, allocations or garbage collected per request, database queries per request, and p99 latency at that fixed rate. A 20% capacity loss shows up as roughly a 25% rise in cost per request, which is far outside normal run-to-run noise if the harness is controlled. Query count per request is especially valuable because it catches the classic N+1 regression deterministically, with no timing noise at all. Then decide the policy: gate the release train on it, keep the full ramp weekly, and track requests-per-core in production as the long-run backstop. I would not hard-block every commit on a benchmark that flakes, because people learn to re-run until green.

go deeper

for a junior

Know that performance can regress silently and that comparing a new build against a recorded baseline under identical conditions is how you notice, rather than waiting for production to slow down.

for a middle

Explain why cost per request detects a capacity change with a shorter and less noisy run than measuring the ceiling, and name concrete per-request counters — CPU time, allocations, downstream queries, bytes — you would track.

for a senior

Show that you have thought about variance control and about what the gate does to people: which metrics are stable enough to block on, which only warn, and how the report is made diagnosable enough that failures get fixed rather than re-run.

for a principal

Own the layered strategy and its cost: per-change counters, a full ramp on a slower cadence, and production efficiency per release, plus who pays for dedicated benchmark hardware and what the organisation does when the number regresses on a deadline.

## Why maximum throughput is the wrong metric for a gate The maximum sustainable rate is the number you publish, but it is an awful regression detector. Finding it needs a long ramp plus a confirming hold; it sits at the noisiest point on the curve, where small perturbations produce large swings; and it depends on the whole environment being stable. Running that on every change is both slow and unreliable. The insight is that capacity is a derived quantity. Throughput ceiling ≈ available resource ÷ resource consumed per request. You do not need to find the ceiling to detect that it moved — you can measure the denominator directly, cheaply, at low load, where variance is small. ## Measure cost per request at a fixed rate Hold a fixed, modest offered rate — commonly 30–60% of the known limit — for a few minutes after discarding warm-up, and record: - **CPU-seconds per request** (total CPU time consumed ÷ requests completed). The single best proxy for capacity. A 20% capacity loss appears as roughly a 25% rise here. - **Allocated bytes or GC work per request**, which catches allocation regressions before they show up as latency. - **Downstream calls and database queries per request.** Deterministic, timing-free, and it catches the classic N+1 query regression exactly. - **Bytes in and out per request**, catching payload bloat such as a newly added field on a hot response. - **p99 latency at that fixed rate**, which is noisier but is the user-facing check. The first four are counters, not timings, so their run-to-run variance is a fraction of what a ceiling measurement suffers. That is what makes a short run statistically meaningful. ## Control the noise, or the gate is worthless A performance gate lives or dies on run-to-run variance. The controls that matter: dedicated, consistently-specified hardware rather than shared burstable instances; the same dataset restored to the same state each run; identical duration and warm-up discard; pinned dependency versions; and comparison of distributions from several repetitions rather than one run against one run. Establish the baseline's own noise band empirically — run the unchanged build ten times and see how much it varies — and set the threshold outside that band. A threshold tighter than the noise floor generates false failures; a threshold much wider than the real regression you care about generates false confidence. ## Decide the policy deliberately This is the part that separates a real answer from a tooling answer, and it is a decision with a cost: - **Hard block on every commit** catches regressions earliest, but a gate that fails one run in five trains engineers to re-run or override, after which it protects nothing and still costs everyone time. Only block on the metrics whose false-positive rate you have actually measured — the deterministic counters are good candidates; a p99 timing comparison usually is not. - **Warn on commit, block on the release train** is the common compromise: cheap counters run per change and post a comparison on the change itself, while the full ramp runs nightly or per release candidate where a longer, more controlled run is affordable. - **Track it in production.** Requests per core, or cost per million requests, plotted per release, is the ultimate backstop — real traffic, real data, no harness noise. It catches what the lab missed, at the price of catching it after deployment. A progressive rollout makes the exposure survivable while that comparison is being made. ## Make the result diagnosable A gate that says "7% slower" and nothing else gets ignored. Attach the breakdown that turns a failure into an action: the per-request counter that moved, per-endpoint deltas rather than an aggregate, and ideally a profile from the run. When the report says "queries per checkout request went from 4 to 19", the fix takes minutes; when it says "throughput down 6%", it takes a day and usually never happens. ## Do not forget the ceiling entirely Cost-per-request drift is a proxy and it misses regressions that only appear near saturation: a new lock that only contends at high concurrency, a pool that is now too small, a retry added on a path that only fails under load. Keep the full ramp on a slower cadence — weekly, or before any release that changes concurrency, pooling, or a dependency — and re-publish the sustainable-rate figure from it so that sizing decisions stay anchored to a measured number. ## What weak answers sound like "Run the load test in CI" with no acknowledgement of duration or variance. Comparing a single run against a single baseline run. Using a shared runner and blaming flakiness on the tool. And treating a green performance gate as proof of headroom when the gate only ever exercised half the endpoints.

  • Why is CPU-seconds per request a better per-change signal than maximum throughput?
    Because it is the denominator of capacity and it is measurable at low load, where variance is small. Finding the ceiling requires a long ramp at the noisiest point on the curve; cost per request needs a few minutes at a fixed rate. A 20% capacity loss shows as about a 25% rise in cost per request, comfortably outside a controlled harness's noise band.
  • What is the risk of hard-blocking every commit on a performance benchmark?
    If its false-positive rate is meaningful, engineers learn to re-run until green or to override, and the gate stops protecting anything while still costing time. Block only on metrics whose variance you have measured — deterministic counters like queries per request are good candidates — and let noisier timing comparisons warn rather than fail.
  • What class of regression will a fixed-rate benchmark at half load miss entirely?
    Anything that only manifests near saturation: lock contention that appears at high concurrency, a connection or thread pool that is now undersized, retries added to a path that only fails under load, and queue growth. That is why the full ramp still runs on a slower cadence and why production efficiency per release is tracked as the final backstop.

saying these in an interview costs you the question

  • Tries to measure maximum throughput on every commit
  • Compares one run against one baseline run
  • Runs the benchmark on shared burstable infrastructure
  • Sets a threshold tighter than the measured noise band
  • Reports only an aggregate percentage with no breakdown

context