skip to content

A load test measures a long-lived request server during its first thirty seconds - why does that figure understate steady-state throughput?

level: seniorimportance: must knowfreq 58%

answer

  1. the first requests are not the system
  2. thresholds have not been crossed yet
  3. compilation takes the same cores
  4. measure after the plateau, not before
  5. report startup and peak separately

basics

~20 s

Early requests execute interpreted or lightly optimised code while counters climb, and compilation itself takes CPU away from request handling. Peak arrives only after the hot paths have been compiled, so an early measurement times warmup rather than the system that was shipped.

solid answer

~40 s

A runtime that compiles while it runs does not start at its final speed. For the first requests the hot paths are still interpreted, then briefly running quickly-compiled code, while profile data accumulates and compiler threads consume cores the request handlers also want. Throughput climbs as those paths reach the optimising tier and then flattens. Measuring during the climb reports a mixture of interpreted execution, partially optimised execution and compiler overhead - a number that describes a transient state. The honest procedure is to drive representative traffic until per-request time stops improving, discard everything up to that plateau, and measure a long window afterwards. Report startup and peak as two separate numbers, because they answer two different questions and no single figure is true of both.

go deeper

for a junior

Recall that a program which compiles itself while running is slower at the start, so the very first measurements are not what the system will do once it has been running a while.

for a middle

Explain the named causes of the ramp - low tiers still executing, a thin profile, compiler threads taking CPU - and describe the warm, discard, measure procedure that produces an honest steady-state figure.

for a senior

Demonstrate that you read the curve rather than an average: spotting a fall-back-and-recompile dip, ruling out non-runtime causes, sizing the measurement window, and knowing when the ramp itself is the number the deployment cares about.

for a principal

Own what the benchmark is for. Decide which figure the business is buying - time to first useful work, peak capacity, or capacity-under-scale-out - and make the measurement regime and the deployment strategy answer that same question.

## What warmup is **Warmup** is the interval between a process starting and its performance reaching a stable plateau. In a runtime that compiles while it runs, it is not a vague settling effect; it has specific causes that can be named and measured: - **Hot paths are still in a low tier.** The counters that trigger compilation have not yet crossed their thresholds, so the code doing the work is interpreted or only cheaply compiled. - **The profile is still thin.** Aggressive optimisation needs observations - which branches run, which types appear - and those accumulate only by executing. - **Compilation competes for CPU.** Compiler work happens inside the same process on the same cores. During the ramp, the machine is doing two jobs, and request handling gets less than all of it. - **First-execution costs are being paid.** Code is being located and prepared, lookup structures are being filled, caches at every level are cold. These are separate from compilation but land in the same window. ## Why the early figure is not the system you shipped A long-lived server spends essentially all of its life in the plateau. A number taken during the ramp describes a state the system passes through once, and it is wrong in a specific direction: it is pessimistic about steady state and it hides the shape of the curve. Two systems with identical peak throughput can have very different warmup profiles, and a single early average tells you nothing about which you have. It also mixes two effects that should be reported separately. Some of the deficit is code quality, which compilation fixes; some is compiler threads taking cores, which is a resource cost during the ramp only. Averaging them into one figure loses the distinction that would tell you what to do about it. ## Measuring honestly 1. **Drive representative traffic.** The profile is built from what actually executes, so warming with one code path and then measuring another means measuring a cold path again. Include the mix of request shapes that production sees. 2. **Warm until the curve flattens.** Track per-request time or throughput over time and wait until improvement stops - do not warm for a fixed count picked by habit. The plateau is an observation, not a constant. 3. **Discard the ramp.** Everything up to the plateau is excluded from the reported figure rather than averaged into it. 4. **Measure a long window afterwards.** Long enough that occasional recompilation, rebalancing and background work are represented rather than accidentally excluded. 5. **Repeat in fresh processes.** Compilation decisions vary between runs, so a single process gives no sense of variance. 6. **Report both numbers.** Time-to-plateau and steady-state throughput, side by side, with the shape of the ramp. ## Reading the curve | Observation | Most likely reading | |---|---| | Per-request time falls steadily, then flattens | ordinary warmup: hot paths reaching higher tiers | | Falls, flattens, then worsens and recovers | an assumption stopped holding; code fell back and was recompiled | | Never flattens across a long run | the workload keeps reaching new code, or thresholds are never crossed on the paths that matter | | Flattens immediately at a poor figure | the bottleneck is not code speed - look outside the runtime | ## When warmup is the number that matters The plateau is the right figure for a process that lives for days. It is the wrong figure whenever processes are short-lived or frequently replaced: - A job that starts, processes a batch and exits within seconds may never leave the ramp at all; its **wall-clock time including the interpreted phase** is the only meaningful measure. - A service redeployed many times a day pays the ramp on every instance, every time. - A fleet that scales out under a traffic spike adds instances precisely when it is least able to absorb their warmup, and a new instance given a full share of traffic immediately will show worse latency than its peers. A sound benchmark answers the question the deployment actually asks. For a long-lived server that is the plateau; for a short-lived job it is the total including the ramp; for an elastic fleet it is both, plus how long an instance needs before it deserves full traffic. ## The trap in the other direction Warming is not licence to measure something the system never does. If the warm-up phase runs a narrower set of paths than production, the profile that the optimising tier built is a profile of the benchmark. The measured plateau is then real for the benchmark and optimistic for production, which is the same error as the early measurement with its sign flipped.

  • The warmed figure was stable for ten minutes and then dropped sharply. What are the likely causes?
    Most often an assumption the optimising tier relied on stopped holding - new input drove a path that had never run, or a call site saw a type it had not seen - so the specialised code was abandoned and the method ran slower until it was recompiled. Traffic-shape changes, background work and resource contention produce the same signature and have to be ruled out.
  • How do you decide, from measurements alone, that warmup has finished?
    Plot per-request time or throughput against elapsed time rather than averaging. Warmup shows as monotone improvement that decelerates and then stops; treat the plateau as reached when the trend is flat within run-to-run noise over a window several times longer than the sampling interval. A fixed iteration count is not evidence, because the right count depends on the workload.
  • A fleet scales out during a traffic spike. What does warmup cost you there?
    Each new instance starts on the ramp, so it serves its first traffic at reduced capacity exactly when the fleet is short of it. Sending a new instance a full share immediately shows up as a latency spike attributed to the wrong cause. The usual answers are ramping traffic to an instance gradually, starting instances before the load arrives, or reducing how much speed depends on in-process warmup.

saying these in an interview costs you the question

  • Reports the first run's timing as the system's throughput
  • Thinks warmup is only the cost of loading code from storage
  • Believes a warmed process never falls back to slower code
  • Assumes discarding one iteration constitutes warmup
  • Warms with one code path and then measures a different one
  • Quotes a single figure for both startup and peak