skip to content

questions

5

In a performance run, what are ramp-up, steady state and warm-up, and which window do you measure?

level: juniorimportance: must knowfreq 62%

answer

  1. Not every second of the run counts
  2. The load climbs, holds, then falls
  3. Caches and pools still filling early
  4. Cut where the trend goes flat
  5. Report only the held-constant stretch

basics

~20 s

Ramp-up is the stretch while offered load climbs to the target level; steady state is where that load is held constant; warm-up is the early portion where caches, pools and compiled code settle. Report the steady state only.

solid answer

~50 s

A performance run has a shape over time. During **ramp-up** the generator adds load gradually towards the target rate or population, so the system is never asked for the full workload at once. During **steady state** the load is held flat long enough for queues, caches and pools to reach a stable regime; this is the only window whose numbers describe the target load. **Warm-up** is the early part of the run — often overlapping ramp-up and the first minutes of steady state — where lazily built caches fill, connections open, buffers grow and hot code paths get optimised, so latency is high and falling. Warm-up samples are excluded from the reported percentiles, and the exclusion is decided from an observed stabilisation point, not a habit. Ramp-down at the end is likewise excluded. If the run never reaches a flat stretch, it has no result to report.

code

pseudocode · 14 lines
pseudocode
schedule = [
  phase(name: "ramp",        from_rate: 0,   to_rate: 240, duration: 12.min),
  phase(name: "steady",      rate: 240,                    duration: 45.min),
  phase(name: "ramp_down",   from_rate: 240, to_rate: 0,   duration: 3.min)
]

samples  = run(schedule)
settled  = first_time_when(samples, metric: p99_per_minute, trend: flat_for(5.min))
reported = samples.where(t >= settled and t < start_of("ramp_down"))

report(p50: percentile(reported, 50),
       p99: percentile(reported, 99),
       window_start: settled,
       window_length: length(reported))

go deeper

for a junior

Be ready to name the three stretches of a run and say plainly that only the held-constant one is reported. Knowing that early samples are unrepresentative because caches and pools start empty is enough at this level.

for a middle

Explain the mechanics behind warm-up — cold caches, pools opening connections on demand, lazily built components, optimisation of hot code paths — and describe how you pick the exclusion point from the latency trend instead of copying a fixed number.

for a senior

Show the judgement of verifying steadiness from the data, sizing the window against the system's slow cycles, and treating a run that never flattens as inconclusive rather than as a result you round off and publish.

for a principal

Own the convention: how run shape, exclusion rules and the definition of a valid window are standardised across teams so results are comparable, and how you keep a short pipeline run honest when it cannot afford a full stabilisation period.

### The run has a shape, and only part of it is the result A performance run is not a single uniform block of traffic. Offered load rises, is held, and then falls, and the system under test behaves very differently in each of those stretches. Reporting one number over the whole run mixes three regimes together and produces a figure that describes no operating condition the system will ever actually be in. **Ramp-up** is the period during which the generator increases what it asks for — either a rising arrival rate or a growing population of simulated users — until the target workload is reached. Ramping matters for two reasons. First, it is closer to how real traffic arrives: production rarely jumps from nothing to peak in one instant, and a system given that jump can fail in ways that say more about the test than about production. Second, a ramp lets you watch behaviour as a function of offered load rather than at a single point, which is often where the interesting inflection shows up. **Steady state** is the flat stretch afterwards, where offered load is held constant. This is the measurement window. It has to be long enough for the system to settle into a stable regime: queues stop growing or stabilise at a length, caches reach their working set, memory reaches a plateau rather than a slope, background jobs have had a chance to interleave. Whether a stretch is genuinely steady is something you verify from the data — flat throughput, flat percentiles, no visible trend — and not something you declare because the schedule said so. **Warm-up** is the sub-stretch at the beginning where the system is still assembling itself. Connection pools open their connections on demand. Caches at every layer start empty and populate. Lazily initialised components build on first use. Runtimes that optimise hot code do so after enough executions. Storage engines have cold buffers and read from disk what they will later serve from memory. During warm-up, response times are high and falling; including those samples drags every reported statistic upward and, worse, makes the result depend on the length of the run rather than on the system. ### Choosing the exclusion, not guessing it The correct way to set the warm-up exclusion is to plot the metric over time and cut at the point where the trend goes flat, then state the cut in the report. A fixed convention copied between projects is a common error: a service whose caches fill in ninety seconds and a service that needs half an hour to reach a stable plateau cannot share a warm-up rule. If you do choose a fixed window because the pipeline needs determinism, verify periodically that the system still stabilises inside it, and treat a run that has not flattened by the end of the window as inconclusive rather than as a result. Ramp-down is excluded for the mirror-image reason: as the generator withdraws load, queues drain and latency falls for reasons that have nothing to do with the behaviour you are trying to characterise. ### A worked example An insurance quote engine is exercised by a nightly 6-hour run. The schedule is a 12-minute ramp, then a flat stretch at the target rate, then a short ramp-down. The team reports one mean and one 99th-percentile figure over the whole file of samples. Two things go wrong. The reported p99 of 2,180 ms is far worse than anything seen after the first quarter hour, because the first several thousand quote requests each triggered a cold rating-table load. And when someone shortens the run to 90 minutes for a quick check, the same system reports a materially worse p99 — not because it changed, but because the fixed warm-up cost is now a larger share of the samples. After cutting the first 14 minutes, based on where the latency trend flattened, the two runs agree with each other. A second, subtler trap on the same run: the driver stepped through a pool of pre-generated quote records by index, and an off-by-one boundary in the loop made every iteration after the first pass reuse the final record. The run looked beautifully steady because a single record stayed hot in cache the entire time. Steadiness of the graph is necessary but not sufficient — the workload itself has to still be varied when the graph goes flat. ### What to say about the phases A defensible report states the ramp shape, the steady-state window and its duration, the warm-up exclusion and the evidence for it, and the percentiles computed over the retained window only. Anything averaged across the whole file is a number about the test schedule, not about the system.

  • How long should the steady state be?
    Long enough that the behaviour you care about has room to appear and that the sample count supports the percentile you report. A high percentile computed from a few hundred requests is noise, so the window has to produce enough observations for the tail to be meaningful. It also has to outlast the system's slow cycles — cache expiry, scheduled background work, log rotation — otherwise you have measured only the quiet gaps between them.
  • Is warming the system up before measuring not cheating, since real users hit a cold system after every deploy?
    Both questions are legitimate but they are different runs. The steady-state run characterises sustained behaviour; a separate cold-start run characterises what the first minutes after a restart or deploy cost, and that one deliberately measures the warm-up. The mistake is silently blending them into one number, which answers neither question.
  • What tells you a stretch is not actually steady even though the offered load is flat?
    A metric with a visible trend rather than a level: latency percentiles still drifting, memory climbing without plateau, queue depth growing, error rate creeping. Any of those means the system is still moving toward some other regime, and the numbers you take from that stretch describe a transient. Extend the run or investigate the drift rather than averaging across it.

Timing a runner from the moment they leave the changing room gives you a slower time than the race — you start the clock at the line, not at the door.

saying these in an interview costs you the question

  • Averaging the whole file, ramp-up included, into one number
  • Applying a fixed five-minute warm-up to every system without checking
  • Jumping straight to peak load with no ramp at all
  • Calling a stretch steady while latency is still visibly trending
  • Ending the run before the system has stabilised and reporting anyway
  • Excluding warm-up but never stating the exclusion in the report

context

open as a page

What is the difference between an open and a closed workload model in a load test?

level: middleimportance: must knowfreq 54%

basics

~20 s

In an open model the generator sends requests at a chosen arrival rate regardless of how the system responds. In a closed model a fixed population of simulated users each waits for a response, then thinks, then sends again, so slowness throttles the load.

open as a page

Why can a load generator's reported p99 be far better than the latency real users would have seen?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Because a generator that waits for each response stops sending while the system is stalled, so the requests that would have been slowest are never issued and never measured. This is coordinated omission: the missing samples are exactly the bad ones.

open as a page

How do you tell whether the load generator, not the system under test, is the bottleneck in a run?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compare achieved load against offered load and instrument the generator hosts: processor use, memory pressure, port and descriptor exhaustion, network saturation. If the achieved rate falls short while the system under test sits idle, the rig is the limit.

open as a page

How would you build a workload model from real production traffic and keep it representative over time?

level: principalimportance: should knowfreq 42%

basics

~20 s

Derive the transaction mix, arrival shape and data profile from observed traffic rather than opinion, state the reference period, and keep the model a versioned, owned artefact that is re-derived on a schedule instead of ageing into fiction.

open as a page