skip to content

Why can a load generator's reported p99 be far better than the latency real users would have seen?

level: seniorimportance: should knowfreq 38%

answer

  1. The measurement pauses when things go wrong
  2. The missing samples are not random
  3. Latency from intended start, not actual send
  4. One long sample stands for many
  5. Throughput dips are the fingerprint

basics

~20 s

Because a generator that waits for each response stops sending while the system is stalled, so the requests that would have been slowest are never issued and never measured. This is coordinated omission: the missing samples are exactly the bad ones.

solid answer

~50 s

A generator whose simulated users block on each response cannot issue a request during a stall. If the system freezes for 900 ms, a real arrival stream would have delivered requests throughout that freeze, each one experiencing a wait from its own arrival moment; the blocked generator delivers none. The samples omitted are not random — they correlate exactly with the worst moments — so the resulting distribution is optimistic precisely in the tail you were trying to measure. This is **coordinated omission**. The fixes are to measure latency from each request's *intended* start time rather than from when the generator got around to sending it, to back-fill the missing observations a stall implies, or to drive the workload with an open arrival schedule that keeps issuing during a stall. Without one of those, a clean-looking high percentile is evidence about the generator, not about the system.

code

pseudocode · 14 lines
pseudocode
interval = 1 / target_rate
next_due = now()

loop:
    intended = next_due
    wait_until(intended)              // may already be in the past after a stall
    sent     = now()
    response = send(request)
    done     = now()

    record(biased_latency:   done - sent)       // hides the backlog
    record(honest_latency:   done - intended)   // includes the wait a real arrival had

    next_due = next_due + interval

go deeper

for a junior

Be ready to state the idea in plain terms: if the tool stops sending while the system is stuck, the slowest requests are never recorded, so the reported tail looks better than reality. Recognising the name is enough at this level.

for a middle

Explain the mechanism concretely — a blocked simulated user contributes one long sample where a real arrival stream would have produced many — and name the corrections: intended-start-time timing, back-filling, or an open arrival schedule.

for a senior

Show you can diagnose it in someone else's results: throughput dips without matching latency spikes, an implausibly smooth tail, a maximum far above the reported percentile, and an independent probe used as the tiebreaker.

for a principal

Own it as a measurement-validity policy: what every performance report must state about arrival model and timing basis, which results are refused, and how the same bias is prevented in production telemetry where the observer shares fate with the subject.

### The measurement stops exactly when the system is worst Coordinated omission is a sampling bias in latency measurement. It arises whenever the act of experiencing a slow response prevents the measuring process from taking further samples. The classic setting is a closed-loop generator: each simulated user sends a request, blocks until the response arrives, then sends the next. Suppose the system stalls for 900 ms — a stop-the-world pause, a lock convoy, a failover, a cache stampede. During that stall, each blocked user contributes exactly one sample, the long one. Meanwhile a real arrival stream would have delivered dozens of requests during those 900 ms, each waiting for the remainder of the stall plus its own service time. So the run records a handful of long samples where reality would have produced many, and it records them as if each request had waited only from the moment it happened to be sent. The distribution is missing observations, and the missing ones are systematically the bad ones — they are *coordinated* with the outage rather than randomly scattered. The reported 99th percentile can look excellent while the system was, for part of the run, unusable. ### Why the ordinary intuitions fail here Two intuitions get people into trouble. The first is that more samples fix it. They do not: a longer run at the same offered load reproduces the same bias, because the omission mechanism is unchanged. The second is that percentiles are inherently robust. Percentiles are robust to *how extreme* the tail values are, not to *whether the tail was sampled at all*. Coordinated omission attacks the second property. The bias is also invisible in the usual charts. Throughput dips during the stall, but the dip is easy to read as an ordinary fluctuation, and the latency chart afterwards looks calm. Nothing in the output announces that observations are missing; you have to know to look. ### The three corrections **Measure from the intended send time.** Give every request a scheduled arrival moment computed from the workload plan, and compute latency as *response received minus scheduled arrival*, not *response received minus actual send*. A request that sat in the generator's own backlog for 620 ms because a previous request was stuck now carries those 620 ms, which is exactly what a real arrival would have experienced. **Back-fill the implied samples.** When a response takes far longer than the expected inter-arrival gap, synthesise the observations that a real arrival stream would have produced during the stall, with decreasing residual waits. This is the correction applied by latency-recording tooling that offers an expected-interval parameter, and it turns one 900 ms sample into the run of samples the stall really implied. **Drive the load open.** An open arrival schedule keeps issuing during a stall, so the observations exist naturally. This only works if the generator can actually keep to its schedule; if it runs out of threads, sockets or CPU and starts blocking, the open run silently degrades into a closed one and the bias returns. ### A worked example An insurance quote engine is exercised by a 6-hour nightly run with a fixed population of simulated brokers. The report shows a p99 of 214 ms, comfortably inside the agreed threshold, and the run is signed off for weeks. Support meanwhile logs periodic complaints of quotes hanging for several seconds. Re-analysing with intended-start-time latency turns the same raw data into a p99 of 3,120 ms. The explanation is a nightly rating-table refresh that briefly blocks quote evaluation. During each block, the whole simulated population was parked, contributing one long sample apiece and issuing nothing else; the corrected analysis restores the requests that a real broker stream would have queued up behind the block. The team then found a second, unrelated distortion in the same run: an off-by-one boundary in the driver's request index meant one request per iteration was skipped rather than sent, so the achieved rate was quietly 4% under the configured one. That is an ordinary bug, not coordinated omission, and it is worth separating the two in a discussion — omission is a property of *how latency is measured under stalls*, not simply of any missing request. ### What to check before you trust a tail number Ask how latency was timed, and whether the generator was free to send during a stall. Compare achieved throughput against offered throughput over time, because a dip is a fingerprint of blocked issuance. Compare the generator-side distribution against a distribution recorded from a wholly independent path — a low-rate probe that fires on its own schedule regardless of the run, or server-side timings — since the probe cannot be parked by the system's own slowness in the same way. If they disagree in the tail, believe the one that could not be silenced.

  • Does an open workload model make coordinated omission impossible?
    No — it removes the usual cause, but only while the generator keeps to its schedule. If issuing threads block, sockets run out or the generator host saturates, arrivals stop during exactly the stall you care about and the bias returns. That is why an open run should chart achieved rate against offered rate continuously and treat a shortfall as an invalidating condition, not a footnote.
  • How would you detect the bias in results someone else produced?
    Look for a throughput dip with no matching latency spike, a suspiciously smooth tail, or a maximum that is orders of magnitude above the reported high percentile with almost nothing in between. Then ask how latency was timed. An independent low-rate probe running on its own schedule during the same window is the strongest check, because it keeps sampling while the main workload is parked.
  • Is coordinated omission relevant to production monitoring, or only to load tests?
    It applies wherever measurement is coupled to the thing being measured. A client that measures only requests it managed to send, a monitoring agent starved of CPU by the same pressure that slowed the service, or an in-process timer that cannot run during a global pause will all under-report the worst intervals. The mitigation is the same: an observer whose sampling schedule is independent of the subject's health.

A road survey that counts cars only when the surveyor can walk to the roadside will record almost no traffic during the jam that trapped them indoors — and will then report the road as quiet.

saying these in an interview costs you the question

  • Believing high percentiles are automatically robust to missing samples
  • Assuming a longer run averages the bias away
  • Timing latency from when the generator actually sent the request
  • Reporting a clean tail without stating the arrival model
  • Dismissing a throughput dip during a stall as ordinary noise
  • Treating any dropped request as coordinated omission

context