What is coordinated omission in latency benchmarking, and why does a naive 'measure each request after the previous one finishes' loop drastically under-report tail latency?
answer
- Closed-loop send waits for prior response -> omits requests during a stall
- One slow sample recorded instead of many that queued behind it
- Tail (p99/p99.9) collapses; always optimistic
- Fix: latency = completion - INTENDED start time
- Tools: HdrHistogram recordValueWithExpectedInterval, wrk2, open-loop
basics
~20 sIf your benchmark sends the next request only after the previous one finishes, then when one request stalls, you simply don't send the requests that should have arrived during the stall. You record one slow sample instead of many. So your worst-case (tail) latency looks far better than what real users would see.
solid answer
~60 sCoordinated omission is a measurement bias in latency benchmarks where the load generator's send rate is 'coordinated' with the system under test: when a response is slow, the next request isn't issued until the slow one completes, so requests that should have been sent during the stall are silently omitted. The result is that a single long pause (a GC, a lock, a stall) is recorded as one slow sample, when in reality every request that would have arrived during that window also suffered — they queued behind it. This collapses the tail: p99/p99.9 latency looks dramatically better than production. The fix is to measure against an intended, fixed schedule: each request has an expected start time; if the system was busy, the recorded latency is measured from when the request should have started, not when it actually got dispatched. Tools like HdrHistogram and wrk2 correct for this; JMH's throughput modes and SampleTime help but you must still drive load at a fixed rate, not closed-loop, to avoid it.
go deeper
Aware that a benchmark that sends one request at a time can hide how bad slow requests really are; trusts purpose-built latency tools.
Can describe that pausing sends during a stall drops samples and makes tail latency look too good; knows to use percentiles and proper load tools.
Explains closed-loop vs open-loop, why the tail collapses, and uses HdrHistogram/wrk2 with intended-start-time correction.
Frames it as a systemic, always-optimistic bias against SLO validation and capacity planning; designs load tests with realistic arrival processes and back-filled/corrected histograms, and reviews others' results for the CO signature.
## Closed-loop vs open-loop load There are two ways to drive load at a system: - **Closed-loop (synchronous) loop:** a fixed set of client threads, each doing `send request → wait for response → send next`. The next request's start time *depends* on the previous response. This is what a naive benchmark loop does. - **Open-loop:** requests arrive on an **independent schedule** (e.g. 10,000/sec, governed by an external clock or arrival process), regardless of how fast the system responds. This models real users/traffic, who don't wait politely for the server to recover before clicking again. Real-world traffic is fundamentally open-loop: users and upstream callers send requests on *their* schedule. Your server doesn't get to throttle the arrival rate by being slow. ## What coordinated omission is **Coordinated omission** (a term coined by Gil Tene) is the bias that arises when the **measurement is coordinated with the thing being measured** — specifically, when a closed-loop generator *pauses sending* because the system is slow. Concretely: Suppose you intend ~1 request/ms and one request hits a **100 ms GC pause**. In a closed-loop loop: - That one request records ~100 ms latency. - But during those 100 ms, you *did not send* the ~100 other requests that, on your intended schedule, should have gone out. You **omitted** them — and that omission was *coordinated* with the very stall you're trying to measure. - In reality those 100 requests would have *queued behind* the stall and experienced latencies of ~100 ms, ~99 ms, ~98 ms, … down to near zero. So you recorded **one** 100 ms sample instead of ~**one hundred** samples in the 0-100 ms range. Your histogram is missing exactly the samples that define the tail. ## Why the tail collapses Latency SLOs live in the **tail**: p99, p99.9, p99.99. Coordinated omission removes precisely the high-latency samples a stall produces, so: - The mean barely moves (it's dominated by the many fast samples). - The **percentiles look fantastic** — p99 might read 2 ms when real users see 90 ms. - The bigger the stall and the higher the intended rate, the worse the under-reporting. This is *systematically optimistic*: it always makes the system look better than it is, and worst exactly where it matters (the tail), so it's especially dangerous for capacity planning and SLO validation. ## The correction The fix is to measure latency against an **intended start time**, not the actual dispatch time: 1. Define an arrival schedule (fixed rate, or a realistic arrival distribution). Each request *i* has an **expected start time** `t_i`. 2. The recorded latency is `completion_time - t_i` (when it *should* have started), **not** `completion_time - actual_dispatch_time`. If the system was busy and you dispatched late, that lateness is correctly counted as latency. 3. When a stall happens, you either keep sending on schedule (a true open-loop generator) or **back-fill the omitted samples**: after a long response, synthesize the latencies of the requests that *would* have been issued during the stall (HdrHistogram's `recordValueWithExpectedInterval` does exactly this). **Tools that handle it:** `wrk2` (constant-throughput, CO-corrected), `HdrHistogram` (`recordValueWithExpectedInterval`), and load generators that decouple send schedule from responses. JMH's `SampleTime` mode plus a fixed-rate driver helps, but a plain closed-loop JMH `@Benchmark` measuring per-call latency in a tight loop is still vulnerable. ## Why it belongs under microbenchmark pitfalls It's the latency analogue of the throughput pitfalls: a hand-rolled timing loop, by construction, coordinates its own sending with the system's responsiveness, so it *cannot* see the queuing a real open-loop arrival process would create. The fix requires an external, schedule-driven notion of "when this request should have happened" — which is precisely why purpose-built harnesses/tools exist rather than a for-loop with `nanoTime()`. ## How to spot it Red flags: a benchmark reports near-flat percentiles despite the system having *known* multi-millisecond pauses (GC, lock contention); the p99 of a closed-loop load test is suspiciously close to the median; or latency is measured as `end - sendTime` in a loop where `sendTime` is captured *after* the previous response returned.
- Does coordinated omission affect throughput numbers as much as latency percentiles?It primarily corrupts latency percentiles, especially the tail. Average throughput is roughly preserved (the work eventually gets done), but the latency distribution is systematically optimistic. That's dangerous because SLOs are usually expressed as tail latency, not mean throughput, so the metric that matters most is exactly the one most corrupted.
- How does HdrHistogram's recordValueWithExpectedInterval correct for it?When you record a latency larger than the expected inter-arrival interval, it back-fills synthetic samples representing the requests that should have been sent during the stall — e.g. a 100 ms sample at a 1 ms expected interval generates additional samples at 99 ms, 98 ms, … — so the histogram reflects the queuing the closed-loop generator failed to produce.
saying these in an interview costs you the question
- Thinking the mean is enough — coordinated omission corrupts the tail, not the mean.
- Measuring latency as completion minus actual dispatch time instead of intended start time.
- Assuming a closed-loop load test reflects real open-loop user traffic.
- Trusting suspiciously flat p99s from a system with known GC/lock pauses.