skip to content

Why is a service's mean garbage-collection pause a poor predictor of what its slowest requests experience?

level: seniorimportance: should knowfreq 48%

answer

  1. averages hide concentration
  2. a pause lands whole, not spread
  3. request window plus pause length
  4. percentiles, never the mean
  5. fan-out multiplies the chance

basics

~20 s

Pauses are rare and arrive whole. They miss most requests entirely and land in full on a few, so an average spreads a concentrated cost evenly and describes nobody's experience; the tail keeps the concentration users actually feel.

solid answer

~50 s

A collection pause is not a small tax on every request, it is a total stall for the requests unlucky enough to be in flight. Take a service that pauses 50 ms every 2 seconds while serving 20 ms requests: a request is caught whenever a pause begins anywhere in a window of `pause + request` around it, so about `(50 + 20) / 2000` = **3.5%** of requests are hit, and those take up to 70 ms instead of 20 ms. The mean impact is under 2 ms and is meaningless; the 99th percentile is set entirely by the collector. Fan-out makes it worse: a request touching 10 replicas independently has roughly a **30%** chance of hitting at least one paused replica. Judge collectors on the pause distribution at the percentile your service is measured on.

code

pseudocode · 16 lines
pseudocode
# P = pause length, T = interval between pauses, D = request duration
# a request is caught if a pause STARTS inside a window of (P + D)

function share_of_requests_hit(P, T, D):
    if P + D >= T:
        return 1.0               # pauses arrive faster than requests clear
    return (P + D) / T

share_of_requests_hit(50, 2000, 20)        -> 0.035

# a request fanning out to k replicas escapes only if all k escape
function fan_out_hit(P, T, D, k):
    p = share_of_requests_hit(P, T, D)
    return 1 - (1 - p) ^ k

fan_out_hit(50, 2000, 20, 10)              -> 0.30

go deeper

for a junior

Remember that a collection pause stops the program completely for its duration, so it does not slow every request a little - it stops a few requests entirely.

for a middle

Explain the arithmetic: the share of requests caught is roughly the pause plus the request duration over the interval between pauses, which is why percentiles and the mean disagree so sharply.

for a senior

Show that you measure the distribution and the maximum, translate it into the percentile the service is judged on, and account for fan-out amplification across replicas.

for a principal

Connect the pause distribution to the latency objective the organisation has committed to, and decide how much throughput the fleet will pay to move a percentile - or whether to change the architecture's fan-out instead.

## An average is the wrong summary for a rare, whole-sized event Averaging works when a cost is spread thinly across a population. A collection pause is the opposite kind of event: it is rare and it is indivisible. During the stall, requests in flight make no progress at all - not slower progress, none. Requests that arrive and finish between stalls pay nothing. So the population splits in two: an unaffected majority and a fully affected minority. The mean reports a number that describes neither group, and the smaller and longer the pauses, the more misleading it gets - because the same mean is consistent with 'everyone slightly slower' and with 'one request in thirty is three times slower'. ## Working out who gets hit For pauses of length `P` arriving every `T`, and requests taking `D`: - A request is affected if a pause **starts** at any point from `P` before it begins to `D` after - a window of `P + D`. - So the fraction affected is about `(P + D) / T`. With `P` = 50 ms, `T` = 2000 ms and `D` = 20 ms, that is `70 / 2000` = **3.5%**. Those requests stretch from 20 ms toward 70 ms. Meanwhile the arithmetic mean of added latency across all requests is roughly `0.035 x 50` ms, under 2 ms - a number a capacity review would wave through and a number no user ever experiences. Read the same figures as percentiles instead and the picture inverts: with 3.5% of requests degraded, everything above the 96.5th percentile is a collection pause. If the service is judged on its 99th percentile, the collector **is** the service's latency for that objective. ## Fan-out amplifies the minority A request that consults several replicas in parallel is only as fast as its slowest branch. If each replica independently degrades 3.5% of the time, a request touching `k` of them escapes only if all `k` escape: `hit = 1 - (1 - 0.035)^k` - `k` = 1 -> 3.5% - `k` = 5 -> about 16% - `k` = 10 -> about **30%** A per-replica effect that looked negligible has become nearly a third of all requests once the architecture fans out. This is why collector pauses are argued about far more in systems with wide fan-out than the per-process numbers alone would justify. ## What to measure instead 1. **The pause distribution, not the mean** - the maximum and the high percentiles of pause length, plus how often pauses occur, since frequency drives the share of requests caught. 2. **The percentile the service is actually judged on**, measured end to end from the caller, which is where any fan-out amplification is visible. 3. **Both together.** Two collectors with identical total processor cost can deliver it as many small stalls or a few large ones; that distinction is invisible in total collector time and decisive for a request-serving service. ## The traps - **Shorter pauses do not automatically shrink the tail.** If buying them costs throughput, utilisation rises, and queueing delay can replace stall time in the tail. The tail is a property of the whole system, not of the collector alone. - **A rare very long pause dominates.** One multi-second stall an hour may not move any percentile below the 99.9th and can still be the worst thing the service does all day, because it can trip health checks and shed traffic. - **Averaging across replicas hides it.** A fleet-wide mean pause tells you almost nothing; the question is what the worst replica did while it held a request. The summary that earns credit: a pause is a concentrated cost, an average is a de-concentrating summary, and the two are the wrong pair. Report the distribution, and report it at the percentile someone has promised to meet.

  • Two collectors burn identical processor time. How do you tell which suits a request-serving service?
    Compare the shape, not the total. Look at pause length at the high percentiles and how often pauses occur, then translate that into the share of requests caught and the percentile the service is judged on. Equal total cost delivered as many short stalls beats one long stall.
  • Why can shortening pauses leave the latency tail exactly where it was?
    Because the tail may not be the collector's. If buying short pauses costs throughput, utilisation and queueing delay rise to fill the gap; and if the tail was always set by a slow dependency or lock contention, collector work was never the binding constraint. Confirm the cause before buying the cure.
  • How does a very long, very rare stall fit this analysis?
    It may sit above every percentile anyone reports and still be the worst event of the day, because a long enough stall fails health checks, drops connections and causes traffic to be shed. Track maximum pause length separately from percentiles.

A drawbridge that lifts for one minute every hour delays almost nobody on average. Everyone who arrives while it is up waits the entire minute.

saying these in an interview costs you the question

  • Reports mean pause time as the service's latency impact.
  • Assumes a pause missing most requests cannot matter.
  • Ignores that one paused replica slows a fanned-out request.
  • Treats total collector processor time as equivalent to pause length.
  • Believes shorter pauses always shrink the latency tail.