skip to content

One replica's p99 latency doubles at peak while its average CPU use sits at half its ceiling — why?

level: seniorimportance: should knowfreq 54%

answer

  1. the enforced window is short
  2. quota per period, not per minute
  3. bursts stall, idle periods dilute
  4. charged to the group, not per thread
  5. count throttled periods, not usage

basics

~20 s

A CPU ceiling is a quota per short repeating period, not a long-run average. A burst exhausts one period's quota and the process waits out the rest of that period, so requests stall in slices while the averaged number stays low.

solid answer

~50 s

Averaging hides the mechanism. The ceiling is enforced as an amount of CPU time granted per short repeating period — typically a fraction of a second — and the accounting resets every period. A request that arrives in a burst can consume the whole quota early in a period and then sit **runnable but unscheduled** until the next refill. That wait is pure added latency with no error, and it lands on exactly the requests that needed the most CPU, which is why it shows in the tail rather than the mean. Meanwhile the average over a minute is diluted by every quiet period the replica barely used. Parallel work makes it worse: CPU time is charged for the whole group, so several runnable threads drain a period's quota in a fraction of its wall-clock length and then all stall together.

code

pseudocode · 15 lines
pseudocode
# assumptions: ceiling = 0.4 cores, enforced as 40 ms of CPU time per 100 ms period
#              one request needs 60 ms of CPU on a single runnable thread

t = 0     period 1 begins, quota = 40 ms
t = 0..40    request runs; quota reaches 0 at t = 40
t = 40..100  runnable but not scheduled  ->  60 ms of pure waiting
t = 100   period 2 begins, quota = 40 ms
t = 100..120 request finishes its remaining 20 ms

wall_clock_for_the_request = 120 ms      # for 60 ms of actual work
cpu_used_over_both_periods = 60 ms / 200 ms = 0.3 cores = 75% of the ceiling

# same ceiling, four runnable threads instead of one:
quota_drain_time = 40 ms of CPU / 4 threads = 10 ms of wall clock
stall_per_period = 100 ms - 10 ms = 90 ms with every thread waiting

go deeper

for a junior

Hold on to the core fact: the ceiling is checked over a very short repeating window, so a burst can be cut off even when the hour-long average looks modest. Averages and quotas measure different things.

for a middle

Walk the arithmetic out loud — quota per period, exhausted early, stalled for the remainder, resumed next period — and note that CPU time is charged for the whole container, so more runnable threads drain the quota proportionally faster.

for a senior

Demonstrate the diagnosis: correlate cut-short periods with the latency tail, rule out dependencies because nothing errors, and explain why usage under the ceiling is the expected observation rather than a contradiction.

for a principal

Treat it as a sizing doctrine. Ceilings set from average consumption systematically under-serve bursty request workloads, and the headroom you tell teams to leave above the average is a fleet-wide trade of cost against the latency tail.

## The mismatch between the window you look at and the window that is enforced A CPU ceiling is not enforced as an average over the window you happen to be graphing. It is enforced as a **quota per short repeating period**: the container is granted an amount of CPU time at the start of each period, spends it, and is cut off until the next one begins. The period is short — a fraction of a second on typical configurations — while the number a human reads is usually an average over a minute or more. Those two windows disagree in one specific way. An average over a minute can only tell you the **total** CPU time consumed. It cannot tell you **when** inside that minute it was consumed, and the ceiling only cares about when. ## What a stall actually looks like Take a ceiling of 0.4 cores enforced as 40 ms of CPU time per 100 ms period, and a request that needs 60 ms of CPU on a single runnable thread: 1. The period begins and the quota refills to 40 ms. 2. The request runs and exhausts the quota after 40 ms. 3. For the remaining 60 ms of the period the process is runnable but not scheduled. It is not blocked on I/O, not waiting on a lock, not doing anything — it is simply not allowed to run. 4. The next period begins, the quota refills, and the request finishes its last 20 ms. The request took **120 ms of wall clock for 60 ms of work**. Over those two periods the container used 60 ms out of 200 ms, which is 0.3 cores — 75% of its ceiling. Add the idle periods between bursts and the minute-average falls to half the ceiling or lower, which is exactly the number in the question. ## Why it lands in the tail Throttling is not spread evenly across requests. It is charged to whichever requests happen to be in flight when the period's quota runs out, and those are disproportionately the expensive ones and the ones that arrived during a burst. So: - Cheap requests early in a period are untouched; the median barely moves. - Expensive or late-arriving requests absorb the whole remainder of the period; p99 jumps by roughly the leftover of a period, and by a multiple of it when a request needs several periods' worth of quota. - The effect is **load-correlated**, appearing at peak and vanishing off-peak, which makes it look like a downstream dependency problem. ## Parallelism is the multiplier CPU time is charged against the whole accounting group, not per thread. A container with a ceiling of 0.4 cores and four runnable threads drains that 40 ms of quota in about 10 ms of wall clock — and then **all four threads stall together** for the remaining 90 ms. This is why a process that decides its own parallelism from the host's visible CPU count, rather than from its own ceiling, behaves so badly under a small ceiling: it creates far more runnable work than the quota can feed, converting what would have been steady progress into a stutter. ## Confirming it rather than guessing The direct evidence is not the usage number. The accounting group counts **how many periods were cut short and how long the container spent waiting in them**, and that counter is what distinguishes throttling from every other cause of tail latency. Reason about it like this: - Latency rises with load, no errors, no dependency slowdown → consistent with throttling. - Throttled-period count rises in step with the latency → that is the confirmation. - Usage well under the ceiling at the same time → expected, not contradictory; that is the whole point of the mismatch. ## What actually helps - Raise the ceiling so a burst fits inside a single period's quota. - Reduce the parallelism the process creates so the quota feeds it steadily instead of in bursts. - Shrink the per-request CPU cost so a request no longer needs more than one period's quota. - Where the platform allows it, lengthen the accounting period so short bursts can borrow across a longer window — some platforms expose this and others do not, and it trades responsiveness of enforcement against burst tolerance. Raising the **reservation** does not fix throttling by itself: the reservation governs placement, and the ceiling is what cuts the process off. It is the ceiling that has to move.

  • Would raising the reservation instead of the ceiling relieve the throttling?
    No. The reservation decides which node the replica lands on and how much capacity is claimed there; it is not what the runtime enforces on the process. The cut-off comes from the ceiling, so only moving the ceiling changes when the quota runs out. A higher reservation may change placement, and thereby contention, but throttling against your own quota is unaffected.
  • Why does the same workload stop stalling when its ceiling is raised to a whole core, even though it never uses one?
    Because what matters is whether a burst fits inside a single period's quota, not the long-run consumption. With a whole core's worth of quota per period, a request needing 60 ms of CPU finishes inside one period and is never cut off, so the tail collapses while the average consumption barely moves. Headroom above the average is what buys burst tolerance.
  • How would you separate CPU throttling from contention for something no ceiling partitions?
    Throttling is self-inflicted and visible as cut-short periods against your own quota, so it tracks your own load and persists on an otherwise quiet host. Contention for unpartitioned resources tracks the host's other tenants instead, shows no throttled periods, and moves when the co-tenants move. Checking whether the stall correlates with your load or with the neighbours' is the discriminator.

saying these in an interview costs you the question

  • Assumes usage under the ceiling means no throttling occurred
  • Reads the minute average as if the ceiling were enforced over it
  • Thinks each thread is granted the quota independently
  • Blames a downstream dependency because nothing errored
  • Says raising the reservation will stop the throttling