skip to content

A service's processor utilisation averages 40% of its ceiling, yet its stopped-for-quota counter climbs and latency spikes — why?

level: seniorimportance: should knowfreq 52%

answer

  1. the average is not the enforcement
  2. enforcement happens window by window
  3. allowance spent early, stopped late
  4. latency does not average away
  5. count the stops, not the mean

basics

~20 s

A processor ceiling is enforced as run time allowed per short repeating window, while utilisation is averaged over a far longer one. The burst spends its allowance early and is stopped until the next window.

solid answer

~50 s

The two figures are measured over different spans. The ceiling is enforced window by window: the workload gets an allowance of run time in each short accounting window, and once it is spent, everything in the container is stopped until the next window begins. Utilisation, as a dashboard shows it, is an average over seconds or minutes, so short stops vanish into surrounding idle time and the mean settles well below the ceiling. Latency does not average away — every request that was in flight when the stop landed carries the pause. The signal that does not hide it is the throttling counter itself: a rising tally of windows in which the allowance ran out. Treat that as the saturation signal, and treat the average as a capacity figure that cannot see inside a window.

code

pseudocode · 19 lines
pseudocode
# assumed: ceiling worth 2 processors, accounting window 100 ms
window          = 100 ms
allowance       = 2 * window          # 200 ms of run time per window
runnableThreads = 16

at window start:
    spent = 0 ms
    while spent < allowance and work remains:
        run all runnable threads
        spent = spent + (runnableThreads * elapsedWallClock)

    # allowance gone after 200 / 16 = 12.5 ms of wall clock
    if spent >= allowance:
        stop every thread in the container
        stoppedWindows = stoppedWindows + 1
        stoppedTime    = stoppedTime + (window - 12.5 ms)   # 87.5 ms
        wait for next window            # no carry-over of leftovers

# minute-wide average over mostly idle time still reads ~40%

go deeper

for a junior

Recall that a processor ceiling grants run time per short window, and that a workload which spends its allowance early waits for the next window. It is slowed, not ended.

for a middle

Explain why an average over a minute cannot reveal a stop that lasted milliseconds, and name the throttling tally as the signal that does show it.

for a senior

Work the trace with numbers, connect the stops to the latency the workload owes, and choose between lowering concurrency, pacing the burst and raising the ceiling on evidence rather than instinct.

for a principal

The call you own is where the default ceiling and the default concurrency sit for workloads across the estate, and how much throttling is acceptable for a batch tier against a latency-bound tier.

## Utilisation is an average; the ceiling is enforced in windows A processor ceiling is not a speed setting. It is an **allowance of run time granted in each repeating accounting window**. Within a window the workload runs at full speed until the allowance is gone; then everything in the container is stopped until the next window opens, and it starts again with a fresh allowance. The utilisation number on a dashboard is a different shape entirely: run time consumed over a reporting interval, divided by the run time the ceiling would have allowed over that same interval. That is a perfectly correct capacity figure and a perfectly blind saturation figure, because it cannot see inside a window. Two workloads with an identical 40% average can have completely different experiences: one spends 40% of every window, never hitting the allowance; the other spends its whole allowance in the first eighth of some windows and sits stopped for the rest. ## A trace, with the numbers stated Assume a ceiling worth **2 processors** and an accounting window of **100 ms**, so the allowance is **200 ms of run time per window**. A request arrives that makes **16 threads runnable at once**: 1. All 16 threads run. Together they consume run time 16 times faster than wall-clock time. 2. The 200 ms allowance is therefore exhausted after **12.5 ms** of wall-clock time. 3. Every thread in the container is stopped for the remaining **87.5 ms** of the window. 4. The next window opens, work resumes, and if the request is not finished the cycle repeats. 5. The request carries **87.5 ms of pure waiting** for every window boundary it crossed — and nothing in it was blocked on a dependency. Now suppose bursts like this occupy one second in every five and the container is nearly idle in between. Averaged over a minute, utilisation lands near 40%. Averaged over 12.5 ms of a burst window, it was 100%. ## Which signal shows what | Signal | Span it summarises | What it reveals | |---|---|---| | Average utilisation against the ceiling | seconds to minutes | how much of the allowance is used overall — a capacity figure | | Peak or high-percentile utilisation | the reporting interval | that demand is uneven, but still not whether the allowance ran out | | Stopped-for-quota tally and stopped time | window by window | that the allowance genuinely ran out, and how much wall-clock time was lost to it | The third row is the saturation signal for a processor ceiling. It is a running tally, so what matters is whether it is increasing while the workload is serving — and, more usefully, how much wall-clock time the stops added. ## What actually helps - **Reduce how much is runnable at once.** If the burst puts fewer threads on the run queue, demand fits inside the window and the stops disappear without changing the ceiling at all. This is the fix that costs nothing. - **Flatten the burst.** Work that can be queued and paced across windows — background maintenance, batch flushes, warm-ups — need not compete with a latency-sensitive request inside one window. - **Raise the ceiling** when the demand is real and the latency matters. That is a capacity decision, and a rising stopped tally under a real latency objective is the evidence for it. - **Do not add workers.** More concurrency against the same allowance makes the exhaustion arrive sooner, not later. - **Alert on the stops, not on the average.** An alert keyed to average utilisation on a bursty workload will not fire until the workload is in serious trouble. ## Two directions to keep straight - **A processor ceiling stops a workload; a memory ceiling ends it.** Exceeding an allowance of run time slows a workload and it stays alive; exceeding a memory ceiling kills the process. Describing throttling as "the container was killed for CPU" is the classic reversal. - **Unused allowance does not carry over.** Each window opens with a fresh allowance and the remainder of the previous one is gone. This is why a workload that is idle for 90 ms of a window still cannot spend 400 ms in the next one. ## Is throttling always a problem? No, and this is the judgment part. A stopped tally climbing on a batch worker with no latency obligation is the ceiling doing exactly the job it was declared for — the work takes longer and nothing is harmed. The same signal on a request-serving workload with a latency objective is a direct cause of the latency it is missing. The question is never "is it throttled?" but "what did the stops cost against what this workload owes".

  • Does a climbing stopped-for-quota tally always mean the ceiling is too low?
    No. It means demand exceeded the allowance inside some windows. On a batch worker with no latency obligation that is the ceiling working as declared. On a request-serving workload with a latency objective it is a direct cause of missed latency. Judge the stops against what the workload owes, not against zero.
  • Two replicas show the same average utilisation, but only one is being stopped for quota. What differs?
    The shape of demand inside the window. One spreads its run time evenly and never reaches the allowance; the other makes many threads runnable at once and exhausts the allowance early, then waits. Same mean, different distribution — which is exactly what an average over seconds cannot show.
  • Why does raising concurrency make throttled latency worse rather than better?
    Because the allowance per window is fixed. More runnable threads consume it faster, so the stop arrives earlier in the window and more requests are in flight when it lands. The extra threads also add switching cost, which consumes part of the same allowance.

A prepaid power meter that allows a fixed number of minutes of electricity each hour: burn them in the first five minutes and the lights are off until the next hour begins, even though the hour's average draw looks modest.

saying these in an interview costs you the question

  • Says a low average utilisation proves there is no processor shortage
  • Claims exceeding a processor ceiling kills the container
  • Thinks an unused allowance carries over into the next window
  • Adds workers to fix latency that comes from being stopped for quota
  • Reads a pause as a memory problem because the process stopped
  • Treats any throttling at all as an incident regardless of the workload