A Docker container's CPU usage sits below its limit yet p99 latency spikes. How do you confirm CPU throttling?
answer
- Averages hide what happens inside a period
- The counter is absent from docker stats
- Look for a per-period enforcement counter
- cpu.stat carries periods, throttled and time
- Sample twice and diff the monotonic counters
basics
~20 sRead the container's cgroup cpu.stat twice a minute apart and diff it. A rising nr_throttled against nr_periods, plus growing throttled_usec, proves the CPU quota is being exhausted inside periods even when average usage looks low. docker stats cannot show this.
solid answer
~40 sAverage CPU usage is the wrong instrument. A CPU quota is enforced per scheduling period, so bursty work can exhaust its slice early in the period, be stopped until the period rolls over, and still average well under the limit. The evidence lives in the container's cgroup: `cpu.stat` carries `nr_periods`, `nr_throttled` and `throttled_usec` (cgroup v1 calls the last one `throttled_time`, in nanoseconds). Sample it twice, diff the counters, and compute the throttled share of periods; a few percent is noise, a third or more explains user-visible latency. Correlate the diff window with the latency spike. Then distinguish the causes: throttling shows quota exhaustion with idle-looking averages, saturation shows usage pinned at the limit continuously, and neither appears when the process is actually blocked off-CPU on I/O or a lock.
code
bash · 7 linesdocker exec reranker cat /sys/fs/cgroup/cpu.stat
sleep 60
docker exec reranker cat /sys/fs/cgroup/cpu.stat
# nr_periods 9143
# nr_throttled 8412
# throttled_usec 21437905go deeper
Know that a CPU limit works as a budget refilled every scheduling period, so a container can be stopped mid-burst while its average usage still looks low. Remember that the throttling counter lives in the cgroup, not in docker stats.
Be ready to name the cpu.stat fields, explain why they must be diffed rather than read once, and turn the diff into a throttled-share percentage that you can defend as significant or not.
Demonstrate the triage: correlate the counter diff with the latency window, and separate throttling from saturation and from off-CPU waiting before proposing any change. Say clearly what evidence would make you drop the CPU hypothesis.
Own the standing capability rather than the incident. Decide that throttling counters are collected for every container by default, define the threshold at which a service is considered mis-limited, and weigh the cost of headroom against the risk of tail latency across the estate.
### Why average usage hides the problem CPU limits on a container are enforced as bandwidth over a repeating scheduling period, not as a speed governor. The container's tasks are allowed a budget of CPU time per period; when the budget is gone, every runnable thread in the group is stopped until the next period begins. A request that arrives just after the budget ran out waits for the remainder of the period before it gets any CPU at all. That is why the averages lie. Take a recommendation re-ranker written in Scala on a JDK base image: it is idle between requests, then wants several cores for a few milliseconds of scoring. Averaged over a minute it might use 38% of its allowance and look comfortably provisioned, while individual periods are being cut off repeatedly. The symptom is a latency distribution with a healthy median and a p99 that jumps - in one incident, a p99 of 41 ms moved to 380 ms with no change in request rate and no change in average CPU. ### The counter that proves it The evidence is in the container's own cgroup, in `cpu.stat`: - `nr_periods` - how many enforcement periods have elapsed for this group - `nr_throttled` - how many of those periods ended with the group throttled - `throttled_usec` - total time the group spent stopped (cgroup v1 names this `throttled_time` and counts nanoseconds) These are monotonic counters since the container started, so a single reading tells you almost nothing - a container up for three weeks accumulates throttling from one bad afternoon and looks permanently sick. Always sample twice around a window you care about and diff: ``` docker exec reranker cat /sys/fs/cgroup/cpu.stat # t0 sleep 60 docker exec reranker cat /sys/fs/cgroup/cpu.stat # t1 ``` Over 60 seconds with a 100 ms period you expect roughly 600 periods per CPU of accounting. In the incident above the diff showed 8,412 throttled periods out of 9,143, and about 21.4 seconds of `throttled_usec` accumulated in a 60-second window - unambiguous. As a rule of thumb, a throttled share under a few percent is normal for bursty services, ten to twenty percent is worth acting on, and anything above a third is already visible to users. On a cgroup v2 host with the default private cgroup namespace, `/sys/fs/cgroup/cpu.stat` inside the container is the container's own file. The identical file is readable from the host under the container's cgroup path, which is how you take the measurement when the image has no shell, and how a metrics agent collects it continuously. ### What `docker stats` will and will not tell you `docker stats` shows a CPU percentage derived from consumed CPU time; it has no throttling column. A throttled container looks *quiet* there, which is exactly why teams misread the situation and go hunting in application code. Use `docker stats` to see whether usage is pinned at the ceiling; use `cpu.stat` to see whether the ceiling is being hit inside periods. ### Telling the three failure shapes apart Once you have both numbers, the diagnosis is mechanical: 1. **Throttling.** Average usage well below the limit, `nr_throttled` climbing fast, latency spiky rather than uniformly slow. The work arrives in bursts that do not fit in a single period. 2. **Saturation.** Usage sitting flat against the limit for the whole window, `nr_throttled` also climbing but with the group genuinely runnable the entire time. Everything is slow, not just the tail. This is a capacity problem, not a burst-shape problem. 3. **Off-CPU waiting.** Latency is bad, usage is low, and `nr_throttled` barely moves. The process is blocked on a database call, a lock, a slow disk, or a DNS lookup. No amount of CPU tuning helps, and this is the case where reading `cpu.stat` earns its keep by *excluding* the CPU hypothesis in thirty seconds. A fourth shape is worth naming because it is commonly confused with the first: a runtime that has sized its thread pools from the host's CPU count will create far more runnable threads than the quota can feed, converting a modest limit into constant throttling. The counter looks identical; the fix is in the application's configuration rather than in the limit. ### Reporting it Bring three things to the review: the diffed counters with the window they cover, the throttled share as a percentage of periods, and the latency series over the same window. That triple distinguishes "this container is limited too tightly", "this workload is too bursty for the period" and "the CPU was never the problem" - and it keeps the conversation off guesswork. Choosing the replacement limit is a separate, deliberate exercise in sizing; the diagnosis stops at proving which of the three you are in.
- You read cpu.stat once and see a huge nr_throttled. Why is that not yet evidence of a problem?The fields are monotonic counters accumulated since the container started. A container that has been up for weeks carries the throttling from every past incident, so a single large number says nothing about now. Take two samples around the window you are investigating, diff them, and express nr_throttled as a share of nr_periods over that window.
- Latency is bad, CPU usage is low, and nr_throttled is not moving. What does that combination rule out?It rules out the CPU limit entirely. The container is neither exhausting its quota nor saturating what it has, so the threads are blocked off-CPU - waiting on a downstream call, a lock, a disk, or name resolution. That is a thirty-second exclusion that stops a team from re-tuning limits that were never the constraint.
- How do you take this measurement when the image ships no shell?Read it from the host. The container's cgroup directory exists on the host and contains the same cpu.stat, so no entry into the container is needed; the cgroup path depends on the daemon's cgroup driver. This is also how a collector agent gathers the counter continuously, which is preferable to sampling by hand during an incident.
It is like a toll road that lets you through only sixty cars a minute: measured over an hour the traffic looks light, but every car that arrives after the sixtieth still waits for the clock to tick over.
saying these in an interview costs you the question
- Concludes there is no CPU problem because average usage is low
- Looks for a throttling column in docker stats
- Reads cpu.stat once and treats the total as a rate
- Confuses throttling with the OOM killer or a crash
- Assumes throttling means the limit was exceeded on average
- Blames application code before excluding the quota