A Kubernetes pod running a search-autocomplete service is slow but never restarts or shows OOMKilled; how do you confirm CPU throttling is the cause?
answer
- slow but alive
- limit becomes a per-period budget
- averages hide 100ms stalls
- throttled periods over total periods
- nr_throttled in cpu.stat
basics
~10 sCheck that the container has a CPU limit, then compare its throttled CFS periods with its total periods: cAdvisor's container_cpu_cfs_throttled_periods_total against container_cpu_cfs_periods_total, or nr_throttled in cpu.stat. A rising ratio confirms throttling.
solid answer
~40 sThrottling never shows up as a restart or an OOMKilled reason, so I look for it directly. First I confirm the container has `resources.limits.cpu` set, because without a CPU limit the kubelet sets no CFS quota and there is nothing to throttle. Then I read the kernel's CFS counters, which the kubelet's embedded cAdvisor exposes: `container_cpu_cfs_throttled_periods_total` divided by `container_cpu_cfs_periods_total` over a window is the fraction of scheduling periods in which the container ran out of quota. I can also run `kubectl exec` and read `nr_throttled` and `throttled_usec` from `/sys/fs/cgroup/cpu.stat` on a cgroup v2 node. I do not trust `kubectl top pod` for this: it shows an average over tens of seconds, and throttling happens inside 100ms periods. A ratio that climbs with the latency, at a usage well under the limit, confirms it.
code
bash · 3 lineskubectl get pod autocomplete-7c9f5d-x2k8q -o jsonpath='{.spec.containers[*].resources}'
kubectl exec autocomplete-7c9f5d-x2k8q -c api -- cat /sys/fs/cgroup/cpu.stat
kubectl get --raw "/api/v1/nodes/worker-11/proxy/metrics/cadvisor" | grep 'container_cpu_cfs_.*autocomplete-7c9f5d-x2k8q'go deeper
Remember that throttling never restarts anything; it only slows the pod. Name the two cAdvisor counters and say that their ratio, not kubectl top, is the evidence.
Explain why an average over tens of seconds cannot show a stall inside a 100ms period, and show how to compute the ratio from two samples of cumulative counters.
Correlate the throttling ratio with p99 over the same window, check the serving container rather than a sidecar, and rule throttling out quickly when the counters stay flat.
Decide which throttling ratio is acceptable per workload class, since batch jobs tolerate what latency-sensitive services cannot, and make that signal part of every service's standard view.
## What CPU throttling is, and why nothing restarts In Kubernetes, a container's `resources.limits.cpu` is not a speed cap in the everyday sense. The kubelet turns it into a **CFS quota**: a budget of CPU time the container may spend in each **scheduling period**, which defaults to **100ms** (the kubelet's `cpuCFSQuotaPeriod`). When the container's threads together spend that budget before the period ends, the kernel stops scheduling them until the next period starts. That pause is **throttling**. Throttling is the opposite of a memory kill: - nothing is terminated, so the **restart count stays flat**; - there is no `OOMKilled` reason in `kubectl describe pod`; - no Event is recorded; - the only visible effect is that requests wait, so **tail latency (p99) rises**. That is why a slow, healthy-looking pod is the classic presentation, and why you have to go looking for the signal. ## Why `kubectl top pod` does not show it `kubectl top pod` reads from metrics-server, which scrapes the kubelet and computes a **rate over its scrape window** (the stock install manifest sets `--metric-resolution=15s`; the flag's own default is 60s). A container that burns its whole quota in the first few milliseconds of every period and then sits blocked can still average well below its limit. Averaging over 150 periods smooths the stalls away. `kubectl top` answers "how much CPU, on average"; throttling is a question about **what happened inside each 100ms slice**. ## The signal: CFS period counters The kernel counts, per cgroup, how many enforcement periods elapsed while the cgroup was active and how many of those ended with the quota exhausted. The kubelet's built-in cAdvisor publishes those counters on the node's `/metrics/cadvisor` endpoint: | Metric | What it counts | How to use it | |---|---|---| | `container_cpu_cfs_periods_total` | Enforcement periods in which the container was active | The denominator | | `container_cpu_cfs_throttled_periods_total` | Periods in which it hit its quota | Divide by periods: the **throttling ratio** | | `container_cpu_cfs_throttled_seconds_total` | Total time its threads spent throttled | How much wall-clock stall there was | | `container_cpu_usage_seconds_total` | CPU time actually consumed | Usage, **not** throttling | The same numbers are visible inside the container on a cgroup v2 node, in `/sys/fs/cgroup/cpu.stat`, as `nr_periods`, `nr_throttled` and `throttled_usec`. ## Confirming it step by step 1. **Check that a CPU limit exists.** `kubectl get pod <name> -o jsonpath='{.spec.containers[*].resources}'`. No `limits.cpu` means no quota, so throttling is ruled out and you look elsewhere. 2. **Rule out the loud failures.** `kubectl describe pod` should show no recent restarts and no `OOMKilled` last state. 3. **Read the counters twice.** Sample `cpu.stat` (or the cAdvisor series) a few minutes apart; the counters are cumulative, so only the **difference** matters. 4. **Compute the ratio.** Throttled periods divided by total periods over the same window. 5. **Line it up with latency.** Throttling that rises and falls with p99 is the cause; throttling that is flat while latency moves is a coincidence. ## Reading the numbers Suppose the search-autocomplete container has `limits.cpu: 1500m` and `kubectl top pod` reports 410m. Over 15 minutes (900 seconds, so up to 9,000 periods of 100ms) the counters move by 9,000 periods and 2,187 throttled periods. The ratio is 2,187 / 9,000 = **24.3%**: roughly one period in four ended with the container frozen, even though its average use was about 27% of its limit. That is the signature: **low average, high throttled ratio, bad tail latency**. A few cautions when reading it: - a small ratio (a few percent) on a batch job may be harmless; on a latency-sensitive service even a modest ratio can be the whole p99 story; - the ratio is **per container**, so look at the container that serves traffic, not a sidecar; - throttled seconds can exceed wall-clock seconds, because stall time is summed across the CPUs the container ran on. ## What throttling is not Throttling is not an eviction, not a probe failure and not a scheduling problem. If the counters stay near zero, the slowness lives elsewhere: a slow dependency, DNS, garbage collection or memory reclaim. The CFS counters let you rule CPU limits in or out in minutes, which is why they are the first thing to check when a pod is slow but alive.
- Why must you sample the CFS counters twice instead of reading them once?`container_cpu_cfs_periods_total` and `container_cpu_cfs_throttled_periods_total` are cumulative since the container started. A single reading mixes last week's traffic with today's incident. Take two samples a known interval apart, subtract, and divide the throttled delta by the periods delta. That gives the throttling ratio for the window you care about, which you can then line up against the latency graph.
- The throttling ratio is 0% but the pod is still slow. What does that tell you?CPU limits are not the bottleneck: either the container has no CPU limit at all, or it never exhausts its quota. Move on to other causes, such as a slow downstream dependency, DNS lookups, garbage-collection pauses, memory reclaim near the memory limit, or the node itself being overcommitted so the container simply gets little CPU under contention.
saying these in an interview costs you the question
- A throttled container would show restarts or an OOMKilled reason
- kubectl top pod below the limit proves there is no throttling
- container_cpu_usage_seconds_total shows throttling directly
- A container without a CPU limit can still be CFS-throttled by its request
- A single reading of the cumulative counters gives the current throttling rate