skip to content

In Kubernetes, how can a container with a 1500m CPU limit be throttled while kubectl top pod shows it using only 410m?

level: middleimportance: must knowfreq 68%

answer

  1. budget per slice, not average
  2. milliCPU times period over 1000
  3. all threads share one cgroup budget
  4. quota does not hide cores
  5. 150ms gone in 3ms

basics

~20 s

The kubelet turns a 1500m limit into 150ms of CPU per 100ms period, shared by all threads. Forty-eight busy threads spend it in about 3ms and then stall for 97ms, while the average stays low.

solid answer

~40 s

The kubelet does not enforce the limit as an average. It converts `limits.cpu` into a CFS quota: `milliCPU × period / 1000`, so 1500m with the default 100ms `cpuCFSQuotaPeriod` becomes 150ms of CPU time per 100ms of wall-clock time. That budget is shared by every thread in the container. If the runtime sized its worker pool from the node's 48 cores and all 48 threads run in parallel, they spend 150ms in about 3.1ms and then sit frozen for the other ~97ms. A request that arrives during the stall waits for the next period, so p99 jumps. Across a 15-second window, though, the container might only use 6.15 CPU-seconds, which `kubectl top pod` shows as 410m. The limit bounds CPU per period, not per minute, and parallelism decides how fast the period's budget goes.

code

yaml · 17 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: autocomplete-api
  labels:
    app: search-autocomplete
spec:
  containers:
  - name: api
    image: registry.example.internal/search/autocomplete:3.14.2
    resources:
      requests:
        cpu: 600m
        memory: 1536Mi
      limits:
        cpu: 1500m
        memory: 2Gi

go deeper

for a junior

Remember that a CPU limit is a budget refilled every 100ms, and that all threads in the container share it.

for a middle

Work the arithmetic aloud: limit to quota, quota divided by parallel threads, and the stall that remains in each period. Explain why the average hides it.

for a senior

Connect throttling to runtime thread pools sized from host cores, and recognise the pattern after a move to bigger nodes with unchanged limits.

for a principal

Weigh changing the node-wide CFS period against per-service fixes, and set platform guidance on how runtimes should size parallelism inside limits.

## From a CPU limit to a CFS quota Kubernetes expresses CPU in **millicores**: `1500m` means one and a half cores' worth of CPU time. Linux does not enforce that as a number of cores. The kubelet translates the container's `resources.limits.cpu` into two cgroup values for the **Completely Fair Scheduler (CFS)** bandwidth controller: - the **period**: how often the budget is refilled, from the kubelet's `cpuCFSQuotaPeriod` setting, default **100ms**; - the **quota**: how much CPU time the container may use in each period, computed as `milliCPU × period / 1000`, with a floor of **1ms**. For `1500m` that is `1500 × 100ms / 1000 = 150ms` per 100ms period. The kubelet only does this when its `cpuCFSQuota` setting is true, which is the default, and only for containers that set a CPU limit. ## The budget is shared by every thread The quota belongs to the **cgroup**, not to a thread. Every thread of every process in the container draws from the same 150ms. How fast it drains depends on how many threads run at once: | Threads running in parallel | Time to spend 150ms | Stall until the refill | |---|---|---| | 1 | never (at most 100ms per period) | none | | 2 | 75ms | 25ms | | 8 | 18.75ms | 81.25ms | | 48 | 3.125ms | 96.875ms | With one thread, a 1500m limit can never bind. With 48 threads, the container spends its whole period's budget in about **3ms** and is frozen for about **97ms**. Any request that is mid-flight, or arrives during the stall, waits for the next period. That added wait is what shows up as a doubled p99. ## Why the average looks harmless `kubectl top pod` shows a **rate averaged over a window** of tens of seconds (metrics-server's scrape interval; the stock install manifest sets 15s). A bursty service is mostly idle and occasionally very busy: 1. a burst of autocomplete queries arrives; 2. 48 worker threads wake at once and spend the quota in 3ms; 3. the threads stall for the rest of the period, and latency spikes; 4. the burst ends and the container is idle for many periods. Over a 15-second window the container may use only 6.15 CPU-seconds of a possible 22.5 (1.5 cores × 15 seconds). That is **410m**, about 27% of the limit, and it looks healthy. The average is correct; it just answers a different question from the one throttling asks. ## Where the 48 threads come from Many language runtimes and libraries size their thread pools from **the number of CPUs they can see**. A CFS quota does **not** hide cores: a process in a container with a 1500m limit on a 48-core node can still see and schedule on all 48 cores. So a runtime that asks "how many CPUs do I have?" may answer 48 and create 48 workers, garbage-collection threads or parallel-stream workers. Some runtimes now read the cgroup limit instead, but you have to check what yours reports inside the pod rather than assume. The result is a mismatch: **the parallelism was sized for the node, the budget was sized for the container**. This is the most common reason a service throttles far below its limit, and it often appears after a move to larger nodes, since the thread count grows with the core count while the limit stays the same. ## Levers that change the arithmetic - **Fewer parallel threads** spend the budget more slowly: sizing the pool to roughly the limit keeps bursts inside the period. - **A higher limit** raises the budget per period. - **No limit** removes the quota entirely; the container's CPU request still sets its share under contention. - **A different period**: the kubelet's `cpuCFSQuotaPeriod` accepts 1ms to 1s. The `CPUCFSQuotaPeriod` feature gate reached GA in Kubernetes 1.36. It is a **node-wide** setting, so it is a platform decision, not a per-service fix. ## Takeaways - A CPU limit is a **budget per period**, not a ceiling on the average. - **Parallelism** decides how quickly the budget drains. - A low `kubectl top` value and heavy throttling are **not contradictory**; they measure different time scales.

  • Would lowering the CFS period to 10ms on the node help this service?
    It shortens each stall: 1500m becomes 15ms per 10ms, so a burst stalls for at most a few milliseconds instead of ~97ms. But it is a node-wide kubelet setting (`cpuCFSQuotaPeriod`) that affects every pod, adds scheduler overhead, and does not reduce total throttling for a genuinely over-parallel workload. Fixing the thread count or the limit is usually the better lever.
  • The same service throttled less on 16-core nodes than on 48-core nodes with an identical limit. Why?
    The quota is the same, 150ms per 100ms, but a runtime that sizes its pools from the visible core count created 16 workers on the smaller node and 48 on the larger. Sixteen parallel threads spend the budget in about 9.4ms; forty-eight spend it in about 3.1ms, so the stall per period is longer and more requests land in it.

It is like a data allowance refilled every tenth of a second: forty-eight downloads at once burn it instantly and you are offline until the refill, even though your daily average is tiny.

saying these in an interview costs you the question

  • A 1500m limit means the container is pinned to one and a half cores
  • Throttling only starts once average usage reaches the limit
  • Each thread in a container gets its own separate CPU quota
  • A CPU limit hides the node's extra cores from the application
  • Throttling is proof the container needs more CPU on average