skip to content

cgroups & Resource Enforcement

The resource-control half of containers: how cgroup controllers meter CPU, memory and IO, and how docker run flags land in the cgroup files the kernel reads. Interviewers pair it with namespaces to test the full 'what makes a container' answer.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

5

Explain the difference between the container options --cpus, --cpu-shares and --cpuset-cpus, and why an application can be slow while its CPU usage looks low.

level: middleimportance: must knowfreq 48%

answer

  1. --cpus = hard CFS quota per 100 ms period (cpu.max)
  2. shares/weight = relative, only under contention, never throttles
  3. cpuset = pinning, not metering
  4. burst exhausts quota early -> stall until next period
  5. proof: cpu.stat nr_throttled / throttled_usec

basics

~20 s

--cpus is a hard quota per scheduling period (cpu.max), --cpu-shares is a relative weight that only matters under contention (cpu.weight), and --cpuset-cpus pins the container to specific cores. A container that burns its quota early is throttled until the next period, so latency spikes while average utilization looks modest.

solid answer

~50 s

**--cpus=N** sets CFS bandwidth control: quota N x period, written to cpu.max, default period 100 ms. It is an absolute ceiling and applies even when the host is idle. **--cpu-shares** (cpu.weight on cgroup v2) is proportional: it decides how contended CPU time is divided between cgroups and has no effect when there is spare capacity. It cannot cause throttling. **--cpuset-cpus=0-3** pins the container's tasks to specific CPUs. That bounds parallelism and helps cache locality and NUMA placement, but it meters nothing - two pinned containers still contend. The slow-but-idle-looking symptom is CFS throttling. A multi-threaded process with `--cpus=1` can consume its 100 ms of quota in the first 25 ms of a 100 ms period across four runnable threads, then every thread is descheduled for 75 ms. Average utilization reads modest over a coarse window, while p99 latency shows fixed-size stalls. Evidence is in cpu.stat: nr_throttled and throttled_usec climbing. Fixes: raise the quota, right-size thread and GC pools to the quota, or prefer weights for latency-critical services.

code

bash · 6 lines
bash
docker run -d --cpus=1.5 app          # cpu.max = 150000 100000 (hard ceiling)
docker run -d --cpu-shares=2048 app   # relative weight, only under contention
docker run -d --cpuset-cpus=0-3 app   # pinned to CPUs 0-3

docker exec app cat /sys/fs/cgroup/cpu.stat
# nr_periods 5921  nr_throttled 1840  throttled_usec 137004221

go deeper

for a junior

Distinguish the three flags in one sentence each and know that --cpus is the hard limit.

for a middle

Explain quota-and-period mechanics, why bursts cause throttling, and where to find nr_throttled.

for a senior

Diagnose tail-latency incidents from throttling counters, right-size concurrency to the quota, and argue quota versus weight per workload class.

for a principal

Set fleet policy on limits versus weights, node density and overcommit, and the tradeoff between predictable isolation and utilization.

## Three different mechanisms **Bandwidth (--cpus, or --cpu-quota/--cpu-period).** The Completely Fair Scheduler's bandwidth control gives a cgroup a quota of CPU time per period. cpu.max holds `"<quota> <period>"` in microseconds; `--cpus=2` is `200000 100000`. The group may run on any CPU, but the sum of the time its tasks consume within one period cannot exceed the quota. When it does, all its tasks are throttled - made unrunnable - until the period rolls over. This is a hard ceiling: an idle 64-core host does not let a `--cpus=1` container use more. **Weight (--cpu-shares, cpu.weight).** A relative number used only when runnable tasks compete for the same CPU. Two containers with weights 100 and 200 get roughly a 1:2 split *while both are busy and the CPU is saturated*; if one is idle, the other uses everything. Weights never throttle and never waste idle capacity. Docker converts the legacy shares value into a v2 cpu.weight, so the number in the file will not match the flag. **Affinity (--cpuset-cpus, --cpuset-mems).** Restricts which CPUs (and NUMA memory nodes) the tasks may run on. It caps parallelism by core count and improves cache and NUMA locality, which matters for latency-sensitive or memory-bandwidth-bound workloads. It is not accounting: several containers pinned to the same cores still fight for them. ## Why throttling looks like "slow but idle" Quota is consumed by *all* runnable threads in aggregate, but the period is fixed at 100 ms. A service with 8 worker threads and `--cpus=2` has 200 ms of CPU per 100 ms of wall clock; if a burst of requests makes all 8 threads runnable, they can exhaust it in about 25 ms and then sit throttled for 75 ms. The reported average CPU may be well under the limit over a one-minute window, yet every request that lands in a throttled stretch pays tens of milliseconds of pure waiting. The signature is bimodal latency: normal p50, a p99 with a hard floor near the throttle duration, and a workload that gets *worse* as you add threads. Garbage collection makes it sharper: a stop-the-world phase across many GC threads can blow the whole period's quota at once, so the pause is inflated by the throttle. ## Proving it Read cpu.stat in the container's cgroup: `nr_periods`, `nr_throttled` (periods in which the group was throttled) and `throttled_usec` (total time spent throttled). A non-trivial nr_throttled / nr_periods ratio is the evidence; time-series it and correlate with latency. Platform metrics expose the same counters under different names; the interpretation is identical wherever you read them from. ## Remedies 1. **Raise the quota** to cover bursts. Throttling is caused by the burst-to-quota ratio, not by average load, so a limit sized to average usage will throttle constantly. 2. **Match concurrency to the quota.** Thread pools, GC threads and connection pools sized from the host's core count are the usual culprits; a container with a 1-CPU quota on a 64-core host may build a 64-thread pool. Modern runtimes read the cgroup quota (the JVM's available-processors derives from it), but library defaults and explicit configuration often do not - set them from the quota. 3. **Use weights instead of quotas** for latency-critical services when the goal is fairness rather than a billing-style cap. Weights protect against noisy neighbours without introducing artificial stalls, at the cost of less predictable capacity. 4. **Increase burst tolerance** where supported: newer kernels expose cpu.max.burst, letting a group bank unused quota to absorb short spikes. 5. **Consider cpuset pinning** for the small set of workloads that need deterministic cache and NUMA behaviour, keeping in mind it is a placement tool, not a limit. ## Choosing Batch and untrusted workloads: quota, so they cannot take more than they paid for. Latency-sensitive services on dedicated capacity: generous quota (or weights) plus right-sized concurrency. Everything: measure nr_throttled before and after, because the flag you set and the behaviour you get are only connected through the burst pattern of the real workload.

  • A service has --cpus=1 and its p99 latency is terrible, but average CPU utilization is only 40%. What do you check first?
    cpu.stat in the container's cgroup, specifically nr_throttled and throttled_usec. Bursty multi-threaded work can exhaust a 100 ms quota in the first few tens of milliseconds and then stall until the period rolls over, which averages out to modest utilization but adds fixed stalls to tail latency. If throttling is confirmed, raise the quota or reduce concurrency so the burst fits inside the period.
  • Why can a thread pool sized from the host's core count be a problem in a quota-limited container?
    The pool size determines how fast the group burns its quota: many runnable threads consume the entire period's budget almost immediately and then all of them are throttled together. Older runtimes and many libraries read the host's CPU count rather than the cgroup quota, so a 1-CPU container on a large host builds a huge pool and throttles constantly. Size pools, GC threads and parallel streams from the quota instead.

saying these in an interview costs you the question

  • Believing --cpu-shares caps CPU usage
  • Thinking --cpus pins the container to a number of specific cores
  • Assuming a container cannot be throttled while the host has idle CPUs
  • Sizing the quota to average utilization and being surprised by tail latency
  • Ignoring that library and runtime thread pools may be sized from host core count

context

open as a page

What are Linux control groups (cgroups), and what actually happens in the kernel when a container is started with `--memory=512m --cpus=1.5`?

level: middleimportance: must knowfreq 58%

basics

~20 s

cgroups are the kernel feature that limits and accounts resources for a group of processes. The runtime creates a cgroup per container and writes the flags into its control files: on cgroup v2, memory.max becomes 536870912 and cpu.max becomes '150000 100000' (150 ms of CPU per 100 ms period).

open as a page

A container exits with code 137 in production. How do you confirm it was a cgroup memory kill rather than something else, and how do you find the cause?

level: seniorimportance: must knowfreq 54%

basics

~20 s

137 means the process died from SIGKILL (128+9), which can be a cgroup OOM kill or an external kill such as a stop timeout. Confirm with docker inspect .State.OOMKilled, the kernel log line 'Memory cgroup out of memory', and the cgroup's memory.events oom_kill counter, then compare peak usage against memory.max.

open as a page

What changed between cgroup v1 and cgroup v2 on Linux, and how does that affect container resource limits?

level: seniorimportance: should knowfreq 38%

basics

~20 s

v1 had a separate hierarchy per controller with inconsistent interfaces; v2 has one unified hierarchy, one consistent API (memory.max, cpu.max, io.max), pressure metrics, and coordinated memory-plus-I/O accounting so buffered writes are attributed correctly. Some v1-only options, such as disabling the OOM killer, no longer exist.

open as a page

How would you choose memory and CPU limits for a latency-sensitive containerized service, given that exceeding the memory limit kills the process while exceeding the CPU quota only delays it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Treat the two asymmetrically. Memory is a cliff: size it from observed peak plus headroom and cap the runtime below it so failures are diagnosable. CPU is a slope: quota only delays work, so set it above burst demand, size thread pools to it, and watch throttling counters instead of shaving it to average usage.

open as a page