Explain the difference between the container options --cpus, --cpu-shares and --cpuset-cpus, and why an application can be slow while its CPU usage looks low.
answer
- --cpus = hard CFS quota per 100 ms period (cpu.max)
- shares/weight = relative, only under contention, never throttles
- cpuset = pinning, not metering
- burst exhausts quota early -> stall until next period
- proof: cpu.stat nr_throttled / throttled_usec
basics
~20 s--cpus is a hard quota per scheduling period (cpu.max), --cpu-shares is a relative weight that only matters under contention (cpu.weight), and --cpuset-cpus pins the container to specific cores. A container that burns its quota early is throttled until the next period, so latency spikes while average utilization looks modest.
solid answer
~50 s**--cpus=N** sets CFS bandwidth control: quota N x period, written to cpu.max, default period 100 ms. It is an absolute ceiling and applies even when the host is idle. **--cpu-shares** (cpu.weight on cgroup v2) is proportional: it decides how contended CPU time is divided between cgroups and has no effect when there is spare capacity. It cannot cause throttling. **--cpuset-cpus=0-3** pins the container's tasks to specific CPUs. That bounds parallelism and helps cache locality and NUMA placement, but it meters nothing - two pinned containers still contend. The slow-but-idle-looking symptom is CFS throttling. A multi-threaded process with `--cpus=1` can consume its 100 ms of quota in the first 25 ms of a 100 ms period across four runnable threads, then every thread is descheduled for 75 ms. Average utilization reads modest over a coarse window, while p99 latency shows fixed-size stalls. Evidence is in cpu.stat: nr_throttled and throttled_usec climbing. Fixes: raise the quota, right-size thread and GC pools to the quota, or prefer weights for latency-critical services.
code
bash · 6 linesdocker run -d --cpus=1.5 app # cpu.max = 150000 100000 (hard ceiling)
docker run -d --cpu-shares=2048 app # relative weight, only under contention
docker run -d --cpuset-cpus=0-3 app # pinned to CPUs 0-3
docker exec app cat /sys/fs/cgroup/cpu.stat
# nr_periods 5921 nr_throttled 1840 throttled_usec 137004221go deeper
Distinguish the three flags in one sentence each and know that --cpus is the hard limit.
Explain quota-and-period mechanics, why bursts cause throttling, and where to find nr_throttled.
Diagnose tail-latency incidents from throttling counters, right-size concurrency to the quota, and argue quota versus weight per workload class.
Set fleet policy on limits versus weights, node density and overcommit, and the tradeoff between predictable isolation and utilization.
## Three different mechanisms **Bandwidth (--cpus, or --cpu-quota/--cpu-period).** The Completely Fair Scheduler's bandwidth control gives a cgroup a quota of CPU time per period. cpu.max holds `"<quota> <period>"` in microseconds; `--cpus=2` is `200000 100000`. The group may run on any CPU, but the sum of the time its tasks consume within one period cannot exceed the quota. When it does, all its tasks are throttled - made unrunnable - until the period rolls over. This is a hard ceiling: an idle 64-core host does not let a `--cpus=1` container use more. **Weight (--cpu-shares, cpu.weight).** A relative number used only when runnable tasks compete for the same CPU. Two containers with weights 100 and 200 get roughly a 1:2 split *while both are busy and the CPU is saturated*; if one is idle, the other uses everything. Weights never throttle and never waste idle capacity. Docker converts the legacy shares value into a v2 cpu.weight, so the number in the file will not match the flag. **Affinity (--cpuset-cpus, --cpuset-mems).** Restricts which CPUs (and NUMA memory nodes) the tasks may run on. It caps parallelism by core count and improves cache and NUMA locality, which matters for latency-sensitive or memory-bandwidth-bound workloads. It is not accounting: several containers pinned to the same cores still fight for them. ## Why throttling looks like "slow but idle" Quota is consumed by *all* runnable threads in aggregate, but the period is fixed at 100 ms. A service with 8 worker threads and `--cpus=2` has 200 ms of CPU per 100 ms of wall clock; if a burst of requests makes all 8 threads runnable, they can exhaust it in about 25 ms and then sit throttled for 75 ms. The reported average CPU may be well under the limit over a one-minute window, yet every request that lands in a throttled stretch pays tens of milliseconds of pure waiting. The signature is bimodal latency: normal p50, a p99 with a hard floor near the throttle duration, and a workload that gets *worse* as you add threads. Garbage collection makes it sharper: a stop-the-world phase across many GC threads can blow the whole period's quota at once, so the pause is inflated by the throttle. ## Proving it Read cpu.stat in the container's cgroup: `nr_periods`, `nr_throttled` (periods in which the group was throttled) and `throttled_usec` (total time spent throttled). A non-trivial nr_throttled / nr_periods ratio is the evidence; time-series it and correlate with latency. Platform metrics expose the same counters under different names; the interpretation is identical wherever you read them from. ## Remedies 1. **Raise the quota** to cover bursts. Throttling is caused by the burst-to-quota ratio, not by average load, so a limit sized to average usage will throttle constantly. 2. **Match concurrency to the quota.** Thread pools, GC threads and connection pools sized from the host's core count are the usual culprits; a container with a 1-CPU quota on a 64-core host may build a 64-thread pool. Modern runtimes read the cgroup quota (the JVM's available-processors derives from it), but library defaults and explicit configuration often do not - set them from the quota. 3. **Use weights instead of quotas** for latency-critical services when the goal is fairness rather than a billing-style cap. Weights protect against noisy neighbours without introducing artificial stalls, at the cost of less predictable capacity. 4. **Increase burst tolerance** where supported: newer kernels expose cpu.max.burst, letting a group bank unused quota to absorb short spikes. 5. **Consider cpuset pinning** for the small set of workloads that need deterministic cache and NUMA behaviour, keeping in mind it is a placement tool, not a limit. ## Choosing Batch and untrusted workloads: quota, so they cannot take more than they paid for. Latency-sensitive services on dedicated capacity: generous quota (or weights) plus right-sized concurrency. Everything: measure nr_throttled before and after, because the flag you set and the behaviour you get are only connected through the burst pattern of the real workload.
- A service has --cpus=1 and its p99 latency is terrible, but average CPU utilization is only 40%. What do you check first?cpu.stat in the container's cgroup, specifically nr_throttled and throttled_usec. Bursty multi-threaded work can exhaust a 100 ms quota in the first few tens of milliseconds and then stall until the period rolls over, which averages out to modest utilization but adds fixed stalls to tail latency. If throttling is confirmed, raise the quota or reduce concurrency so the burst fits inside the period.
- Why can a thread pool sized from the host's core count be a problem in a quota-limited container?The pool size determines how fast the group burns its quota: many runnable threads consume the entire period's budget almost immediately and then all of them are throttled together. Older runtimes and many libraries read the host's CPU count rather than the cgroup quota, so a 1-CPU container on a large host builds a huge pool and throttles constantly. Size pools, GC threads and parallel streams from the quota instead.
saying these in an interview costs you the question
- Believing --cpu-shares caps CPU usage
- Thinking --cpus pins the container to a number of specific cores
- Assuming a container cannot be throttled while the host has idle CPUs
- Sizing the quota to average utilization and being surprised by tail latency
- Ignoring that library and runtime thread pools may be sized from host core count