skip to content

questions

4

Compare Docker's `--cpus`, `--cpu-shares` and `--cpu-quota`/`--cpu-period` flags: which actually caps a container's CPU usage, which only matters under contention, and how would you tell that a container is being CPU-throttled?

level: middleimportance: must knowfreq 62%

answer

  1. shares = weight, only under contention
  2. --cpus = quota/period sugar (CFS bandwidth)
  3. quota burns in parallel => stall to end of period
  4. cpu.stat: nr_throttled / throttled_usec
  5. cpuset = pinning, not proportion

basics

~20 s

--cpus=N is a hard cap — shorthand for --cpu-quota/--cpu-period (CFS bandwidth), e.g. 150000/100000 for 1.5 CPUs. --cpu-shares is only a relative weight that applies when the host is contended; idle CPU is still free to take. Throttling appears as nr_throttled/throttled_usec in the cgroup's cpu.stat.

solid answer

~50 s

**`--cpu-shares`** (default 1024) sets a *relative weight*. It does nothing on an idle host — a 512-share container will happily use every core available. It matters only when runnable containers compete, deciding the split ratio. Good for prioritisation, useless as a guarantee or a cap. **`--cpu-quota` with `--cpu-period`** is Linux CFS bandwidth control: the container may consume `quota` microseconds of CPU time in every `period` (default 100 ms). Exhaust the quota early and every thread is stopped until the next period opens. **`--cpus=1.5`** is simply a wrapper that writes quota=150000, period=100000. It is the flag you normally use. `--cpuset-cpus=0-3` is different again: it pins the container to specific cores, which matters for NUMA and cache locality rather than proportional sharing. Throttling is invisible in `docker stats` — the container just sits at its cap. Read `cpu.stat` in its cgroup: rising `nr_throttled` and `throttled_usec` mean real stalls, and a multi-threaded service can burn a whole period's quota in milliseconds and then freeze, wrecking tail latency.

code

bash · 9 lines
bash
docker run -d --cpus=1.5 myapp:1.0
docker run -d --cpu-period=100000 --cpu-quota=150000 myapp:1.0

docker run -d --cpu-shares=512 batch-job:1.0

docker exec myapp cat /sys/fs/cgroup/cpu.stat
# nr_periods 41230
# nr_throttled 9877
# throttled_usec 812440000

go deeper

for a junior

Know --cpus=N is the hard cap you normally set, and that CPU pressure throttles rather than kills.

for a middle

Explain shares-as-weight versus quota-as-bandwidth, the quota/period arithmetic behind --cpus, and where throttling counters live.

for a senior

Diagnose latency caused by per-period stalls, tune quota, period and worker counts, and explain how the quota reconfigures runtime parallelism such as GC thread counts.

for a principal

Set fleet policy: caps everywhere to bound blast radius, weights to express workload priority, pinning reserved for latency-critical NUMA-sensitive services, with utilisation targets chosen against throttling risk.

## Two different ideas CPU controls in Docker come in two families that candidates constantly conflate: **proportional weight** (who wins when there is contention) and **hard bandwidth** (an absolute ceiling regardless of idle capacity). Getting that distinction right is the whole question. ## Weight: `--cpu-shares` `--cpu-shares` maps to the cgroup CPU weight (`cpu.shares` on v1, `cpu.weight` on v2). The default is 1024 and only the *ratio* matters: two containers at 1024 and 512, both runnable and competing for the same CPU, get roughly a 2:1 split. The critical property is that shares are meaningless when the CPU is not saturated. A container with 100 shares on an idle 16-core host can use all 16 cores. Shares are therefore a priority mechanism — never a guarantee, never an isolation boundary, never something you can bill against. They are the right tool when a batch job should yield to a latency-sensitive service without being capped when the box is quiet. ## Bandwidth: `--cpu-quota` and `--cpu-period` The Completely Fair Scheduler's bandwidth control gives each cgroup a budget of CPU time per interval. `--cpu-period=100000` (microseconds — the default 100 ms) with `--cpu-quota=50000` means: across all its threads, the container may consume 50 ms of CPU time in each 100 ms window, i.e. half a core. When the budget is spent, every runnable thread in the cgroup is de-scheduled until the next period begins. This is a genuine cap: it applies even on a completely idle host. `--cpus=0.5` is exactly this pair expressed once, and is the flag you should normally reach for; `--cpus=2` means quota 200000 per 100000 period — two cores' worth of time, spendable in parallel across many cores. ## Why quotas hurt latency The subtlety is that quota is consumed *in parallel*. A service with 20 threads under `--cpus=1` can burn its entire 100 ms budget in roughly 5 ms of wall-clock time if all threads run at once — and then the whole container is frozen for the remaining ~95 ms. Average CPU usage looks like a comfortable 1.0 while p99 latency contains ~95 ms of pure stall. This is the classic "my container is only at 40% CPU but requests are slow" report. Mitigations: raise the quota so bursts fit; reduce thread or worker counts so parallelism matches the quota; shorten the period (`--cpu-period=10000`) so stalls are smaller though more frequent; or drop the hard cap for latency-critical services and rely on weights plus capacity planning. ## Detecting throttling `docker stats` will not tell you. A throttled container simply reports usage at its ceiling and looks fine. The evidence lives in the cgroup: `cpu.stat` exposes `nr_periods`, `nr_throttled` and `throttled_usec` (on cgroup v1, `throttled_time`). A steadily climbing `nr_throttled` as a fraction of `nr_periods` — anything sustained above a few percent for a latency-sensitive service — means the cap is actively hurting. Read it via `docker exec` or from the host's cgroup tree. ## Pinning: `--cpuset-cpus` `--cpuset-cpus=0-3` restricts the container to those specific CPUs. This is neither proportional nor a time budget — it is placement. It helps with NUMA locality, cache warmth for pinned workloads, and keeping noisy neighbours off cores reserved for a critical service. It also implicitly caps parallelism at the number of assigned cores. ## Interaction with runtimes Language runtimes read the quota to size their internal parallelism. A JVM derives its active processor count from quota divided by period, which drives GC thread counts, JIT compiler threads, the common ForkJoinPool and any library calling `availableProcessors()`; recent Go runtimes behave similarly, while older versions read host core count and over-schedule. Setting `--cpus=0.5` therefore does not merely slow a runtime, it reconfigures it — typically down to a single reported "processor". ## Choosing A workable default: set `--cpus` on everything so no container can monopolise a host, size it above measured peak burst rather than average, and use shares only to express priority between co-located classes of work. Reserve `--cpuset-cpus` for genuinely latency-critical, NUMA-sensitive services. And whenever someone reports "slow but low CPU", check `cpu.stat` before anything else.

  • A service averages 45% of its CPU limit but has terrible p99 latency. What do you suspect?
    CFS throttling. Many threads can drain the whole per-period quota in a few milliseconds, after which the entire cgroup is frozen until the next period, so averages stay low while individual requests absorb tens of milliseconds of stall. Confirm with `nr_throttled` and `throttled_usec` in `cpu.stat`, then raise the quota, cut worker parallelism, or shorten the period.
  • Two containers have 1024 and 512 shares and the host is idle. How much CPU does each get?
    As much as it can use — shares impose no ceiling. The 512-share container can saturate every core while the other sits idle. The 2:1 ratio takes effect only when both are runnable and competing for the same CPUs.

Shares are seniority in a queue — irrelevant when nobody else is waiting. A quota is a prepaid data plan per 100 ms: spend it in five milliseconds and you are cut off until the clock rolls over.

saying these in an interview costs you the question

  • Describing `--cpu-shares` as a cap or a guarantee
  • Not knowing `--cpus` is shorthand for quota over period
  • Believing a CPU-limited container gets killed the way a memory-limited one does
  • Expecting `docker stats` to reveal throttling
  • Confusing `--cpuset-cpus` pinning with proportional sharing

context

open as a page

What actually happens when a process in a Docker container allocates past the container's `--memory` limit, and what does the `--memory-swap` flag change about that behavior?

level: middleimportance: must knowfreq 70%

basics

~20 s

The kernel first reclaims what it can; if that is not enough, the cgroup OOM killer SIGKILLs a process inside that container only. If the victim is PID 1 the container dies with exit code 137 and State.OOMKilled=true. --memory-swap sets the combined memory+swap ceiling; setting it equal to --memory disables swap.

open as a page

A JVM service runs in a container capped at 512 MB and keeps getting killed by the kernel, even though its heap usage looks small and it never throws `java.lang.OutOfMemoryError`. How does a modern JVM discover container CPU and memory limits, and how would you size it so this stops happening?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Modern JVMs read cgroup limits (container support is default-on since JDK 10, backported to 8u191) and default the heap to about 25% of the container limit. The kill comes from total process memory: heap plus metaspace, thread stacks, code cache, direct buffers and native allocations. Size with -XX:MaxRAMPercentage and leave real non-heap headroom.

open as a page

How do you check what CPU and memory running Docker containers are using right now, and what does each column of the `docker stats` output mean?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Run docker stats. It live-streams per-container CPU % (summed across cores, so it can exceed 100%), MEM USAGE / LIMIT and MEM %, NET I/O, BLOCK I/O and PIDS, read from each container's kernel cgroup counters. Add --no-stream for a single snapshot.

open as a page