Why does `nproc` in a Docker container report the host's core count when the container was run with `--cpus`?
answer
- Two flags restrict CPU in different ways
- One caps time, the other caps where
- nproc follows the scheduling affinity mask
- Quota does not narrow that mask
- cpuset does; /proc/cpuinfo never does
basics
~20 snproc reports the CPUs in the process's scheduling affinity mask. --cpus limits CPU bandwidth through the cgroup and leaves affinity untouched, so all host CPUs remain visible. --cpuset-cpus does change affinity, so nproc reflects that one.
solid answer
~50 sThere are two different ways to restrict CPU on a container and they are visible to different things. `--cpus` sets a cgroup bandwidth quota: the container may use that much CPU time per period, but it is still allowed to run on every core, so its scheduling affinity mask is unchanged and `nproc` returns the host count. `--cpuset-cpus` pins the container to named CPUs by narrowing the affinity mask, and `nproc` then returns the smaller number. `/proc/cpuinfo` always lists every host CPU regardless, since procfs is not cgroup-aware. This matters because anything that auto-sizes from the observed core count - worker pools, build parallelism such as `make -j$(nproc)`, library thread pools - will provision for the host and oversubscribe a quota-limited container, which shows up as heavy throttling and context switching. Pass the intended parallelism explicitly instead.
code
bash · 4 lines# 64-core host
docker run --rm --cpus=2 alpine nproc # prints 64
docker run --rm --cpuset-cpus=0,1 alpine nproc # prints 2
docker run --rm --cpus=2 alpine nproc --all # prints 64go deeper
Remember that a CPU limit set with --cpus does not change how many CPUs the container can see, and that build or worker parallelism should be given as a setting rather than read from nproc.
Be ready to contrast bandwidth limiting with affinity pinning, explain that nproc follows the affinity mask while /proc/cpuinfo is host-wide, and predict what each docker run flag makes nproc print.
Show the incident pattern: auto-sized parallelism against a quota produces throttling, memory pressure and wildly variable durations. Demonstrate confirming it with the cgroup throttling counters before and after fixing the configuration.
Own the standard across the fleet. Decide that effective CPU is injected as configuration rather than discovered, judge whether pinning is worth its utilisation cost, and treat a compatibility shim as a bounded exception with a named owner.
### Two different limits, two different observers Docker offers two unrelated mechanisms for constraining CPU, and confusing them is what produces this question. `--cpus` is a **bandwidth** limit. It configures the CPU cgroup controller so the container's tasks may consume a fixed amount of CPU time in each scheduling period. Nothing stops those tasks running on any core; they are simply stopped once the budget is spent. `--cpuset-cpus` is an **affinity** limit. It restricts which CPUs the container's tasks may be scheduled on at all, by narrowing the affinity mask the kernel applies to them. `nproc` from GNU coreutils reports the number of processing units available to the calling process, which it determines from the process's scheduling affinity. So `--cpuset-cpus=0,1` gives `nproc` a value of 2, while `--cpus=2` on a 64-core host still gives 64. (`nproc --all` deliberately ignores affinity and reports the installed total in both cases.) `/proc/cpuinfo` is a host-wide procfs file with no cgroup awareness, so it always lists all 64 - which is why counting its entries is an even worse way to size anything. ### Why anyone cares The count matters because so much software auto-sizes from it: worker pools, parser and compression thread pools, connection pools, and above all build parallelism. Consider a 47-node CI fleet where each node has 64 cores and jobs run in containers created with `--cpus=2` so that many builds can share a node. A build step of `make -j$(nproc)` inside such a container launches 64 compiler processes against two cores' worth of quota. The result is not a clean slowdown: memory use scales with the number of concurrent compilers, so the job is far more likely to hit its memory limit and be killed; the CPU cgroup is throttled in nearly every period; and the wall-clock time of the step gets worse, not better, because of context switching and cache pressure. The fleet's symptom was a step whose duration varied between 4 and 26 minutes depending on how many jobs happened to share the node - variance without an obvious cause, which is the signature of oversubscription. The same shape appears outside CI in any runtime whose default pool sizing comes from the observed processor count. A modern container-aware runtime avoids this by reading the cgroup CPU quota rather than the affinity mask or procfs, and deriving its own effective processor count from it. Whether a given library does that is worth checking rather than assuming. ### The three fixes, in order of preference 1. **Pass the number explicitly.** Configure parallelism from the same source of truth that set the limit - the deployment or job definition - rather than letting the process guess. `make -j4`, an explicit pool-size setting, an environment variable read at startup. It is the only fix that works regardless of runtime, library or kernel behaviour, and it makes the intended concurrency reviewable in configuration. 2. **Use `--cpuset-cpus` where pinning is acceptable.** Narrowing affinity makes the observed count correct for free, and can improve cache locality. The costs are real: pinned containers cannot use idle capacity elsewhere on the machine, and someone has to assign non-overlapping CPU sets, which becomes an allocation problem of its own as density rises. 3. **Mount lxcfs.** This FUSE filesystem synthesises cgroup-aware versions of `/proc/cpuinfo`, `/proc/meminfo`, `/proc/stat` and friends and bind-mounts them over the container's real files, so legacy tools see container-shaped numbers. It does not change the affinity mask, so it corrects tools that read procfs while `nproc` itself still follows affinity. It suits system-container images carrying agents you cannot modify, and it adds a host daemon and a FUSE hop that application containers rarely need. ### Closing the loop After changing parallelism, verify with the counter rather than with a feeling: diff the container's cgroup `cpu.stat` across a run and check that the throttled share of periods has fallen. Oversubscription and a genuinely undersized limit produce the same throttling signature, and only the experiment distinguishes them. State the rule plainly in review: inside a container, the number of CPUs you may use is a configuration input, never something to be discovered by asking the machine.
- If `--cpuset-cpus` makes the count correct, why not use it everywhere instead of `--cpus`?Pinning trades flexibility for accuracy. A container restricted to two named CPUs cannot borrow idle capacity elsewhere on the host, so utilisation drops, and someone must allocate non-overlapping sets as density grows, which is scheduling work you have taken on manually. A bandwidth quota lets the kernel place work wherever there is room. Use cpuset when locality or isolation genuinely matters, not as a reporting fix.
- How would you confirm that oversubscription, rather than a limit set too low, caused a slow build step?Diff the container's cgroup cpu.stat across the step, then repeat the run with the parallelism set explicitly to match the quota. If the throttled share of periods collapses and wall-clock time improves at the same limit, the process was creating more runnable work than the quota could feed. If throttling stays high with modest parallelism, the limit itself is too small.
- Does mounting lxcfs make `nproc` report the reduced count?Not by itself. lxcfs replaces procfs files such as /proc/cpuinfo and /proc/meminfo with cgroup-aware versions, which fixes tools that read those files. nproc asks the kernel for the process's scheduling affinity, which lxcfs does not change, so a bandwidth-limited container still reports the host count there. That gap is one reason to configure parallelism explicitly rather than to rely on a shim.
A bandwidth quota is a monthly spending limit and affinity is the list of shops you are allowed to enter. Counting the shops on the high street tells you nothing about the limit on your card.
saying these in an interview costs you the question
- Believes --cpus hides host CPUs from the container
- Thinks nproc counts the lines in /proc/cpuinfo
- Uses make -j$(nproc) inside a quota-limited container
- Says the runtime always detects the container limit
- Assumes lxcfs corrects nproc as well as procfs files
- Treats resulting throttling as proof the limit is too small