What does os.cpu_count() report, and why is it the wrong number inside a container?
answer
- It measures the machine, not your slice
- The kernel throttles; it hides nothing
- A second function honours the affinity mask
- The quota lives outside the os module
basics
~10 sos.cpu_count() reports the logical processors the host kernel exposes. A container's CPU limit is enforced as scheduling bandwidth rather than by hiding processors, so the count still says 64 on a two-CPU budget.
solid answer
~40 s`os.cpu_count()` returns the number of logical CPUs on the **machine** — cores times hardware threads — or `None` when the platform cannot determine it. A container CPU limit is a control-group quota: the kernel grants the group so much CPU time per period and stops its threads once that is spent. It does not remove processors from view, so the count is unchanged and code that sizes a process pool from it starts sixty-four workers on a two-CPU budget. Since **Python 3.13** there is a second function, `os.process_cpu_count()`, which reports the CPUs the calling thread may actually run on, honouring a CPU affinity mask; it still does not see a quota. The safe number is the smaller of that count and the quota, floored at 1.
code
python · 14 linesimport os
from pathlib import Path
host = os.cpu_count()
usable = os.process_cpu_count()
quota = None
cpu_max = Path("/sys/fs/cgroup/cpu.max")
if cpu_max.exists():
limit, period = cpu_max.read_text().split()
if limit != "max":
quota = int(limit) / int(period)
print(host, usable, quota)go deeper
Know that os.cpu_count() reports the machine's logical processors and that this is the number most tutorials feed to a pool. Be able to say out loud that a container does not change it.
Explain the mechanism: a CPU limit is a time quota enforced by the scheduler, not a smaller set of processors, so nothing in the os module notices it. Name os.process_cpu_count() and say what extra thing it honours.
Show how you would derive a usable number in production — reconcile the host count, the affinity-aware count and the control-group quota, floor it at 1, and describe the throttling sawtooth an oversubscribed pool produces in the latency tail.
Own where that number comes from across a fleet: whether it is injected as PYTHON_CPU_COUNT at deploy time from the same field that sets the limit, or computed by a shared helper, and what each choice costs in drift and debuggability.
### The number `os.cpu_count()` reports `os.cpu_count()` asks the operating system how many *logical* processors the machine has: physical cores multiplied by hardware threads, summed across sockets. It returns an `int`, or `None` when the platform cannot work it out — which is why careful code writes `os.cpu_count() or 1` rather than doing arithmetic on the result directly. What the function never reports is how much of that machine the calling process is entitled to consume. That distinction is invisible on a laptop, where the two questions have the same answer, and it is the whole story on a scheduler-managed host, where they do not. ### A CPU limit is bandwidth, not a smaller machine When a process runs inside a control group with a CPU limit, the kernel does not build it a smaller computer. It gives the group a quota and a period — say 200 ms of CPU time out of every 100 ms of wall clock, which is "two CPUs' worth". Threads in that group are still scheduled onto any processor on the host; when the group has burned its quota for the current period, every runnable thread in it is stopped until the next period starts, then all of them resume at once. Because nothing is hidden, everything that enumerates processors keeps reporting the host: the kernel's own CPU listing, and therefore `os.cpu_count()`. On a 64-way host it returns 64 whether the container was granted thirty-two CPUs or a tenth of one. There is no exception, no warning and no `None` — the number is simply answering a different question than the one the caller meant to ask. ### What that costs when a pool is sized from it The damage is done by defaults. In Python 3.14, `concurrent.futures.ProcessPoolExecutor()` with no `max_workers` uses `os.process_cpu_count()`, and `multiprocessing.Pool()` with no argument does the same; `concurrent.futures.ThreadPoolExecutor()` uses `min(32, cpu + 4)`. Before **3.13** all of those defaults came from `os.cpu_count()`. Either way, none of them consults a quota, so a service deployed with a two-CPU budget onto a large host cheerfully starts dozens of worker processes. The result is not a crash but a slow, confusing degradation. Every worker is a full interpreter with its own memory. All of them are runnable, so the group burns its quota early in each period and then every worker is stopped for the remainder — latency arrives in a sawtooth, and the tail of the latency distribution gets far worse while median throughput barely improves. Background work sharing the process, such as a periodic refresh thread, is throttled along with everything else and can miss its schedule entirely. ### Getting a number that means something There are three sources to reconcile: 1. **`os.cpu_count()`** — the host's logical processors. Useful for reporting hardware, almost never for sizing. 2. **`os.process_cpu_count()`** — added in **3.13**; the CPUs the calling thread may run on, which on Linux reflects the process's CPU affinity mask, as set by a pinning tool. Correct when a scheduler pins work to a subset of processors; still blind to a quota. 3. **The control-group quota** — readable from the group's `cpu.max` file on cgroup v2 (a limit and a period, or the word `max` when unlimited), or the older `cpu.cfs_quota_us` and `cpu.cfs_period_us` pair on v1. Take the smallest of whichever apply and floor it at 1, since a fractional grant such as half a CPU must still yield one worker. Two extras help: **3.13** added the `-X cpu_count` interpreter option and the `PYTHON_CPU_COUNT` environment variable, which override the value *both* functions return — set that from the same manifest field that sets the CPU limit and every stdlib default lands on the right number without touching application code. ### The trap in the other direction A host that is not containerized can lie the other way. The count is also unhelpfully large on a machine that has been narrowed by pinning: a batch scheduler that assigns a job four processors out of sixty-four leaves `os.cpu_count()` reporting sixty-four, and workers beyond the fourth can never be scheduled anywhere useful at all. That case at least has a standard-library answer, since `os.process_cpu_count()` reflects it. `os.cpu_count()` on a shared build machine reports every processor even though a dozen other jobs are competing for them; the process gets its share, not the machine. The honest framing is that `os.cpu_count()` answers "how big is this computer", and pool sizing needs the answer to "how much of it is mine".
- When does os.cpu_count() return None, and how should calling code handle it?It returns `None` when the platform cannot determine the count — rare, but real on unusual builds and embedded targets. `os.process_cpu_count()` can return `None` for the same reason. Any code doing arithmetic on the result must write `os.cpu_count() or 1`, or it will hit a `TypeError` in exactly the environment nobody tests on. Treating an unknown count as 1 is the conservative choice: it under-parallelizes instead of oversubscribing.
- What actually goes wrong when sixty-four worker processes run under a two-CPU quota?Throughput does not improve, because the quota, not the worker count, is the ceiling. Each period the group burns its allowance early and then every worker is stopped until the period rolls over, which turns a smooth latency curve into a sawtooth and wrecks the tail. Each worker is also a full interpreter with its own memory, so resident size scales with the wrong number, and the extra context switching costs real CPU out of the same allowance.
- Does the same problem affect thread pools, or only process pools?It affects both, differently. `concurrent.futures.ThreadPoolExecutor()` defaults to `min(32, cpu + 4)`, so the over-count is capped at 32 and threads are far cheaper than processes — but the quota still applies to the whole group, so extra threads add contention and context switches without adding CPU time. The memory blow-up is specific to processes; the throttling is common to both.
It is like reading the capacity of a motorway and concluding you may drive all the lanes: the road really is eight lanes wide, but your permit is for two, and the enforcement happens after you set off.
saying these in an interview costs you the question
- Believes a container CPU limit shrinks os.cpu_count()
- Thinks os.cpu_count() reports physical cores, not logical CPUs
- Sizes a process pool from the host count and expects a speedup
- Confuses a CPU quota with a limit on how many threads may exist
- Assumes os.cpu_count() reads a cgroup file
- Ignores that os.cpu_count() can return None