skip to content

Pool Sizing and Oversubscription

How many workers a pool should really have: the core count this process may use rather than the machine's, queue depth on an I/O pool, and nested pools inside libraries multiplying the total.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why is os.cpu_count() the wrong number to size a worker pool, and what does os.process_cpu_count() report instead?

level: middleimportance: must knowfreq 45%

answer

  1. The machine's cores are not your cores
  2. Two similar os functions, two questions
  3. Permission to run, not hardware inventory
  4. An affinity mask narrows one of them
  5. os.process_cpu_count versus os.cpu_count

basics

~20 s

os.cpu_count() reports every logical CPU on the machine, whether or not this process may use them. os.process_cpu_count(), added in Python 3.13, reports how many the calling thread is actually allowed to run on, which is the number to size from.

solid answer

~50 s

`os.cpu_count()` answers a hardware question: how many logical CPUs the machine has. A process restricted to a subset of them by a CPU affinity mask still only runs on that subset, so sizing a pool from the machine total oversubscribes the cores you actually own and buys context switches instead of throughput. `os.process_cpu_count()`, new in **Python 3.13**, returns how many logical CPUs the calling thread may run on, and it is what `ProcessPoolExecutor` and `multiprocessing.Pool` default to and what `ThreadPoolExecutor` derives its `min(32, n + 4)` default from since 3.13. Either function can return `None` when the platform cannot tell, so write `(os.process_cpu_count() or 1)`. Neither reflects a fractional CPU *quota* — a time-slice allowance shows up as throttling, not as a smaller count — which is why the interpreter also takes an explicit override.

code

python · 9 lines
python
import os

machine = os.cpu_count()
usable = os.process_cpu_count() or 1

print("logical CPUs on this machine:", machine)
print("logical CPUs this thread may use:", usable)
print("ProcessPoolExecutor default workers:", usable)
print("ThreadPoolExecutor default workers:", min(32, usable + 4))

go deeper

for a junior

Know that Python has two different CPU counts and that the bigger one describes the machine, not your process. Be able to say why running far more CPU-bound workers than cores does not make the work finish sooner.

for a middle

Be ready to explain the mechanics: os.process_cpu_count() reflects the calling thread's CPU affinity, it was added in 3.13, the executor and multiprocessing defaults moved to it in the same release, and either function may return None so sizing code needs an or 1.

for a senior

An interviewer expects you to show where detection breaks in production: a time-slice quota does not shrink either count, so you allow an explicit override per pool, log the number the process actually chose, and verify it on the real deployment target rather than on a laptop.

for a principal

Own the policy angle. Decide whether services detect their own width or are told it once at process start via PYTHON_CPU_COUNT, and defend that choice against the cost of every library in the process sizing itself independently from a number that is wrong.

### Two different questions `os.cpu_count()` and `os.process_cpu_count()` look interchangeable and are not. The first is an inventory question — *how many logical CPUs does this machine have?* — and it answers with the whole box, including CPUs this process will never be scheduled onto. The second, added in **Python 3.13**, is a permission question: *how many logical CPUs may the calling thread actually run on?* On Linux that is the size of the thread's CPU affinity mask, which a launcher, a batch scheduler or an orchestration layer may have narrowed. On a 64-CPU host where the process has been pinned to four of them, `os.cpu_count()` says 64 and `os.process_cpu_count()` says 4. Both can return `None`: the count is genuinely undeterminable on some platforms, and code that writes `os.process_cpu_count() + 4` will then raise `TypeError` at start-up on exactly the machine you cannot debug. The idiom is `(os.process_cpu_count() or 1)`. ### Why the difference bites when sizing a pool For CPU-bound work the whole point of a pool is to keep every core you own busy and no more. Threads or processes beyond that count do not add throughput — they add context switches, cache pressure, and for processes, memory. Sizing 64 workers onto a 4-CPU allowance means each unit of work is interleaved sixteen ways: total completion time is unchanged at best, and worse once the scheduler and the caches start paying for the churn. Memory is the harsher penalty for a process pool, where every worker is a full interpreter. The standard library moved with this in 3.13. `ProcessPoolExecutor` and `multiprocessing.Pool` default their worker count to `os.process_cpu_count()`; before 3.13 they used `os.cpu_count()`. `ThreadPoolExecutor`'s default became `min(32, (os.process_cpu_count() or 1) + 4)` — the `+ 4` is a small allowance for I/O-bound tasks that spend most of their time blocked, and the cap of 32 stops a very large machine from spawning an absurd number of threads for what is usually I/O work. On Windows a process pool is additionally capped at 61 workers by an operating-system wait limit. So on 3.14, code that simply omits the worker count already gets the better number. Code that hardcodes `ThreadPoolExecutor(max_workers=os.cpu_count() * 2)` — a very common line — does not. ### What the functions do *not* tell you Affinity is a *set* restriction: which CPUs you may run on. A quota is a *rate* restriction: how much CPU time you may consume per period. Only the first narrows `os.process_cpu_count()`. A process allowed a half CPU of time but permitted to run anywhere still sees the full machine count, sizes a pool for it, saturates every core in bursts and then gets throttled hard. This is the single most common surprise in constrained deployments, and it is why the number sometimes has to be told rather than detected. Python 3.13 added exactly that knob: `-X cpu_count=n` on the command line, or the `PYTHON_CPU_COUNT` environment variable. Either one overrides what **both** functions report for the whole process, so every library that sizes itself from them follows without being patched individually. Passing the literal value `default` restores normal detection. Note what the override does not do: it changes what CPython *reports*, not what the operating system permits — it is a sizing hint, not a sandbox. ### Physical versus logical Neither function distinguishes physical cores from hardware threads; both count logical CPUs. For workloads whose per-core work is memory-bound or already vector-saturated, two hardware threads on one physical core do not give two cores' throughput, and a pool sized to the logical count can measure slower than one sized to the physical count. Python does not expose a physical-core count in the standard library, so if that distinction matters for your workload, it belongs in configuration measured on the target machine rather than in a detection call. ### The practical rule Size CPU-bound pools from `(os.process_cpu_count() or 1)`, never from `os.cpu_count()`. Let the executor defaults do it when you have no better information, since since 3.13 they already use the right function. Allow an explicit override in configuration for every pool, because detection cannot see a quota. And treat `os.cpu_count()` as what it is: a fact about the hardware, useful for logging and diagnostics, and almost never the right input to a sizing decision.

  • What do the -X cpu_count option and the PYTHON_CPU_COUNT environment variable actually change?
    Both arrived in Python 3.13 and override the value CPython reports: os.cpu_count() and os.process_cpu_count() then both return the number you gave, and anything sizing itself from them, the concurrent.futures and multiprocessing defaults included, follows without being patched. Passing the literal value default restores normal detection. It changes what the interpreter reports, not what the operating system permits, so it is a fleet-wide sizing hint rather than any kind of enforcement.
  • Does os.process_cpu_count() shrink when a process is limited to a fraction of a CPU's time?
    No. It reports how many logical CPUs the calling thread is permitted to run on, so a set restriction such as an affinity mask narrows it, but a time-slice quota does not: the process still sees every CPU it is allowed onto and simply gets throttled once it exceeds its share. If a quota rather than an affinity mask is what bounds your workload, set the number explicitly with PYTHON_CPU_COUNT instead of trusting detection.
  • Why can these functions return None, and what should sizing code do about it?
    The count is not determinable on every platform, and both functions say so with None rather than guessing. Any expression like os.process_cpu_count() + 4 then raises TypeError at start-up. Write (os.process_cpu_count() or 1) so an unknown count degrades to a single worker, which is slow but correct, and log the fallback so the wrong-looking throughput has an explanation.

Asking os.cpu_count() how wide to run is like sizing a work crew from the number of desks in the building rather than the number of desks your team was actually given.

saying these in an interview costs you the question

  • Treats os.cpu_count() as the count of usable CPUs
  • Sizes a CPU-bound pool from the machine's logical CPU total
  • Thinks os.process_cpu_count() accounts for a fractional CPU quota
  • Never handles None from either count function
  • Assumes more workers than cores raises CPU-bound throughput
  • Believes multiprocessing.cpu_count() reports usable CPUs

context

open as a page

An email-digest sender's ThreadPoolExecutor backs up at a 1,200-send-per-minute peak; why does raising max_workers not fix it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Worker count is not the bottleneck. Threads added past the point where the downstream service saturates only wait somewhere else, while the executor's unbounded work queue hides the growing backlog. Size from arrival rate times service time, and bound admission instead.

open as a page

How do you set and enforce worker-pool sizing policy across a fleet of Python services with different CPU allotments?

level: principalimportance: should knowfreq 30%

basics

~20 s

Give every pool a sized default, an override in configuration, and a stated owner. Tell the process its width once at start-up rather than patching each library, and size I/O pools from load and downstream limits, never from cores.

open as a page

Why can a ThreadPoolExecutor of 8 workers end up with hundreds of live threads in one process?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Because pools nest and multiply. If each of the 8 tasks calls code that starts its own pool sized to the machine, the process runs 8 times that many threads, and a process pool repeats it inside every child.

open as a page