A feature-flag service sizes its process pool from os.cpu_count() inside a two-CPU container and misses its 92nd-percentile latency budget. How do you pick the right worker count?
answer
- The median looks fine; the tail does not
- Runnable workers outnumber the allowance
- Check how many periods were throttled
- Reconcile the affinity count with the quota
- Floor the derived value at one
basics
~20 sDerive the count from the smaller of os.process_cpu_count() and the control group's CPU quota, floored at 1. The host count starts dozens of workers on a two-CPU allowance, and the resulting throttling is what breaks the tail latency.
solid answer
~40 sFirst confirm the diagnosis: the pool was sized from the host's processor count, so the container runs far more workers than its two-CPU allowance can serve. All of them are runnable, the control group burns its quota early in each period, and every worker is then stopped until the period rolls over — a sawtooth that barely moves the median but destroys the 92nd percentile. Check the group's throttling counters to confirm. The fix is to derive the number: take `os.process_cpu_count() or 1`, read the quota from the group's `cpu.max` file, take the smaller, floor at 1, and pass it as `max_workers`. Better still, inject `PYTHON_CPU_COUNT` at deploy from the same field that sets the limit, so every standard-library default lands on the right number. Log the derived value at startup.
code
python · 21 linesimport os
from pathlib import Path
def quota_cpus():
cpu_max = Path("/sys/fs/cgroup/cpu.max")
if not cpu_max.exists():
return None
limit, period = cpu_max.read_text().split()
return None if limit == "max" else int(limit) / int(period)
def usable_cpus():
counts = [os.process_cpu_count() or 1]
quota = quota_cpus()
if quota is not None:
counts.append(max(1, int(quota)))
return max(1, min(counts))
print(usable_cpus())go deeper
Take away the rule of thumb: never size a pool from os.cpu_count() in a container. Ask what CPU limit the deployment sets and make the worker count follow it.
Explain the mechanism behind the sawtooth — the group spends its allowance early in each period and is then stopped — and write the derivation that takes the smaller of the affinity-aware count and the quota.
Drive the diagnosis from evidence: throttling counters, the median-versus-tail split, and the knock-on starvation of background refresh work. Then make the fix survive the next deployment rather than just the current one.
Frame it as a contract between the deployment system and the runtime: one field sets the limit and the worker count derives from it, with alerting when the two disagree, so no service can be sized by an assumption about the host.
### Confirm the shape of the failure before changing anything The symptom — median latency roughly unchanged, the 92nd percentile far over budget — is the signature of periodic stalling rather than slow work. The specific evidence to gather is the control group's CPU throttling counters: how many enforcement periods were throttled and how much total time was spent stopped. A service whose throttled-period count is a large fraction of elapsed periods is not compute-bound, it is quota-bound, and adding workers has made it worse rather than better. The cause is a number, not an algorithm. The pool was sized from `os.cpu_count()`, which reports the *host's* logical processors. A CPU limit is enforced as bandwidth on the control group — so much CPU time per period — and removes no processors from view, so on a large host that count can be an order of magnitude above the real allowance. ### Why oversubscription hurts the tail specifically With far more runnable workers than the quota can serve, the group consumes its whole allowance in the first slice of each enforcement period. Every thread in the group is then stopped until the next period begins, including ones that had nothing to do with the burst. Requests that arrive during the stalled remainder wait out the rest of the period before they are even looked at, which adds a floor to their latency that no amount of code optimization removes. Median requests, which often land early in a period, look fine — so a dashboard showing averages hides the problem entirely. There is a second-order effect worth naming for this kind of service. A flag-evaluation process typically refreshes its rule set on a background schedule and serves evaluations from an in-memory cache in between. That refresh is a thread inside the same throttled group: when the workers monopolize the allowance, the refresh misses its interval and evaluations keep answering from a stale cached value long past its intended lifetime. The visible bug is then "a flag flip took ten minutes to take effect", which sends people to look at the distribution layer rather than at the pool size. ### Deriving a number that respects both restrictions Three sources have to be reconciled: - **`os.process_cpu_count()`** — added in **Python 3.13**; the processors this thread may be scheduled onto, honouring the CPU affinity mask. Use `or 1`, since it may be `None`. - **The control-group quota** — on cgroup v2, the group's `cpu.max` file holds a limit and a period, or the word `max` for unlimited; dividing gives the allowance in whole CPUs. On v1 the same information lives in `cpu.cfs_quota_us` and `cpu.cfs_period_us`. - **The floor** — a fractional grant such as half a CPU must still yield one worker, so clamp the result to at least 1. Take the smallest of whichever apply. That single helper, called once at startup, is the whole fix; everything else is about making sure nothing bypasses it. ### Making the number stick Passing `max_workers` explicitly fixes the pool you edited. It does not fix the next pool someone adds, or the ones inside libraries that size themselves from the standard functions. Since **3.13** there is a better lever: setting the `PYTHON_CPU_COUNT` environment variable, or the `-X cpu_count` option, substitutes the value that *both* `os.cpu_count()` and `os.process_cpu_count()` return for the whole interpreter. Emitting that variable from the same deployment field that sets the container's CPU limit makes every standard-library default — `concurrent.futures.ProcessPoolExecutor`, `multiprocessing.Pool`, `concurrent.futures.ThreadPoolExecutor` — correct without any code change. Whichever lever you use, log the derived count at startup with the inputs that produced it. The most expensive version of this bug is the one where the pool size is right in one environment and wrong in another and nothing in the logs says which number the process chose. ### Verifying the fix Re-measure the throttled-period counters, not just the latency: the goal is that throttling becomes rare, at which point the tail is governed by real work again. Expect throughput to stay roughly the same or improve slightly — the quota was always the ceiling — while the tail collapses toward the median and resident memory falls sharply, since each worker process was a full interpreter. If throttling persists at a correct worker count, the service genuinely needs more CPU than it was granted, and that is a capacity conversation rather than a sizing bug.
- How would you prove throttling is the cause rather than slow request handling?Read the control group's CPU pressure and throttling counters: the number of enforcement periods throttled and the cumulative time the group spent stopped. If throttled periods are a large share of elapsed periods and the count rises in step with the latency spikes, the group is quota-bound. Slow handling would show as high CPU time per request with little throttling, and would move the median as well as the tail.
- Why does an oversubscribed pool leave the median latency roughly unchanged?Because the stall is periodic rather than uniform. Requests that arrive while the group still has allowance are served at full speed and land in the median; requests that arrive after the allowance is spent wait out the remainder of the enforcement period before any work happens. That splits the distribution, so averages and medians look healthy while the upper percentiles carry a floor equal to the throttled remainder.
- The service refreshes its rule set on a background schedule. How does the same bug affect that?The refresh thread lives in the same throttled control group, so it competes for the same exhausted allowance and misses its interval. Evaluations keep answering from the in-memory cache, which means the process serves a stale value long past its intended lifetime and a flag change appears not to take effect. It is worth alerting on cache age directly, so this failure names itself rather than presenting as a distribution problem.
- Would the same pool size be right if the container were given four CPUs instead of two?The derivation would be, but not the literal number — that is the point of computing it rather than hard-coding it. A helper that takes the smaller of the affinity-aware count and the quota, floored at 1, tracks the grant automatically. Hard-coding two would leave half the new allowance unused, which is the mirror image of the original bug and just as easy to ship.
saying these in an interview costs you the question
- Adds more workers to fix the latency spikes
- Blames garbage collection without checking throttling counters
- Assumes os.process_cpu_count() already accounts for the quota
- Hard-codes a worker count that drifts from the granted limit
- Rounds a fractional CPU grant down to zero workers
- Never logs the derived worker count at startup