A search indexer sizes its worker pool from the processor count it reads at start-up — why does it crawl under a small CPU ceiling?
answer
- what did the process actually ask
- the answer describes the machine
- the boundary rewrites no hardware inventory
- threads sized for absent processors
- read the ceiling or be told it
basics
~20 sIt asked the system how many processors the machine has, and the host answered. The pool and its per-worker buffers are sized for hardware it never gets, so workers contend for one small run-time allowance.
solid answer
~50 sA container gets its own views of things like the filesystem, the process table and the network, but a question such as "how many processors does this machine have?" is still answered from the shared kernel's view of the hardware. So the indexer reads the host's count, multiplies it into a worker pool, and gives each worker a buffer — all sized for a machine it will never get. The processor ceiling is an allowance of run time, not a number of processors, so the extra workers do not add throughput: they add context switching, queue depth and latency, and the workload spends its time being stopped for quota. The per-worker buffers overshoot the memory ceiling in the same way. The fix is to size from the ceiling — read it from the container's own accounting, or be told it explicitly through configuration.
code
pseudocode · 16 lines# WRONG: sized from what the machine reports
cores = system.reportedProcessorCount() # 64 - the host's
workers = cores * 2 # 128 workers
bufferPer = 32 MB
bufferTotal = workers * bufferPer # 4096 MB
# ceiling is 2 processors' worth of run time and 1024 MB
# RIGHT: sized from this container's own ceiling
cpuCeilingInCores = container.cpuCeilingInCores() # 2
memoryCeiling = container.memoryCeilingBytes() # 1024 MB
if cpuCeilingInCores is absent:
cpuCeilingInCores = system.reportedProcessorCount()
workers = max(1, round(cpuCeilingInCores * 2)) # 4 workers
bufferTotal = min(workers * bufferPer, memoryCeiling / 4)
# min(128 MB, 256 MB) = 128 MBgo deeper
Remember that hardware questions asked from inside a container are answered by the host. A pool built from that answer is built for a machine the workload does not have.
Explain the mechanics: a processor ceiling is run time per window, so extra workers add switching and queueing instead of throughput, and per-worker buffers multiply into the memory charge.
Demonstrate the diagnosis and the fix under load: pool size matching the node's hardware, stopped-for-quota climbing, and the pairing of a smaller pool with higher throughput as proof of cause.
The decision you own is whether concurrency is discovered or declared. Declaring it as configuration next to the ceiling makes the two reviewable together and removes a class of failure that appears only after someone else changes a number.
## What the process actually asked A process that wants to size a pool asks the system a hardware question: how many processors are there, how much memory is installed, how much is free. Those questions are answered from the machine's point of view. The per-resource views a container is given — its own filesystem tree, its own process table, its own network view, its own user identity mapping, its own hostname — do not include a rewritten hardware inventory. Nothing in the boundary makes the machine describe itself as smaller. So the indexer wakes up on a large node, is told it has many processors and a large amount of memory, and sizes itself for that. The ceiling that actually governs it lives in a different place: the kernel's resource-accounting and limiting mechanism, which the runtime configured for this container and which the process never consulted. ## The three overshoots that follow 1. **The worker pool.** A pool derived from the host's processor count can be an order of magnitude too large. Every worker is runnable, so the operating system keeps switching between them inside the small allowance the container gets, and throughput falls rather than rises. 2. **The per-worker memory.** Read buffers, decode buffers and per-worker caches are usually sized *per worker*, so an oversized pool multiplies straight into the memory charge — and a memory ceiling is enforced by ending the process, not by slowing it down. 3. **The background work.** Maintenance threads, batch flushers and parallel housekeeping are frequently sized from the same figure, so the overshoot repeats in parts of the process nobody thought about. ## What the process should have read instead | What the process reads | What it gets | What it should use | |---|---|---| | Processor count | the host's total | its own allowance of run time per window, expressed as processors' worth | | Installed memory | the machine's memory | the memory ceiling declared for this container | | Free memory | node-wide free memory | its ceiling minus its own current charge | The first column is a question about a machine. The third column is a question about a boundary. They are different questions and the process asked the wrong one. ## A processor ceiling is run time, not processors The most useful correction to make here is conceptual. A processor ceiling is not "you get two processors"; it is "you get this much run time per repeating accounting window". That is why: - A ceiling can be a **fraction** of one processor, which no integer core count can express. - More threads never increase the allowance. They only divide the same allowance into smaller, more expensive pieces. - A workload can exhaust its allowance early in a window and sit stopped for the rest of it, which reads as latency rather than as high utilisation. ## Making the process container-aware - **Read the ceiling from the container's own accounting** where the platform exposes it to the workload, and derive the pool from that. - **Or be told.** Passing the intended concurrency in as configuration is the simplest, most portable answer, and it keeps the number under the same review as the ceiling itself — the two should change together. - **Handle the no-ceiling case.** A container need not have a ceiling at all; platforms differ in whether one is required, defaulted or simply absent. Where none exists, the host's figure is the honest answer — until someone adds a ceiling later and the same image starts thrashing without a single line of code changing. - **Round with care.** A fractional ceiling rounded up to a whole number of processors reintroduces the same overshoot in miniature; clamp to at least one worker and treat the remainder as run time, not as another worker. - **Prove it in the numbers.** After the fix, the stopped-for-quota signal should fall and throughput should rise even though the pool got smaller — that pairing is the evidence that the overshoot was the cause. ## Why this failure is so hard to see Nothing errors. The image is identical everywhere, the process starts cleanly, the logs are quiet, and on a developer machine — where the process really does own the hardware it read about — it is fast. The symptom appears only under a ceiling, as latency with an unremarkable average utilisation, which is why it is so often misfiled as a slow dependency or a bad node. The give-away is the shape: the pool size matches the *node's* hardware, and nobody chose it.
- What if the container runs with no processor ceiling at all?Then the host's count is the honest answer and sizing from it is correct, because the workload really can use the machine. The risk is the day someone adds a ceiling: the same image, unchanged, starts building a pool for hardware it no longer has. That is why the ceiling and the concurrency setting should be reviewed together.
- Why does a fractional ceiling break pool sizing worse than a whole one?Because a fraction of a processor cannot be expressed as a count. Rounding up gives at least one worker that expects a whole processor's worth of run time and will not get it. The safe reading is to treat the ceiling as an amount of run time per window and let the pool size follow from that, clamped to a minimum of one.
- How would you confirm that oversized concurrency, rather than a slow dependency, is the cause?Compare the pool size against the container's ceiling: if the pool matches the node's hardware and nobody chose that number, it is the prime suspect. Then reduce concurrency and watch the stopped-for-quota signal fall while throughput rises. A slow dependency shows the opposite shape — waiting workers and low run-time demand.
saying these in an interview costs you the question
- Believes a process inside a container automatically sees its own ceiling
- Thinks extra workers are harmless because the scheduler sorts it out
- Says the platform raises the ceiling to fit the threads the process started
- Sizes per-worker buffers without checking the memory ceiling
- Assumes every container has a processor ceiling to read
- Treats a processor ceiling as a count of processors rather than run time