Why does a containerized service reading the machine's CPU and memory totals see the host's numbers rather than its own ceiling?
answer
- hiding and capping are different jobs
- capacity numbers come from the host
- the ceiling is never reported back
- a time quota, not a processor count
- read your ceiling or be told it
basics
~20 sThe fences hide other workloads; they do not rewrite the interfaces that report machine capacity. Those numbers come from the shared kernel's host-wide view, while the container's ceiling lives in an accounting group nobody told the process about.
solid answer
~50 sBecause the container's fences and its cap are different mechanisms with different jobs. The fences replace *views* — filesystem, process table, network, host identity — and the interfaces that report how much CPU and memory the machine has are not among them, so they keep answering for the whole host. The cap lives in the accounting group the runtime put the process in, and that group is not consulted when the process asks how big the machine is. So a service on a 64-way host with a two-CPU-equivalent ceiling reads `64`, sizes its caches and worker pools for a machine it does not have, and then meets the ceiling the hard way: throttled on CPU, and ended by the kernel when it crosses the memory wall. The fix is to read the ceiling itself, or to have the deployment pass it in.
code
pseudocode · 15 lines# host has 64 processors and 256 GB; this container's ceiling is 2 CPU-equivalents / 4 GB
reported_cpus = machine_processor_count() # -> 64 (host-wide, not fenced)
reported_memory = machine_memory_total() # -> 256 GB (host-wide, not fenced)
own_cpu_ceiling = accounting_group_cpu_quota() # -> 2.0 or unset
own_memory_ceiling = accounting_group_memory_max() # -> 4 GB or unset
if own_memory_ceiling is set:
budget = own_memory_ceiling * 0.7 # <- branch taken in a capped container
else:
budget = reported_memory * 0.7 # only correct when nothing caps us
# the unchanged service skips the check entirely and uses reported_memory:
# budget = 179 GB against a 4 GB wall -> the kernel ends the processgo deeper
Remember the headline: a process inside a container sees the whole machine's capacity, not its own slice. If you are choosing a cache size or a worker count, get the number from configuration rather than from the machine.
Explain the mechanism, not just the symptom: capacity reporting is a host-wide interface that no fence replaces, and the ceiling lives in an accounting group that is never consulted when the process asks how big the machine is.
Show the production consequence and the asymmetry between the two ceilings — CPU overshoot buys latency, memory overshoot buys a kill — and be able to say how you would detect a workload sized against host totals before it reaches a busy environment.
Make it a platform standard rather than a per-service fix: decide whether ceilings are injected as configuration for every workload, and accept that relying on a platform to present corrected capacity values buys convenience at the cost of portability.
## The number the process reads, and where it comes from When a long-running service starts, it commonly asks the machine two questions: how many processors are there, and how much memory is there. Inside a container, the honest answer to both is "that is not a well-formed question", and the interfaces that answer them do not know that. They report the **host's** totals, because they are host-wide interfaces served by the one shared kernel, and nothing in the container's fence set rewrites them. So a search indexer moved into a container unchanged still behaves as though it owns the box. It reads the host's processor count and memory size, sizes its internal structures accordingly, and only finds out later that it was budgeting against someone else's capacity. ## Why the fences do not fix it This is the clearest practical consequence of the fact that hiding and capping are separate mechanisms: - The **fences** change what a process can *see and reach* — its root filesystem, its process table, its network, its host identity. A capacity report is none of those things. It is a description of the machine, and the machine is genuinely that big. - The **accounting and limiting group** changes what a process may *consume*. It is enforcement, not narration. It sets a ceiling; it does not go back and edit what the process reads. Neither mechanism was designed to produce the answer the application wanted, and the gap between them is exactly the bug. | What the process reads | What actually constrains it | |---|---| | the host's total processor count | a CPU quota: run time allowed per period | | the host's total memory size | a hard memory ceiling enforced by the kernel | | the host's load and free memory | its own accounted usage inside its group | | nothing about its own ceiling | the ceiling, which it never asked for | ## A CPU ceiling is a rate, not a count Even if the process could see its ceiling, "how many processors do I have" would still have no clean answer. A CPU ceiling is usually expressed as an amount of run time per period — the equivalent of two and a half processors, say — and a rate does not convert into an integer count. It can be spent as two threads running flat out, or twenty threads each running a tenth of the time. That is why the mistake is not only "the number is too big" but also "the number is the wrong kind of thing". Rounding a fractional quota up gives a count the workload cannot sustain; rounding down wastes what was reserved for it. ## How the two ceilings punish the mistake differently 1. **CPU.** Demand above the quota does not fail; it **waits**. The work still completes, more slowly, and the damage shows up as latency — particularly at the tail, where a request unlucky enough to be running when the quota runs out waits for the next period. 2. **Memory.** There is no equivalent of waiting. When the container crosses its memory ceiling, the kernel **ends a process** inside it. A cache sized for the host's memory does not degrade; it gets the workload killed, usually under exactly the load that made the cache fill. That asymmetry is why a workload sized against the host's totals tends to look fine in a quiet environment and die in a busy one. ## What to do instead - **Read the ceiling, not the machine.** The accounting group's own configured ceiling is readable, and a workload that consults it gets a number that is true for itself. Many modern process runtimes now do this automatically instead of trusting the machine totals; older software and anything written before containers were normal does not. - **Pass it in.** The most portable fix is for whoever deploys the workload to supply the ceiling as configuration, so the application never has to guess. It costs one value in the deployment and removes an entire class of surprise. - **Insert a layer that tells the truth.** Some platforms can present capacity values that reflect the container's ceiling rather than the host's. Where that exists it is convenient, but treating it as guaranteed is how a workload becomes portable only to the platform it was written on. ## The case where there is no ceiling at all If a container is given no CPU or memory ceiling, the host's totals are momentarily *accurate* — and still the wrong planning input. An uncapped container shares the machine with everything else scheduled onto it, so sizing for the full host is sizing for a machine it will never have to itself. The absence of a ceiling makes the workload a noisy neighbour rather than a correct one. ## What interviewers listen for The strong answer connects the symptom to the mechanism in one move: the capacity numbers are not fenced, the ceiling is not reported, and those are two sides of one design where views hide and groups cap. Candidates who blame the image, the base image's contents, or a missing permission have not built that model yet.
- Why is "how many processors do I have" not a well-formed question inside a capped container?Because the cap is a rate, not a count: it grants an amount of run time per period, which can be spent by a few threads running constantly or many threads running occasionally. There is no integer that describes it. Rounding up promises throughput the quota cannot sustain, and rounding down wastes capacity that was already reserved.
- A container is given no CPU or memory ceiling at all. What does it read, and is that number now right?It reads the host's real totals, which are accurate but still the wrong planning input. An uncapped container shares the machine with every other workload placed there, so sizing for the whole host means sizing for capacity it will never hold alone. The absence of a ceiling turns the workload into the neighbour everyone else has to survive.
- Where should a workload get its capacity numbers instead?From its own ceiling, read out of the accounting group it was placed in, or — more portably — from a value the deployment passes in as configuration. Both give a number that is true for this instance. The machine totals are only correct for a process that genuinely owns the machine, which a container never does.
saying these in an interview costs you the question
- Says the fences make a process see only its own share of CPU and memory.
- Thinks a CPU ceiling grants a fixed number of dedicated processors.
- Believes exceeding a CPU ceiling ends the process the way a memory ceiling does.
- Assumes the platform rewrites capacity numbers whenever a ceiling is set.
- Blames the image or a missing permission for the wrong hardware report.