A container reports 92% memory use while its host shows most memory free — which reading decides its fate?
answer
- a percentage needs a denominator
- which boundary was it measured in
- host headroom belongs to the node
- enforced against the same accounting
- compare against the declared ceiling
basics
~20 sThe container's own figure decides: usage is accounted inside its boundary and judged against the ceiling declared for it. The host's free memory is capacity this workload may never use, so 92% means it is near being killed.
solid answer
~40 sThose two numbers have different denominators. The container's `92%` is what the kernel's resource-accounting charged to that boundary, divided by the memory ceiling declared for that container — and the ceiling is what the runtime enforces on the running process. The host's `free` figure is machine-wide and answers a capacity question about the node, not a survival question about this workload. A container does not get to spend the host's spare memory once its own ceiling is reached; exceeding a memory ceiling ends the process regardless of how idle the machine is. So read usage against the ceiling first. The host figure is still worth a glance, but it tells you about the node's situation, which is a different diagnosis.
go deeper
Recall that a container's usage is charged to its own boundary and compared with the ceiling declared for it. The machine's free memory is a different number with a different denominator.
Explain that the same accounting mechanism produces both the usage figure and the ceiling enforcement, which is why that pairing is the meaningful fraction, and why byte counts across replicas are not comparable when ceilings differ.
Show the operational habit: every memory alert names which figure it is built on, thresholds are written against the ceiling, and a node-level reading triggers a node-level investigation rather than a workload verdict.
The trade-off you own is what the platform's dashboards standardise on. A single agreed denominator across teams makes readings comparable; mixing container and host denominators in one view makes every number arguable.
## Two readings, two denominators Every utilisation percentage is a fraction, and it means nothing until you know what sits underneath it. A container platform publishes two families of number that look interchangeable on a dashboard and are not: - **What this workload used**, accounted inside its own boundary, divided by **the ceiling declared for that container**. - **What the machine has**, totalled across everything on it, divided by **the installed hardware**. The question *"is this container about to die?"* is answered only by the first. The question *"can this node take another workload?"* is answered only by the second. Reading one as if it were the other is the single most common mistake in container diagnosis, and it goes both ways: people relax because the host looks idle, and people panic because the host looks full while every individual workload is comfortable. ## Where the container's figure comes from When a platform tells you a container is using a certain amount of memory, it is reading the kernel's resource-accounting and limiting mechanism. That mechanism charges each page of memory to the boundary that caused it to be allocated, and it is the same mechanism the ceiling is enforced against. So the numerator and the denominator of the container figure come from the same place — which is exactly why it is the trustworthy reading. It counts what this workload caused, and it compares it against the number that will be used to decide whether this workload keeps running. The host figure is assembled differently: it totals everything on the machine, including the operating system, the platform's own agents, and every other container. A workload can be 5% of a busy host and still be at 100% of its own ceiling. ## What each reading predicts | Reading | Denominator | What it predicts | |---|---|---| | Container usage against its ceiling | the ceiling declared for that container | whether *this* workload is about to be ended for exceeding its memory ceiling | | Container usage against host memory | the machine's installed memory | almost nothing useful about this workload | | Host memory free | the machine | whether the node as a whole is heading for trouble — a different diagnosis, and not this one | ## Why the host's headroom will not save it A ceiling is a hard boundary on the running process, not a hint. The runtime enforces it whether or not the machine has spare capacity, and that is the whole point: the ceiling exists so one workload's growth cannot become the machine's problem. Two consequences follow: 1. A container at 92% of a small ceiling on an almost-empty host is in more danger than a container at 40% of a large ceiling on a busy host. 2. "There is plenty free" is never a reason to dismiss a container's high reading. The free memory belongs to the node, not to this workload. ## Reading the pair together - Judge a workload against its **own** ceiling. That is the number the platform will act on. - Use the host figure to answer node-level questions, and treat it as a separate investigation with its own owner. - Compare like with like across replicas: two copies of the same workload with different ceilings will show different percentages for identical work, and the percentage — not the byte count — is what tells you which one is close to the edge. - Watch the trend, not the instant. A single sample at 92% may be a transient; the same figure climbing across an hour is a trajectory. - Be explicit about which figure an alert is built on. An alert whose threshold was written against a host-wide reading will either never fire for a small container or fire constantly for a large one. ## When there is no ceiling A container does not have to have one. If no memory ceiling was declared, there is no denominator, and the platform can only show you absolute bytes — the percentage either disappears or is quietly computed against the machine, which is the misleading form of the same mistake. In that state the workload's effective limit is whatever the machine has left, shared with everything else on it, and judging "how close is it?" needs the node's picture rather than the container's. That is a different subject; the point for this reading is simply that a percentage with no declared ceiling behind it is not the survival number you think it is. ## The habit to build When someone shows you a memory graph for a container, ask two questions before you interpret it: *what is the ceiling*, and *is this figure charged to the container or to the machine*. Almost every wrong conclusion drawn from a container memory graph comes from skipping one of them.
- A host-wide process view and the container's own accounting disagree about the same workload. Which do you trust?The container's accounting. It charges what this boundary caused and it is the same figure the ceiling is enforced against, so it is the one that predicts whether the workload keeps running. A host-wide view attributes shared pages by its own rules and knows nothing about the ceiling.
- The container has no memory ceiling declared. What does its utilisation percentage mean then?There is no denominator, so any percentage shown is against the machine and is not a survival number. You are left with absolute bytes, and the effective limit becomes whatever the node has spare and shared with everything else on it — which makes it a node-level question rather than a container-level one.
saying these in an interview costs you the question
- Says the host still has free memory, so the container must be fine
- Divides a container's usage by the host's total memory to get utilisation
- Thinks a container may spend the host's spare memory past its ceiling
- Treats 92% of a ceiling and 92% of a machine as the same risk
- Assumes usage is measured for the node and split evenly between containers