skip to content

Operations & Diagnostics

What a container platform hands you when something is wrong - a captured output stream, usage against a limit, a boundary with no shell inside - and the two failures you read all of it for.

on this pageshow

questions

20

A container reports 92% memory use while its host shows most memory free — which reading decides its fate?

level: juniorimportance: must knowfreq 65%

answer

  1. a percentage needs a denominator
  2. which boundary was it measured in
  3. host headroom belongs to the node
  4. enforced against the same accounting
  5. compare against the declared ceiling

basics

~20 s

The container's own figure decides: usage is accounted inside its boundary and judged against the ceiling declared for it. The host's free memory is capacity this workload may never use, so 92% means it is near being killed.

solid answer

~40 s

Those two numbers have different denominators. The container's `92%` is what the kernel's resource-accounting charged to that boundary, divided by the memory ceiling declared for that container — and the ceiling is what the runtime enforces on the running process. The host's `free` figure is machine-wide and answers a capacity question about the node, not a survival question about this workload. A container does not get to spend the host's spare memory once its own ceiling is reached; exceeding a memory ceiling ends the process regardless of how idle the machine is. So read usage against the ceiling first. The host figure is still worth a glance, but it tells you about the node's situation, which is a different diagnosis.

go deeper

for a junior

Recall that a container's usage is charged to its own boundary and compared with the ceiling declared for it. The machine's free memory is a different number with a different denominator.

for a middle

Explain that the same accounting mechanism produces both the usage figure and the ceiling enforcement, which is why that pairing is the meaningful fraction, and why byte counts across replicas are not comparable when ceilings differ.

for a senior

Show the operational habit: every memory alert names which figure it is built on, thresholds are written against the ceiling, and a node-level reading triggers a node-level investigation rather than a workload verdict.

for a principal

The trade-off you own is what the platform's dashboards standardise on. A single agreed denominator across teams makes readings comparable; mixing container and host denominators in one view makes every number arguable.

## Two readings, two denominators Every utilisation percentage is a fraction, and it means nothing until you know what sits underneath it. A container platform publishes two families of number that look interchangeable on a dashboard and are not: - **What this workload used**, accounted inside its own boundary, divided by **the ceiling declared for that container**. - **What the machine has**, totalled across everything on it, divided by **the installed hardware**. The question *"is this container about to die?"* is answered only by the first. The question *"can this node take another workload?"* is answered only by the second. Reading one as if it were the other is the single most common mistake in container diagnosis, and it goes both ways: people relax because the host looks idle, and people panic because the host looks full while every individual workload is comfortable. ## Where the container's figure comes from When a platform tells you a container is using a certain amount of memory, it is reading the kernel's resource-accounting and limiting mechanism. That mechanism charges each page of memory to the boundary that caused it to be allocated, and it is the same mechanism the ceiling is enforced against. So the numerator and the denominator of the container figure come from the same place — which is exactly why it is the trustworthy reading. It counts what this workload caused, and it compares it against the number that will be used to decide whether this workload keeps running. The host figure is assembled differently: it totals everything on the machine, including the operating system, the platform's own agents, and every other container. A workload can be 5% of a busy host and still be at 100% of its own ceiling. ## What each reading predicts | Reading | Denominator | What it predicts | |---|---|---| | Container usage against its ceiling | the ceiling declared for that container | whether *this* workload is about to be ended for exceeding its memory ceiling | | Container usage against host memory | the machine's installed memory | almost nothing useful about this workload | | Host memory free | the machine | whether the node as a whole is heading for trouble — a different diagnosis, and not this one | ## Why the host's headroom will not save it A ceiling is a hard boundary on the running process, not a hint. The runtime enforces it whether or not the machine has spare capacity, and that is the whole point: the ceiling exists so one workload's growth cannot become the machine's problem. Two consequences follow: 1. A container at 92% of a small ceiling on an almost-empty host is in more danger than a container at 40% of a large ceiling on a busy host. 2. "There is plenty free" is never a reason to dismiss a container's high reading. The free memory belongs to the node, not to this workload. ## Reading the pair together - Judge a workload against its **own** ceiling. That is the number the platform will act on. - Use the host figure to answer node-level questions, and treat it as a separate investigation with its own owner. - Compare like with like across replicas: two copies of the same workload with different ceilings will show different percentages for identical work, and the percentage — not the byte count — is what tells you which one is close to the edge. - Watch the trend, not the instant. A single sample at 92% may be a transient; the same figure climbing across an hour is a trajectory. - Be explicit about which figure an alert is built on. An alert whose threshold was written against a host-wide reading will either never fire for a small container or fire constantly for a large one. ## When there is no ceiling A container does not have to have one. If no memory ceiling was declared, there is no denominator, and the platform can only show you absolute bytes — the percentage either disappears or is quietly computed against the machine, which is the misleading form of the same mistake. In that state the workload's effective limit is whatever the machine has left, shared with everything else on it, and judging "how close is it?" needs the node's picture rather than the container's. That is a different subject; the point for this reading is simply that a percentage with no declared ceiling behind it is not the survival number you think it is. ## The habit to build When someone shows you a memory graph for a container, ask two questions before you interpret it: *what is the ceiling*, and *is this figure charged to the container or to the machine*. Almost every wrong conclusion drawn from a container memory graph comes from skipping one of them.

  • A host-wide process view and the container's own accounting disagree about the same workload. Which do you trust?
    The container's accounting. It charges what this boundary caused and it is the same figure the ceiling is enforced against, so it is the one that predicts whether the workload keeps running. A host-wide view attributes shared pages by its own rules and knows nothing about the ceiling.
  • The container has no memory ceiling declared. What does its utilisation percentage mean then?
    There is no denominator, so any percentage shown is against the machine and is not a survival number. You are left with absolute bytes, and the effective limit becomes whatever the node has spare and shared with everything else on it — which makes it a node-level question rather than a container-level one.

saying these in an interview costs you the question

  • Says the host still has free memory, so the container must be fine
  • Divides a container's usage by the host's total memory to get utilisation
  • Thinks a container may spend the host's spare memory past its ceiling
  • Treats 92% of a ceiling and 92% of a machine as the same risk
  • Assumes usage is measured for the node and split evenly between containers
open as a page

A service writes its log to a file inside the container and the platform's captured log view is empty — why?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A container platform's log contract is the process's standard output and error streams — the runtime captures those. A file written inside the container is outside that capture, so nothing collects it, and it is gone when the instance is replaced.

open as a page

A workload never reached running and is serving nothing — which four distinct startup failures produce that symptom?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Four separate failures look identical from outside: the image never arrived on the node, the start was refused before any code ran, the process started and died, or it runs and never reports ready. Naming which one is the first move.

open as a page

A search indexer sizes its worker pool from the processor count it reads at start-up — why does it crawl under a small CPU ceiling?

level: middleimportance: must knowfreq 60%

basics

~20 s

It asked the system how many processors the machine has, and the host answered. The pool and its per-worker buffers are sized for hardware it never gets, so workers contend for one small run-time allowance.

open as a page

A worker ships on an image with no shell and no package manager and fails in one environment only — what debugging moves are gone?

level: middleimportance: must knowfreq 62%

basics

~20 s

Everything that needs a program inside the container: no interactive session, no install, no file viewing, no probing from within. What is left is a joined debug container, the instance's retained output, copying a file out, and a local rebuild with tools added.

open as a page

Why can a container be killed for running out of memory while its own usage stays under its declared memory ceiling?

level: middleimportance: must knowfreq 62%

basics

~20 s

Two different shortages end in a kill. A ceiling is enforced against one container's own usage; the host's total memory is a separate constraint. When the host runs out, the kernel picks a victim by footprint, and your ceiling grants no immunity.

open as a page

You attach a debug container sharing a shell-less workload's process and network view — why can it still not read the workload's config file?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Because the filesystem root is not one of the views being shared. The debug container sees its own image's files; the target's files are reachable only through the operating system's per-process entry for the target process, which the shared process view makes identifiable.

open as a page

When a node runs short of memory, what decides which workload the node agent evicts first?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Rank, not blame, decides. The node agent evicts by how far each workload's usage sits above what it reserved — a workload that declared no reservation is first in line, and one at or under its reservation is last.

open as a page

A shell-less batch worker already exited, so there is nothing to join — what evidence about that run survives?

level: middleimportance: should knowfreq 44%

basics

~20 s

Three things: the output the platform captured from that instance, the exit status it recorded, and anything the process wrote to a path that outlives the container. Everything in the container's own writable layer, and all of its memory, went with it.

open as a page

A service prints summaries to its output stream and full detail to a file inside the container — what does that split cost you?

level: middleimportance: should knowfreq 46%

basics

~20 s

The split keeps only the half that was never the useful half. Summaries on the output stream are captured and shipped; the detail is collected by nothing, readable only from inside that running instance, and gone the moment it is replaced.

open as a page

A nightly reconciliation batch keeps retrying and its output stream is completely empty — has the process ever run?

level: middleimportance: should knowfreq 52%

basics

~20 s

Probably not, but emptiness is evidence rather than proof. Nothing written by the process favours a failure before it ran — a fetch that failed or a refused start — while a process that dies before its first line can also leave nothing behind.

open as a page

A workload instance has been alive for ten minutes, was never replaced, and still receives no traffic — which startup failure is it?

level: middleimportance: should knowfreq 58%

basics

~20 s

The fourth one: the process runs but has never reported itself ready, so the platform keeps it out of routing. A climbing uptime with no replacement and no termination record rules out a process that started and died.

open as a page

A service's processor utilisation averages 40% of its ceiling, yet its stopped-for-quota counter climbs and latency spikes — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A processor ceiling is enforced as run time allowed per short repeating window, while utilisation is averaged over a far longer one. The burst spends its allowance early and is stopped until the next window.

open as a page

During a burst, an hour of a container's captured output is missing centrally although the collector never went down — what happened?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The node's copy of the captured output is bounded by a size cap and a number of files kept. Under a burst the workload outran the collector, and the oldest files were deleted before it read them — those lines never entered the pipeline at all.

open as a page

A node runs out of disk and starts evicting workloads that barely write anything — where did the space go?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Most of a node's disk is consumed by things no running workload owns: image layers cached from everything ever placed there, and artifacts left by containers already removed. The node reclaims those first, and only then evicts.

open as a page

The same nightly batch reaches running on one node and never starts on another — what does that asymmetry tell you?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Asymmetry points at node-local state rather than the spec: an artifact already held on one node, fetch credentials or a rate limit that differ per node, a mount attachable only in one failure domain, or nodes of differing processor architecture.

open as a page

In a merged view of a container's standard output and error streams, why can two lines appear in the wrong order?

level: middleimportance: nice to knowfreq 28%

basics

~10 s

The two output streams are separate channels with independent buffers, so a line written first can reach the runtime second. Order is preserved within each stream; across the two, the merged view guarantees nothing.

open as a page

A container's reported memory climbs for hours while its live data stays flat — is that a leak?

level: seniorimportance: nice to knowfreq 32%

basics

~10 s

Usually not. The figure charges the container for page cache its own file reads and writes created, and those pages are reclaimable. A rising total with a flat non-reclaimable part is caching, not leaking.

open as a page

Your platform team wants every service on a shell-less image — what must on-call engineers be given before that standard is safe?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

A replacement for every use the removed shell had: a permitted joined-debug-container path with a maintained tool image, diagnostics on the output stream, the terminated instance's output reaching the pager, a file-extraction route, and a one-step local rebuild with tools.

open as a page

How much of every node would you hold back so the platform's ordered eviction, rather than the kernel's kill, chooses the victim?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Hold back enough that reclaim finishes before free memory reaches zero. That headroom buys a victim chosen by declared rank with a recorded reason, instead of one chosen abruptly by footprint. The cost is capacity you never sell.

open as a page