skip to content

A container's reported memory climbs for hours while its live data stays flat — is that a leak?

level: seniorimportance: nice to knowfreq 32%

answer

  1. what is the figure charging for
  2. not all of it is live
  3. file reads leave a footprint
  4. reclaimable pages are dropped first
  5. growth with a flat live set

basics

~10 s

Usually not. The figure charges the container for page cache its own file reads and writes created, and those pages are reclaimable. A rising total with a flat non-reclaimable part is caching, not leaking.

solid answer

~50 s

The headline memory figure for a container is everything the accounting charged to that boundary, and that includes cached file pages the workload caused by reading or writing files. Cache is **reclaimable**: when the container approaches its ceiling, those pages are dropped rather than the workload being ended, so a total that climbs toward the ceiling and plateaus there is normal behaviour for anything that touches files. What you actually need is the non-reclaimable part — the memory the workload is genuinely holding. A leak shows there: the non-reclaimable charge rises and never returns after load falls. Platforms differ in whether the number they show by default is the raw charge or one with reclaimable pages already subtracted, so the first question to ask about any container memory graph is which of the two it is.

go deeper

for a junior

Remember that a container's memory figure includes cached file pages the workload caused, and that those pages can be given back. A rising number is not automatically a problem.

for a middle

Explain the split between reclaimable and non-reclaimable memory, and why the ceiling being enforced against the total still does not mean cache growth ends the process.

for a senior

Show the diagnosis: read the non-reclaimable part, observe a full load cycle, compare the plateau against the ceiling, and refuse the tempting fixes of dropping cache or raising the ceiling.

for a principal

The standard you set is which memory figure the estate's dashboards and alerts are built on. Alerting on the total trains teams to ignore memory alerts, which costs you the one signal that matters.

## What the figure is charging you for When a workload reads or writes a file, the kernel keeps those file pages in memory so the next access does not have to touch storage. That page cache is charged to the boundary that caused it, which means a container's memory figure is not "the memory my program allocated". It is the sum of several things the workload caused, and they do not all behave the same way under pressure. This is why a perfectly healthy file-touching workload — an indexer, a log writer, anything that reads a large corpus — shows memory that climbs steadily for hours and then sits just below its ceiling forever. Nothing is wrong. The workload read files, the pages were cached, and the cache grew to fill the space it was allowed. ## Reclaimable against not | Part of the charge | What it is | Under pressure | |---|---|---| | Cached file pages | copies of file contents the workload read or wrote | dropped and re-read later; costs latency, not the process | | Dirty pages awaiting write-out | written data not yet on storage | must be written before the memory can be released | | The workload's own in-use memory | heap, stacks, buffers it is genuinely holding | cannot be reclaimed at all | The ceiling is enforced against the **total**, which is why hitting it does not immediately end the process: the kernel first reclaims what it can — the top row, and the second once written out. Only when the non-reclaimable part alone cannot fit under the ceiling does the out-of-memory kill follow. That ordering is the whole reason a climbing headline figure is not an emergency by itself. ## So how do you tell a leak from a cache? 1. **Split the figure.** Read the non-reclaimable part rather than the total. A leak lives there; cache does not. 2. **Watch it across a load cycle.** Cache and working memory both rise under load; only a leak fails to fall back after load drops. One quiet period is worth more than an hour of peak-time graphs. 3. **Compare the plateau.** Cache plateaus at whatever headroom the ceiling leaves, so raising the ceiling moves the plateau up by roughly the same amount. A leak ignores the ceiling and keeps climbing until it reaches it. 4. **Check the restart shape.** After a restart, cache rebuilds quickly to the same plateau and stops. A leak climbs at the same slow rate it did before, from a low starting point. ## What to alert on - Alert on the **non-reclaimable** charge against the ceiling, not on the total. An alert on the total fires on every workload that reads files and teaches the team to ignore memory alerts. - Alert on the **outcome** as well: a workload ended for exceeding its memory ceiling is a fact, not an inference, and it is the signal that the earlier inference was wrong. - Track **the trend after load falls**, because that is where the two shapes separate. - Be explicit in dashboards about **which figure is drawn**, since platforms differ here: some show the raw charge for the boundary, some show a figure with reclaimable pages already subtracted, and the same workload looks like two different stories in the two views. ## The tempting wrong fixes - **Dropping the cache to make the number look better.** The pages were free to give up anyway; forcing them out only makes the workload re-read from storage, trading a cosmetic improvement for real latency. - **Raising the ceiling because the graph is near the top.** If the space is filled by cache, the new headroom fills with cache too and the graph looks the same a day later. - **Calling it a leak and hunting in the code.** Hours disappear into a profile that shows a flat live set, because the growth was never in the workload's own memory. ## The judgment to carry A container memory graph answers one question honestly — how much is charged to this boundary — and people use it to answer a different one: how much this workload is holding. The gap between those two questions is the page cache. Before anyone concludes anything from a rising line, establish which figure is plotted and what its non-reclaimable part is doing. That single habit removes most false leak investigations, and it also prevents the opposite error: dismissing a genuine climb as "just cache" when the non-reclaimable part has been rising all along.

  • Does the same 'includes cache' caveat apply to the figure the ceiling is enforced against?
    Yes. The ceiling is checked against the total charged to the boundary, cache included. That is why reaching the ceiling triggers reclaim rather than an immediate kill: the kernel drops what it can, and only a non-reclaimable demand that cannot fit under the ceiling on its own produces the out-of-memory kill.
  • Why is dropping a container's cache to 'fix' the reading a bad idea?
    Because those pages were already available to the kernel at no cost to the workload. Forcing them out only guarantees the next access re-reads from storage, so you exchange a cosmetically lower graph for real latency, and the number climbs back as soon as the workload reads files again.

saying these in an interview costs you the question

  • Calls every rising container memory figure a leak
  • Thinks page cache charged to a container is never reclaimed
  • Believes cache growth alone triggers the out-of-memory kill
  • Assumes the headline figure is only the workload's live data
  • Raises the ceiling without reading the non-reclaimable part first
  • Dismisses a real climb as cache without checking it after load falls