In Kubernetes, kubectl top pod shows 1870Mi against a 2Gi memory limit with no OOMKills; how do working set and RSS differ?
answer
- cache hides inside the top number
- usage minus inactive_file
- anonymous memory cannot be dropped
- kernel reclaims before it kills
- eviction uses working set
basics
~20 sWorking set is the cgroup's memory usage minus inactive file cache, so it still counts active page cache the kernel can reclaim. RSS counts anonymous memory that cannot be reclaimed without swap. RSS near the limit predicts an OOM kill.
solid answer
~40 s`kubectl top pod` shows the **working set**, which cAdvisor computes as the cgroup's total memory usage minus `inactive_file`. That still includes **active page cache**: file pages the container has read or memory-mapped recently. When usage nears the limit, the kernel reclaims clean cache first, so a service that maps a large index can sit at 91% working set for weeks. **RSS** (`container_memory_rss`) is anonymous memory, heap and stacks, which the kernel cannot drop on a node without swap; when that approaches the limit, an OOM kill is close. So I compare RSS, not working set, with the limit to predict kills. Working set still matters, though: the kubelet's node-pressure eviction uses it, and reclaiming active index pages causes refaults that show up as latency.
go deeper
Remember that kubectl top pod shows working set, and that it can include file cache the kernel is able to reclaim.
Explain the working set formula, what RSS counts, and the order the kernel follows: reclaim first, kill only when that fails.
Read RSS for OOM risk and working set for reclaim pressure, and recognise cache refaults as a latency cause that looks like throttling.
Set memory limits so hot data stays cached, and decide which memory metric alerts should use so teams are not paged for healthy cache.
## The number `kubectl top pod` shows `kubectl top pod` reports a container's **working set**, taken from the kubelet's `container_memory_working_set_bytes`. The kubelet's embedded cAdvisor computes it as **working set = cgroup memory usage − `inactive_file`**, where **memory usage** is everything charged to the container's cgroup (anonymous memory, page cache, kernel memory) and **`inactive_file`** is file-backed page cache the kernel considers cold. What remains is roughly "memory the container is actively using", and that still includes **active file cache**. ## RSS, cache and what the kernel can reclaim | Kind of memory | In working set? | In RSS? | Reclaimable at the limit? | |---|---|---|---| | Anonymous memory (heap, stacks) | yes | yes | not without swap | | Active file cache (recently read or mapped files) | yes | no | yes, clean pages can be dropped | | Inactive file cache | no | no | yes, first to go | | Memory-backed `emptyDir` contents | yes | no | no, it cannot simply be dropped | **RSS** here means `container_memory_rss`: the cgroup's anonymous memory (the `anon` figure on a cgroup v2 node). When the cgroup approaches its memory limit, the kernel first **reclaims** memory: it drops clean file pages and, where swap is available, swaps anonymous pages out. Only when reclaim cannot free enough does the kernel's OOM killer terminate a process in the container, and the kubelet records `OOMKilled`. ## The autocomplete example A search-autocomplete container has a 2Gi (2048Mi) memory limit and memory-maps a 1.1Gi prefix index. `kubectl top pod` shows **1870Mi**, about 91.3% of the limit, and the service has never been killed. Breaking it down: - RSS is around 640Mi: the heap and request buffers; - most of the rest is **active file cache** from the mapped index; - inactive cache is excluded from the working set already. The kernel can drop mapped index pages under pressure and read them back later, so the working set sits near the limit without danger. RSS is the number to compare with the limit when you ask "will this be OOMKilled?" ## Why working set still matters Working set is not a vanity metric: 1. **Node-pressure eviction** uses it. The kubelet's `memory.available` signal is node capacity minus the node's working set, so heavy cache on a node counts toward eviction decisions. 2. **Latency** suffers when the kernel reclaims *active* pages. If the index is evicted from cache, the next lookups must read it back from disk. These **refaults** add latency without any kill or restart, which is the same shape as CPU throttling and easy to confuse with it. 3. **Memory-backed volumes** count as used memory and cannot be dropped like cache, so a growing in-memory `emptyDir` pushes towards a real kill. On a node where swap is enabled for the pod, anonymous pages can also be moved out, so RSS alone is no longer the whole story; the reasoning above assumes the common case of no swap for the container. ## How to read a container near its memory limit - Compare **RSS** with the limit to judge OOM risk. - Compare **working set** with the limit to judge reclaim pressure and latency risk. - Check `kubectl describe pod` for an `OOMKilled` last state and a rising restart count; if both are absent, the kernel has been reclaiming successfully. - If p99 rises while the working set is pinned at the limit and CPU throttling is flat, suspect **cache reclaim** and consider a larger memory limit so the hot data stays cached. ## Common misreadings - **"91% working set means it is about to die."** Not if most of it is reclaimable cache. - **"Working set is the same as RSS."** Working set includes active page cache; RSS does not. - **"The OOM killer acts on the working set."** The kernel reclaims what it can first; a kill happens only when what remains, mostly anonymous memory, still exceeds the limit. - **"Cache is free, so the limit can be tight."** A tight limit squeezes hot cache and turns memory pressure into latency. - **"A restart would clear it."** Restarting drops the cache, which then refills as the index is read again, so the working set climbs back to the same place.
- The same container's p99 rises whenever its working set is pinned at the limit, but it is never killed. What is happening?The kernel is reclaiming active file pages, such as the memory-mapped index, to stay under the limit. The next lookups must read those pages back from disk, so latency rises without any kill or restart. Confirm that CPU throttling is flat, then give the container a larger memory limit so the hot index stays cached.
- Which memory figure does the kubelet use to decide it must evict pods from a node?Working set. The kubelet's `memory.available` eviction signal is node memory capacity minus the node's working set, so active page cache counts against the node. That is why a node can report memory pressure while much of the used memory is technically reclaimable cache.
saying these in an interview costs you the question
- Working set and RSS are two names for the same number
- A working set above 90% of the limit means an OOM kill is imminent
- The kernel kills the container as soon as usage touches the limit
- Page cache never counts toward a container's memory limit
- Memory-backed emptyDir data is dropped like ordinary file cache