skip to content

Adding up the RSS that `ps` prints for every process on a Linux box gives a total far larger than the machine's physical memory. Why does that happen, and which per-process figure can you meaningfully add up instead?

level: seniorimportance: nice to knowfreq 32%

answer

  1. shared pages counted many times
  2. one library, forty tenants
  3. divide by the number of mappers
  4. the rollup file, not smaps
  5. walking page tables is not free

basics

~20 s

RSS charges every shared page in full to each process mapping it, so shared libraries, executable text, forked copy-on-write pages and shared memory are counted many times. PSS, in /proc/<pid>/smaps_rollup, splits each shared page across its mappers and can be summed.

solid answer

~50 s

Resident set size answers "how many physical pages is this process currently able to touch", and it makes no attempt to apportion sharing. A page of libc that forty processes have mapped is counted forty times, once in full per process. The same goes for executable text, tmpfs segments mapped by several services, and the copy-on-write pages a freshly forked child still shares with its parent. So summing RSS across a busy box routinely exceeds `MemTotal`, which tells you nothing is wrong — the metric was simply never additive. The additive figure is proportional set size: /proc/<pid>/smaps_rollup reports `Pss`, where each page is charged as its size divided by the number of processes mapping it, so the sum across all processes approximates real usage. The price is that reading smaps_rollup walks the process's page tables, which is far more expensive than reading `VmRSS`, and it still leaves out kernel-side memory such as slab and page tables.

code

bash · 3 lines
bash
ps -eo rss= | awk '{t+=$1} END {print "sum of RSS (kB):", t}'
awk '/^MemTotal:/ {print "MemTotal  (kB):", $2}' /proc/meminfo
grep -E '^(Rss|Pss|Private_Dirty|Swap):' /proc/self/smaps_rollup

go deeper

for a junior

Know that RSS counts shared pages in each process that maps them, so the same library memory appears in many processes at once and the totals cannot simply be added together.

for a middle

Be able to name the sources of the double counting — shared libraries, executable text, copy-on-write after fork, shared memory segments — and to point at PSS as the apportioned alternative.

for a senior

Show that you know where PSS lives, what it costs to read, and which figure you would actually put on a per-service memory chart, including why Private_Dirty is often the most honest one.

for a principal

Own the accounting story end to end: which memory is attributable to services at all, what stays kernel-side and never appears in any process, and what that means for the capacity model and chargeback you ask teams to trust.

## Why the sum overshoots RSS is a per-process count of resident physical pages, and it is deliberately naive about sharing. Every page currently mapped into the process and backed by RAM is counted, in full, regardless of how many other processes are mapping the same physical frame. Four sources of over-counting dominate on a real server: - **Shared libraries.** libc, OpenSSL, libstdc++ — one physical copy, mapped by every process that links them, charged in full to each. - **Executable text.** Fifty worker processes forked from the same binary map one physical copy of that binary's text. - **Copy-on-write after fork.** A child that has forked but not yet written to a page shares the parent's frame. Both show it in RSS; the machine holds one. - **Shared memory.** tmpfs, /dev/shm and System V segments mapped by several processes — a database's shared buffers being the classic case. So the sum of RSS is an upper bound with no defined relationship to physical memory. Comparing it to `MemTotal` is a category error, not a discovery. ## The additive alternative: PSS Proportional set size fixes exactly this. Each resident page is charged to a process as `page size / number of processes mapping it`. A 4 KiB libc page mapped by 40 processes contributes 102 bytes to each. Sum PSS across every process and you get a figure that stays under `MemTotal` and moves sensibly when a workload grows. It lives in /proc/<pid>/smaps_rollup, a single pre-aggregated summary of all the process's mappings: ```bash grep -E '^(Rss|Pss|Private_Clean|Private_Dirty|Swap|SwapPss):' /proc/self/smaps_rollup ``` Useful neighbours in the same file: - `Private_Dirty` — pages this process alone has modified. This is the memory that genuinely disappears only when the process exits, and it is often the most honest single number for "what is this service costing me". - `Swap` and `SwapPss` — the swapped-out equivalents, the latter apportioned the same way. `smaps_rollup` has been available since Linux 4.14. Before it, tools had to parse the full /proc/<pid>/smaps, which lists every mapping separately and is far slower on a process with thousands of VMAs. ## What it costs, and why RSS still exists `VmRSS` in /proc/<pid>/status is a counter the kernel maintains anyway; reading it is nearly free. Computing PSS requires walking the process's page tables and consulting the map count of each page. On a large process — a JVM with tens of gigabytes resident, or a database with a huge shared segment — a single `smaps_rollup` read can take a noticeable amount of CPU and briefly contend with the process's own memory operations. That is why `ps` and `top` still show RSS by default, and why you should not have a monitoring agent scraping PSS for every process every ten seconds. The practical policy: RSS for continuous, cheap sampling and for spotting *trends* in a single process; PSS when you need to attribute a fixed amount of RAM among several processes and have the answers add up. ## The reconciliation never comes out exact Even a perfect PSS sum will not equal `MemTotal - MemFree`, because a process's page tables describe only userspace mappings. Missing from the accounting: - kernel slab caches (`Slab`, `SReclaimable`, `SUnreclaim` in /proc/meminfo) - the page tables themselves (`PageTables`) - kernel stacks (`KernelStack`), network socket buffers, filesystem metadata caches - page cache not currently mapped by anyone — usually the largest item of all So the honest workflow is top-down and bottom-up at once: /proc/meminfo tells you where memory went in aggregate, and per-process PSS tells you which service owns the userspace share of it. When those two disagree by a lot, the answer is usually kernel-side — a runaway slab cache, or an enormous page-table footprint from many processes mapping a huge shared segment. ## Also worth remembering RSS excludes anything swapped out. A process that has been paged out heavily can show a *falling* RSS while its true footprint is unchanged; `VmSwap` from /proc/<pid>/status is the other half of that picture, and a leak hunt that watches only RSS on a swapping box will draw the wrong conclusion.

  • Which single per-process number would you watch to decide whether a service is genuinely leaking?
    `Private_Dirty` from /proc/<pid>/smaps_rollup, or `RssAnon` from /proc/<pid>/status as the cheap approximation. Both isolate memory that belongs to this process alone and cannot be shared away or dropped like clean file pages. A steady rise in either across a stable request rate is the shape of a leak; total RSS moving is not, since it also tracks page cache the process happens to have mapped.
  • Why does `top` show RSS rather than PSS by default?
    Because RSS is a counter the kernel already maintains, so reading it for hundreds of processes several times a second is essentially free. PSS requires a page-table walk per process, which on a large process is expensive enough to perturb the machine you are measuring. `top` optimises for cheap continuous refresh; PSS is a deliberate, occasional query.

saying these in an interview costs you the question

  • Sums RSS across processes and calls it total usage
  • Thinks RSS excludes shared library pages
  • Believes PSS is as cheap to read as RSS
  • Expects per-process figures to reconcile with MemTotal
  • Ignores swapped-out memory when tracking a leak

context