skip to content

On a Linux box running a pre-forking web server with 40 worker processes, adding up the RSS column from `ps aux` gives about 60 GB on a machine with 16 GB of RAM, and the machine is not swapping. Why can that sum exceed physical memory, and what would you measure instead?

level: seniorimportance: nice to knowfreq 32%

answer

  1. per-process counters are not additive
  2. forked workers share the parent's pages
  3. one physical page, forty page tables
  4. divide each page by its sharers
  5. smaps_rollup has the field you want

basics

~20 s

RSS counts every resident page a process maps, including pages shared with other processes, so forked workers each report the same shared libraries and copy-on-write parent pages. Summing RSS double-counts them. Use PSS from /proc/<pid>/smaps_rollup, which divides each shared page among its users.

solid answer

~50 s

RSS is per-process, not per-page-owner: it counts every page of that process's address space that is currently resident in RAM, whether or not other processes are mapping the same physical page. Forty workers forked from one parent share the parent's copy-on-write pages, the same shared libraries, and the same file-backed pages of the binary, so each worker's RSS includes that shared footprint in full. Adding the column up counts those physical pages up to forty times, and the result can exceed the machine's RAM without anything being wrong. What I actually want is PSS — proportional set size — where each page is divided by the number of processes mapping it, so the PSS values do sum to something meaningful. It is exposed per process in `/proc/<pid>/smaps_rollup`, alongside USS-style private counters that tell you how much would really be freed if you killed that one worker.

code

bash · 1 line
bash
sudo grep -H '^Pss:' /proc/*/smaps_rollup 2>/dev/null | sort -t: -k3 -n | tail -n 10

go deeper

for a junior

Know that RSS is the memory a process currently has resident, and that several processes can be resident on the same physical pages — so the column is not something you add up.

for a middle

Be able to explain copy-on-write after fork and shared library mappings as the reason for the double count, and name proportional set size as the field that does sum correctly.

for a senior

Show that you would confirm the pattern (sibling workers, common parent, near-identical RSS), then read PSS and private-dirty from the per-process rollup to separate a real per-worker leak from shared-image accounting.

for a principal

Own the measurement contract: decide which memory figure the organisation alerts on and capacity-plans with, and be able to explain why naive per-process totals overstate usage on any forking or heavily shared workload.

## What RSS counts Resident set size is the number of pages of a process's virtual address space that are currently present in physical memory, converted to kibibytes for the `RSS` column of `ps` and the `RES` column of `top`. The definition is entirely process-centric: the kernel walks that process's page tables and counts resident mappings. It does not ask whether a physical page is also mapped by somebody else. That is fine when you look at one process. It falls apart the moment you add the column up, because on Linux a great deal of memory is legitimately mapped by many processes at once: - **Copy-on-write pages after fork.** `fork()` does not copy the parent's memory; it marks the pages read-only and shares them until one side writes. In a pre-forking server the parent may have loaded a large configuration, template cache, or interpreter heap before forking, and every worker reports all of it as resident. - **Shared libraries.** libc, OpenSSL, the language runtime — one physical copy in the page cache, mapped into every process, counted in every RSS. - **File-backed text pages.** The executable's own code, likewise shared. - **Explicitly shared memory.** `tmpfs` files, `shm` segments, `MAP_SHARED` mappings — one region, many mappers, counted once per mapper. So the 60 GB is not a measurement error, it is the wrong arithmetic on a correct measurement. ## PSS: the number that adds up Proportional set size solves exactly this. For each resident page, PSS adds `page_size / number_of_processes_mapping_it` to the process's total. A page private to one process contributes fully; a page shared by 40 workers contributes a fortieth to each. The consequence is the useful property: **the sum of PSS across all processes approximates real physical memory in use**, without double counting. The kernel exposes it per process in two places: ``` grep -H ^Pss: /proc/*/smaps_rollup 2>/dev/null | sort -t: -k3 -n | tail ``` `/proc/<pid>/smaps_rollup` gives one pre-summed set of counters for the whole address space, which is cheap to read. `/proc/<pid>/smaps` gives the same fields broken down per mapping, which is what you read when you want to know *which* region is large — but it is expensive on a process with thousands of mappings, so prefer the rollup for scanning. Alongside `Pss` you will find `Rss`, `Shared_Clean`, `Shared_Dirty`, `Private_Clean`, `Private_Dirty` and `Swap`. The private counters are what people mean by USS, the unique set size: the memory that would actually be returned to the system if you killed this one process. For a fleet of forked workers, private-dirty is the interesting number — it is the memory each worker has genuinely made its own by writing to shared pages. If private-dirty grows steadily per worker, you have a real per-worker leak; if it stays flat while RSS looks huge, the workers are just sharing the parent's image as designed. Reading these files for another user's processes requires privilege, so this is usually a `sudo` operation, and tools such as `smem` exist to do the walk and present PSS and USS in one table. ## Two more ways RSS misleads - **RSS omits swapped-out pages.** A process being swapped out shows a *falling* RSS, which reads like memory being released when the opposite is happening. The `Swap` line in the rollup is what tells you where those pages went. - **RSS says nothing about reclaimability.** Much of a large RSS can be clean, file-backed pages the kernel can drop instantly under pressure. A 4 GB RSS that is mostly clean page cache mappings is a very different situation from 4 GB of private dirty anonymous memory, and only the breakdown distinguishes them. ## How to answer this at the terminal The practical sequence is short. Confirm the shape with `ps`: ``` ps -eo pid,ppid,rss,args --sort=-rss | head -n 20 ``` If the large entries are siblings with a common parent and near-identical RSS, that pattern alone is the tell — forked workers sharing a parent image. Then read PSS from the rollup for a handful of them, and compare private-dirty across workers. If PSS per worker is a small fraction of its RSS, the sum was an illusion and the machine's memory accounting is fine. If private-dirty is large and climbing, you have found the actual consumer, and the question becomes what each worker is allocating for itself. The interviewer is testing one specific instinct here: whether you know that per-process memory columns are not additive, and whether you know the field that is.

  • Which number tells you how much memory you would actually get back by killing one of those workers?
    The private counters in that process's smaps rollup — private-clean plus private-dirty, often called the unique set size. Those are the pages nobody else maps, so they are freed on exit. The shared portion of its RSS stays resident because the other 39 workers still map it. That is why killing one worker of a pre-forking pool usually frees far less than its RSS suggests.
  • Why does reading /proc/<pid>/smaps for every process on a busy host cost more than reading RSS?
    RSS is a single pre-computed counter the kernel already maintains, while smaps walks every virtual memory area and computes sharing counts per page range. On a process with thousands of mappings that is expensive, and doing it fleet-wide can measurably disturb the machine. The rollup file exists precisely to give you the same totals in one pass, so scan with the rollup and only open the full smaps for the process you have already singled out.
  • A monitoring agent alerts on a process whose RSS is steadily falling. Is that good news?
    Not necessarily. RSS counts only resident pages, so it falls when pages are swapped out or reclaimed as well as when memory is genuinely released. A falling RSS on a host under memory pressure often means the process is being paged out, which is a worse situation than the flat line that preceded it. The swap counter for that process is what distinguishes the two.

Forty flatmates each reporting the full rent of the shared flat: every individual figure is honest, but adding them up invents money that nobody paid.

saying these in an interview costs you the question

  • Adds RSS across processes to estimate total memory use
  • Believes RSS counts only memory unique to that process
  • Assumes killing a worker frees its whole RSS
  • Reads falling RSS as memory being released
  • Thinks shared library pages are copied per process

context