A Linux server shows 3 GB of swap used in `free -h`, but users report no problem and the box feels responsive. Which measurements tell you whether swap is actually hurting this machine, and where do you read them?
answer
- a stock, not a flow
- sample it twice, subtract
- direction of the traffic matters
- stall time beats occupancy
- per-process field in status
basics
~20 sSwap used is a stock, not a rate: it can be cold residue paged out weeks ago. Judge swap by paging traffic instead — the pswpin/pswpout counters in /proc/vmstat over an interval — and by memory stall time in /proc/pressure/memory.
solid answer
~50 sThe `used` figure for swap is cumulative occupancy, not activity. Pages that were paged out during one nightly backup and never touched again still sit there, costing nothing. What hurts is *traffic*, so I measure a rate: sample `pswpin` and `pswpout` from /proc/vmstat a few seconds apart and look at the delta. Near zero means the swapped-out pages are cold and the box is fine; a sustained flow in both directions is thrashing, and it usually comes with `pgmajfault` climbing too. On a kernel with pressure stall information I go straight to /proc/pressure/memory: the `full avg10` figure is the share of the last ten seconds in which *everything* was stalled waiting on memory, and a non-trivial value there is the cleanest confirmation. To find the culprit I read `VmSwap` from /proc/<pid>/status across processes, and `swapon --show` to see which devices are in play and how full they are.
code
bash · 5 linesawk '/^(pswpin|pswpout|pgmajfault) /' /proc/vmstat
sleep 10
awk '/^(pswpin|pswpout|pgmajfault) /' /proc/vmstat
cat /proc/pressure/memory
swapon --showgo deeper
Know that swap being in use is normal and not automatically bad, and that the question to ask is whether the machine is actively paging right now rather than how much swap is occupied.
Be ready to name where the paging rate lives — the pswpin and pswpout counters in /proc/vmstat — and to explain that they are cumulative, so you sample twice and take the difference.
Demonstrate the full triage: rate first, then pressure stall information, then attribution to a process via VmSwap, and a check of which device backs swap. Say what each result would make you do next.
Own the monitoring policy this implies: swap occupancy is a diagnostic, stall time is the alert, and a fleet-wide page on "swap in use" is a defect that trains engineers to ignore memory alerts.
## The distinction that the whole question turns on A stock is how much is stored; a flow is how much moves per second. `free`'s swap `used` is a stock. A page that was evicted six weeks ago during a log rotation, and never referenced since, is *exactly* what swap is for: it freed a frame of RAM for something that actually wanted it. Seeing it counted forever is not a symptom. Thrashing, by contrast, is a flow — the same pages going out and coming straight back because the working set does not fit. That is what makes a machine feel dead while its CPUs look idle, because the time is spent waiting on storage, not computing. So the first move is never to react to the number in `free`. It is to measure a rate. ## Reading the rate /proc/vmstat carries cumulative counters since boot. Two matter here: - `pswpin` — pages read back in from swap - `pswpout` — pages written out to swap Sample them twice and subtract: ```bash awk '/^pswp/ {print}' /proc/vmstat; sleep 10; awk '/^pswp/ {print}' /proc/vmstat ``` A delta of essentially zero over ten seconds, on a box with gigabytes of swap occupied, is the signature of cold residue — nothing to do. Thousands of pages per second in *both* directions is thrashing. Pair it with `pgmajfault`, which counts faults that required disk I/O: a major-fault storm alongside swap-in traffic is the same story told twice. A useful asymmetry: heavy `pswpout` with little `pswpin` is the kernel making room, which may be a one-off and may be fine. Heavy `pswpin` means processes are actively touching memory that is no longer resident, and that is the painful direction. ## Pressure stall information: the cleanest signal On kernels with PSI (4.20 and later; some distributions require the `psi=1` boot parameter), /proc/pressure/memory reports how much time was lost to memory reclaim: ``` some avg10=12.51 avg60=8.02 avg300=3.14 total=98765432 full avg10=4.02 avg60=2.11 avg300=0.90 total=12345678 ``` `some` is the share of wall time in which at least one task was stalled on memory; `full` is the share in which *no* task could make progress. These are percentages over the trailing 10, 60 and 300 seconds. The virtue of PSI is that it measures the harm directly rather than a proxy: it catches file-cache thrashing on a swapless box just as well as it catches swap thrash, and it does not require you to interpret whether a given rate is "a lot". My rule of thumb during triage: `full avg10` sitting above a couple of percent means memory is genuinely costing this machine throughput, whatever `free` says. ## Attributing it to a process Once you know the box is paging, find who. Every process exposes `VmSwap` in /proc/<pid>/status, so a one-line sweep ranks them: ```bash for p in /proc/[0-9]*; do s=$(awk '/^VmSwap:/ {print $2}' "$p/status" 2>/dev/null) [ -n "$s" ] && [ "$s" -gt 0 ] && echo "$s ${p##*/} $(cat $p/comm 2>/dev/null)" done | sort -rn | head ``` For a finer view, /proc/<pid>/smaps_rollup reports both `Swap` and `SwapPss`, the latter dividing shared swapped pages among their users. And `swapon --show` lists each swap device with its `SIZE`, `USED` and `PRIO` — worth checking, because a swap file on the same spindle or the same throttled cloud volume as your database turns modest paging into an outage. ## What you conclude - Swap occupied, no paging traffic, PSI near zero: healthy. Leave it alone. Reclaiming that swap by cycling `swapoff`/`swapon` just forces the cold pages back into RAM you would rather use for cache. - Swap occupied, sustained `pswpin`, PSI `full` non-trivial: the working set does not fit. The fix is less memory demand or more RAM, not more swap. - No swap traffic but PSI still high: the pressure is on the *file* side — the kernel is evicting and refaulting page cache. Same conclusion about memory, different evidence. ## The reporting trap Alerting on "swap used > 0" is one of the most durable sources of false pages in the industry. Alert on stall time or on the paging rate, and let occupancy be a diagnostic detail rather than a trigger.
- Your monitoring pages on "swap used above 1 GB". What would you replace that alert with?With a stall-based or rate-based condition. `full avg10` from /proc/pressure/memory measures the harm directly and fires only when memory is actually costing throughput; a sustained `pswpin` rate from /proc/vmstat is a reasonable substitute on kernels without PSI. Occupancy stays on the dashboard as context for the investigation, but it never triggers a page on its own.
- PSI shows real memory pressure but the paging counters are flat. What is happening?The pressure is on file-backed memory. The kernel is evicting page cache and the workload is faulting the same pages straight back from disk — a refault loop that never touches swap, and which a swapless machine can suffer just as badly. Look at `pgmajfault` climbing and at the `some`/`full` split in /proc/pressure/memory; the remedy is still less memory demand or more RAM.
- Is running `swapoff -a` followed by `swapon -a` a reasonable way to clear accumulated swap?Rarely. `swapoff` must read every swapped page back into RAM before it can release the device, so on a box that is already tight it causes a large allocation burst and can itself trigger an OOM kill. If the pages are cold, you gain nothing by pulling them in; if they are hot, they would not still be in swap. Do it only during a maintenance window with headroom.
saying these in an interview costs you the question
- Treats any swap usage as a problem to fix
- Alerts on swap occupancy instead of paging rate
- Runs swapoff on a memory-tight production box
- Concludes thrashing without sampling a rate
- Ignores which device the swap file sits on