skip to content

A Linux application server with swap enabled becomes extremely slow while its CPUs sit mostly idle and its swap usage climbs. What is the kernel doing, and why does adding more swap not fix a memory shortage?

level: seniorimportance: should knowfreq 50%

answer

  1. working set larger than RAM
  2. evicted then immediately needed
  3. blocked on I/O, not computing
  4. traffic matters, not swap occupancy
  5. swappiness rebalances, it does not add memory

basics

~20 s

The kernel is reclaiming pages that are still in use and immediately faulting them back from disk — thrashing. The CPUs are idle because tasks are blocked on I/O. Swap buys time and absorbs cold pages; it does not add usable memory to a working set that exceeds RAM.

solid answer

~50 s

Under memory pressure the kernel reclaims pages: it drops clean file-backed pages and writes anonymous pages out to swap. That works when the evicted pages are genuinely cold. When the working set is larger than RAM, the pages evicted are ones the application needs again immediately, so each access takes a **major fault** — a disk read on the critical path. The process spends its time blocked in uninterruptible I/O sleep, which is exactly why the machine looks slow with idle CPUs and no obvious runaway process. `vm.swappiness` (default 60) only sets the *relative* cost the kernel assigns to reclaiming anonymous pages versus file pages; it shifts which pages get evicted, it does not change how much memory exists. Swap is best understood as a place to park pages that are truly never touched, and as a shock absorber that turns a hard out-of-memory kill into gradual degradation. Fixing the shortage means shrinking the working set or adding RAM.

code

bash · 3 lines
bash
grep -E '^(pswpin|pswpout|pgmajfault) ' /proc/vmstat
cat /proc/pressure/memory
cat /proc/sys/vm/swappiness

go deeper

for a junior

Know that swap is disk used to hold memory pages the kernel evicted, that reading them back is far slower than RAM, and that swap does not increase how much memory the machine really has.

for a middle

Explain reclaim and major faults: the kernel evicts pages it believes are cold, and if the application needs them again each access becomes a disk read, so throughput collapses while CPUs wait on I/O.

for a senior

Demonstrate diagnosis and judgment — distinguish healthy parked swap from a swap-in storm, cite pressure stalls and major-fault rates, and argue what to change: working-set size, headroom, or placement, not just the swappiness knob.

for a principal

Own the policy question across a fleet: whether hosts run with swap as a shock absorber or without it for crisp failure, how that interacts with cgroup-scoped limits and orchestration, and which pressure signal triggers shedding load or rescheduling.

## What reclaim actually does When free pages fall below the kernel's low watermark, reclaim starts. The kernel walks LRU lists and frees pages in two categories: - **File-backed pages** — page cache. A clean one is dropped for free; a dirty one is written back first. - **Anonymous pages** — heap, stack, and anything with no file behind it. These have nowhere to go except a swap device, so with no swap configured they are simply not reclaimable at all. So swap is not "extra RAM". It is a backing store that makes anonymous pages eligible for reclaim, exactly as a file makes file pages eligible. ## Thrashing, and why the CPUs look idle Reclaim is a bet that the evicted page will not be needed soon. When the working set — the set of pages actually being used in a given interval — fits in RAM, the bet is good. When it does not, the kernel evicts a page and the application touches it milliseconds later. That access is a **major fault**: the faulting task blocks in uninterruptible sleep while the kernel reads the page back, then something else must be evicted to make room. The system converts memory pressure into disk I/O at the worst possible place, on the critical path of every memory access. The signature is distinctive and worth naming in an interview: - throughput collapses, latency goes to seconds, but CPU utilisation is *low*, because the work is blocked, not computing; - swap-in traffic is as high as swap-out — steady-state swap-out alone is healthy, but pages coming back in mean the eviction was wrong (`pswpin`/`pswpout` in `/proc/vmstat` count them); - on kernels with `CONFIG_PSI` (4.20+), `/proc/pressure/memory` shows a large share of time stalled on memory — the cleanest single signal that this is memory pressure and not a slow database or a network problem. The same mechanism explains the classic "why is my swap full but the machine is fine" case: cold pages parked in swap and never read back cost nothing. **Swap usage is not the symptom; swap traffic is.** ## vm.swappiness, correctly stated `vm.swappiness` is routinely described as "how eager the kernel is to swap", which invites the belief that lowering it eliminates swapping. It is better described as the kernel's assumed *relative I/O cost of reclaiming anonymous memory versus file-backed memory*. It biases the split between the two LRU lists during reclaim. The default is 60. Raising it favours evicting anonymous pages and keeping cache; lowering it favours evicting cache and keeping anonymous pages resident. Setting it to 0 does not disable swap — it tells the kernel to avoid anonymous reclaim until it is nearly out of alternatives, which in practice means the machine will chew through its page cache first and then still swap rather than die. On fast NVMe storage, and especially with compressed swap in RAM (`zram`) or a compressed cache in front of the swap device (`zswap`), the cost assumption behind a low swappiness is often wrong, and modest swapping is genuinely cheap. ## Should a server have swap at all? Both extremes are defensible and the interviewer wants the reasoning: - **Swap off** makes behaviour crisp: the machine either has memory or the out-of-memory killer runs, with no long grey zone where everything is alive but unusably slow. It also removes anonymous pages from the reclaim toolkit entirely, so pressure escalates to a kill much sooner, and genuinely idle pages are stuck occupying RAM. - **Swap on** gives the kernel somewhere to put cold anonymous memory, which is a real win for machines running long-lived processes with large initialisation-time allocations, and it converts a cliff into a slope — often the difference between a degraded service you can shed load from and a dead one. A good middle answer is a modest swap device plus an explicit pressure alert, so degradation is detected long before it becomes a stall. What you should not do is size swap as if it were memory: if the working set exceeds RAM, more swap only lengthens the period of thrashing before the inevitable. ## Fixing it Shrink the working set (bound caches and per-request buffers, cap the runtime's heap), give the machine reclaimable headroom, move a tenant off the box, or add RAM. Tuning `vm.swappiness` redistributes the pain; it does not create memory.

  • A server shows 4 GB of swap in use but no performance problem. Is that a concern?
    Not by itself. Occupied swap means pages were evicted at some point; if they are never read back, that memory was cold and the kernel made a good trade. The metric that matters is swap *traffic* — pages faulting back in — together with major-fault counts and memory pressure stalls. Alerting on swap occupancy generates noise; alerting on swap-in rate finds real trouble.
  • Does setting vm.swappiness to 0 disable swapping?
    No. It tells the kernel that reclaiming anonymous memory is very expensive relative to reclaiming file cache, so it will evict page cache aggressively first. When reclaim still cannot keep up it will swap anyway rather than fail. If you truly want no swapping, remove the swap device — and accept that anonymous pages then become unreclaimable, so pressure escalates straight to an out-of-memory kill.
  • How do compressed swap mechanisms change this calculus?
    `zram` presents a compressed block device in RAM as swap, and `zswap` puts a compressed cache in front of a real swap device. Both replace disk I/O with CPU work on the reclaim path, so a swap-in costs microseconds of decompression instead of milliseconds of storage latency. They effectively stretch RAM for compressible anonymous data, at the cost of CPU and of a much smaller safety margin than a real device.
  • Why can the system look idle while it is thrashing?
    Because the tasks are not runnable — they are blocked waiting for page-ins from the swap device, an uninterruptible I/O sleep. CPU utilisation measures time spent executing, so a machine where every task is waiting on storage shows low utilisation and terrible latency simultaneously. Load average, in contrast, rises, because Linux counts uninterruptible sleepers in it.

saying these in an interview costs you the question

  • Treats swap as extra RAM that raises capacity
  • Alerts on swap occupancy rather than swap traffic
  • Claims swappiness 0 disables swapping entirely
  • Concludes CPUs are idle so memory is fine
  • Adds more swap to fix an oversized working set

context