You are setting the memory policy for a fleet of Linux application servers and a colleague proposes running them all with swap disabled. What does the vm.swappiness setting actually control, and how would you decide?
answer
- a bias, not a switch
- two pools to reclaim from
- zero is not off
- the pressure does not disappear
- alert on stall, not occupancy
basics
~20 svm.swappiness is a relative weighting between reclaiming anonymous pages and file-backed pages, not an on/off switch or a percentage of RAM. Disabling swap does not remove memory pressure; it leaves the kernel only file pages to evict, so thrash reappears as executable-text refaults.
solid answer
~50 s`vm.swappiness` biases which side the kernel reclaims from when it needs pages: higher values make it more willing to push anonymous memory to swap, lower values push it toward evicting file-backed page cache. The default is 60, the range is 0-100 on older kernels and 0-200 since Linux 5.8, and 0 does not disable swap — it only tells the kernel to avoid it until it is nearly out of options. Disabling swap entirely does not remove memory pressure; it removes one of the kernel's two reclaim targets, so under pressure it must evict file pages instead, and you get refault thrash on executable text and cache with a completely idle swap device to prove your innocence. My decision is measurement-driven: keep a modest swap area, treat occupancy as a diagnostic rather than an alert, and drive alerting from memory stall time. Go swapless only where something specific demands it and you have accepted that OOM kills become the failure mode.
code
bash · 4 linessysctl vm.swappiness
swapon --show
awk '/^(pgmajfault|pswpin) /' /proc/vmstat
cat /proc/pressure/memorygo deeper
Know that swappiness expresses a preference between reclaiming cached file pages and swapping out anonymous memory, and that setting it to 0 does not turn swap off.
Be able to explain the two reclaim pools and what changes when one of them is removed — a swapless box under pressure evicts and refaults page cache instead, with an idle swap device.
Argue the tradeoff from operational evidence: what the failure looks like with and without swap, why occupancy is a bad alert, and which stall metrics you would put in its place.
Own the fleet-wide position and its consequences: the failure mode you are choosing for out-of-memory conditions, how it is measured, where exceptions are granted, and how you keep the debate anchored in stall data instead of preference.
## What swappiness is, and what it is not When the kernel needs free pages it can reclaim from two pools: - **file-backed pages** — page cache holding file contents, including your executables' text. Clean ones can be dropped instantly; the cost is a future major fault to read them back. - **anonymous pages** — heap, stack, private mappings. Nothing on disk backs them, so the only way to reclaim one is to write it to swap. `vm.swappiness` is the knob that biases the choice between those two pools. It is **not** a percentage of memory, **not** a threshold at which swapping begins, and **not** an enable/disable switch. Key values: - **60** — the default: a moderate willingness to swap anonymous memory. - **0** — do not swap anonymous memory unless reclaim is otherwise failing. It does not disable swap; the kernel will still swap rather than OOM-kill. - **100** — treat the two pools as roughly equal in cost. - **up to 200** — available since Linux 5.8, for cases where swap is on storage fast enough that swapping out anonymous memory is genuinely cheaper than losing cache. ```bash sysctl vm.swappiness sysctl -w vm.swappiness=10 # runtime only; persist under /etc/sysctl.d/ ``` ## The argument for turning swap off, stated fairly It is not a foolish position. The case runs: a latency-sensitive service that has been paged out is effectively down anyway, so a fast, loud OOM kill and restart beats a slow, quiet death; supervision will restart it; capacity planning should have prevented the shortage in the first place; and a swapless machine has one fewer confusing variable during an incident. ## The argument against Disabling swap does not create memory. It removes one of the two reclaim targets, and everything the kernel would have done to anonymous memory it now does to file pages instead. The practical consequences: - **Refault thrash on the wrong pages.** Under pressure the kernel evicts page cache — including the mapped text of your own binaries and libraries. The process then faults it straight back from disk. The machine crawls, CPU sits in iowait, and the swap device is untouched, so anyone who equates "memory trouble" with "swap in use" concludes memory is fine. - **Cold anonymous memory is stranded in RAM.** Every long-lived process holds start-up allocations, one-time configuration structures and idle thread stacks that are never touched again. With swap, those quietly move out and their frames become cache. Without it, they occupy premium RAM permanently. - **The failure mode becomes abrupt.** No swap means the runway between "tight" and "OOM kill" shortens dramatically, and the victim is chosen by the kernel rather than by you. ## The decision I would actually make **Default: keep a modest swap area and instrument the pressure.** Enough to absorb cold anonymous memory rather than enough to run a workload from — sizing swap to "twice RAM" is folklore from a different era. Then: - Alert on **memory stall time** from /proc/pressure/memory (`full avg10`) on kernels 4.20 and newer, never on swap occupancy. Occupancy generates permanent false alarms and teaches people to ignore memory pages. - Keep `vm.swappiness` low but non-zero (10 is a common choice) on latency-sensitive application servers: it prefers dropping cache but leaves the escape hatch open for genuinely cold anonymous pages. - Put swap on local storage, never on a network volume or a throttled cloud disk, where paging turns a slowdown into an outage. - Consider **zram** — a compressed block device in RAM used as swap, managed with `zramctl` — on memory-constrained hosts. Cold anonymous pages compress well and "paging" costs CPU rather than disk I/O. **Go swapless deliberately, in specific cases**: nodes running a workload that already manages its own memory ceiling and expects to be killed on breach; systems whose scheduler or orchestrator explicitly requires it; hosts where the only available swap device is unacceptably slow. **Either way, pin memory that must never be paged.** A process holding secrets or a latency-critical data structure should use `mlock`, which is precise, rather than depending on a machine-wide swap policy to protect it. ## How I would present it to the team The question "swap on or off" is the wrong axis, and that framing is what makes the debate perennial. The real question is *what happens when this machine runs short of memory* — how long the degradation lasts, how visible it is, and who decides which process dies. Swap changes the shape of that failure; it does not change whether it happens. Fix the shortage with capacity and with limits; use swap and swappiness to choose how the edge case degrades, and measure stall time so the choice is evaluated on evidence rather than on preference.
- A team sets vm.swappiness=0 and reports the box still swapped. Were they misled?No, that is correct behaviour. Since the semantics were tightened, 0 means the kernel avoids swapping anonymous memory until reclaim is otherwise failing — it still prefers swapping to an OOM kill. If the intent was truly never to swap, the honest expression of that is disabling swap and accepting OOM kills, or `mlock` on the specific memory that must stay resident.
- On a swapless machine that is short of memory, what would you expect the symptoms to look like?High iowait and poor responsiveness with a completely idle swap device. The kernel can only reclaim file pages, so it evicts mapped executable text and cache and the workload faults it straight back — visible as climbing `pgmajfault` in /proc/vmstat and elevated `some`/`full` values in /proc/pressure/memory. Anyone using swap usage as the memory alarm sees nothing at all.
- How would you size a swap area for an application server today?Small and deliberate — enough to absorb cold anonymous memory, not enough to run a workload out of. A few gigabytes is typical regardless of RAM size; the old twice-RAM rule dates from hibernation requirements and tiny memories. Put it on local storage, and size it by observing how much `VmSwap` accumulates across your fleet during steady state rather than from a formula.
- Where does zram fit into this decision?It gives you a compressed swap device that lives in RAM, managed with `zramctl`, so evicting a cold anonymous page costs CPU and compression ratio instead of disk I/O. On memory-constrained hosts it recovers real capacity from pages that compress well, and it sidesteps slow or network-attached storage entirely. It is not a substitute for capacity — the compressed pages still occupy RAM.
saying these in an interview costs you the question
- Calls swappiness a percentage of RAM before swapping starts
- Thinks vm.swappiness=0 disables swap entirely
- Assumes no swap means no memory pressure
- Sizes swap at twice RAM by reflex
- Alerts on swap usage rather than stall time