On a Linux host, the `available` column of `free -h` comes from MemAvailable in /proc/meminfo, and it is always smaller than MemFree plus Cached. What does the kernel refuse to count as available, and why?
answer
- not free plus cached
- the kernel keeps a reserve
- some cache is pinned
- tmpfs pages live in Cached
- dirty pages need writeback first
basics
~20 sMemAvailable estimates how much memory a new workload could get without swapping. The kernel holds back its free-page watermark and part of the page cache, and excludes cache it cannot simply drop: tmpfs/shmem pages, mlocked pages, and dirty pages awaiting writeback.
solid answer
~50 s`free` is just a formatter over /proc/meminfo: `buff/cache` is `Buffers + Cached + SReclaimable`, and `available` is the kernel's own `MemAvailable` estimate rather than arithmetic done by the tool. The kernel builds that estimate conservatively. It first subtracts the low watermark of free pages it must keep in reserve to avoid stalling allocations, then counts only part of the page cache and part of the reclaimable slab — roughly half, on the assumption that the rest is hot and would immediately refault. On top of that, some memory sitting in `Cached` is not droppable at all: `Shmem` pages (tmpfs, /dev/shm, System V shared memory) have no backing file to be written to, so they can only go to swap; `Mlocked`/`Unevictable` pages are pinned; and `Dirty` pages must be written back before the frames can be freed. So `available` is a deliberately pessimistic answer to "can I start a 4 GB process here", which is exactly why it is the number to trust.
code
bash · 2 linesgrep -E '^(MemTotal|MemFree|MemAvailable|Buffers|Cached|Shmem|Dirty|SReclaimable):' /proc/meminfo
free -w -hgo deeper
Know that the available column, not free, answers "is there room here", and that buff/cache is memory being put to work rather than memory lost. Say plainly that a low free on a busy server is normal.
Be ready to explain where each free column comes from in /proc/meminfo and why the kernel computes MemAvailable conservatively — a reserved watermark plus only part of the page cache and reclaimable slab.
Show that you know which cached memory is not reclaimable at all — Shmem/tmpfs, mlocked and dirty pages — and that you check Shmem and Mlocked when available collapses without used moving.
Own the guidance others follow: which single number goes on the memory dashboard and the alert, why alerting on free generates permanent noise, and what threshold on available is meaningful for the workload profile you run.
## What `free` is actually printing `free` does almost no work of its own — it reads /proc/meminfo and lays the fields out in columns. On procps-ng the mapping is: - `total` = `MemTotal` - `free` = `MemFree` - `buff/cache` = `Buffers` + `Cached` + `SReclaimable` - `available` = `MemAvailable` - `used` = `total - free - buff/cache` `free -w` splits buffers and cache into two separate columns if you want to see them apart. Notice that `used` is a derived leftover, not something the kernel reports, and `available` is the one column that the *kernel* computes rather than the tool. ```bash grep -E '^(MemTotal|MemFree|MemAvailable|Buffers|Cached|Shmem|Dirty|SReclaimable):' /proc/meminfo ``` ## Why the kernel does not just add MemFree and Cached MemAvailable was added precisely because everyone was doing that addition by hand and getting it wrong. The kernel's estimate (`si_mem_available()` in the memory management code) starts from `MemFree`, subtracts the watermark of pages it must keep free so that allocations do not stall or trigger reclaim, and then adds back only a *portion* of the reclaimable pools: - page cache, minus a held-back share (on the order of half, bounded by the low watermark) - reclaimable slab (`SReclaimable`, the dentry and inode caches among others), with the same kind of hold-back The hold-back exists because reclaiming the last of the cache is not free. Those pages are being read; evicting them means the next access is a major fault back from disk. A machine that technically "gets" the memory by evicting its entire working set has not gained anything. ## The cache that cannot be handed over at all This is the part candidates miss. `Cached` is not a synonym for "droppable": - **Shmem** — tmpfs files, /dev/shm, POSIX and System V shared memory. These pages appear inside `Cached`, and there is no file on disk behind them, so they cannot be dropped. Their only exit is swap. A build directory on tmpfs or a large /dev/shm segment inflates `Cached` while lowering `available`. - **Unevictable / Mlocked** — pages pinned by `mlock()`, typical of databases and anything holding key material out of swap. - **Dirty / Writeback** — modified pages that must reach storage before the frame is reusable. They are reclaimable eventually, not now. A quick demonstration: write half a gigabyte into /dev/shm and watch `Cached` and `Shmem` climb together while `MemAvailable` *falls* by the same amount. Page cache from reading a real file behaves the opposite way — `Cached` climbs and `available` barely moves. ## How to use it during an incident The operational rule is: **compare `available` against the size of the thing you are about to start, and ignore `free` entirely.** A 32 GB box with 200 MB free and 20 GB available is a healthy box doing its job. The signals that actually mean trouble are `available` trending toward zero over time, or `available` collapsing while `used` stays flat (something is pinning cache — check `Shmem` and `Mlocked`). For a per-process breakdown, /proc/<pid>/status carries `VmRSS` split into `RssAnon`, `RssFile` and `RssShmem`, plus `VmSwap`. Anonymous memory is the part that only swap or an OOM kill can reclaim, so `RssAnon` growth across restarts is the shape of a genuine leak. ## Traps worth naming out loud - **`echo 3 > /proc/sys/vm/drop_caches` is not a fix.** It makes `free` look better and makes the machine slower, because the working set refaults from disk immediately. It is a benchmarking tool, not an operational remedy. - **`used` is not "what the applications are using."** tmpfs contents land in it, and it is a subtraction rather than a measurement. - **Inside a container, `free` still reads the host's /proc/meminfo.** It reports the whole machine, not the limit the container is running under, so a containerised process can be killed for exceeding its limit while `free` cheerfully shows tens of gigabytes available. Read the container's own accounting rather than `free` in that case. - **MemAvailable is an estimate, not a guarantee.** It says nothing about fragmentation, and a huge-page or large contiguous allocation can still fail with plenty "available".
- A colleague clears the cache with drop_caches every night to keep `free` looking healthy. What do you tell them?That it reliably makes the machine slower. Dropping the page cache does not release memory anything was waiting for — `available` already counted that cache as obtainable. All it does is force the working set to refault from disk, so you trade a cosmetic number for a burst of major page faults and I/O. It is a benchmarking aid, not an operational remedy.
- Someone mounts a large build directory on tmpfs and now `available` is far below what they expect. Where does that memory show up?In `Cached` and, specifically, in `Shmem` in /proc/meminfo — tmpfs pages are accounted as page cache but have no backing file, so the kernel cannot drop them and excludes them from MemAvailable. They can only be pushed to swap. Compare `Shmem` before and after; `df -h` on the tmpfs mount tells you how much of it is actually occupied.
saying these in an interview costs you the question
- Says available is just free plus buff/cache
- Calls page cache leaked memory that must be freed
- Recommends drop_caches as a production fix
- Assumes everything in Cached can be dropped instantly
- Reads used as the applications' memory footprint