A process on Linux allocates a large buffer, the allocation returns successfully, and the process is later killed while writing into that buffer. Why can an allocation succeed when the memory is not actually available, and which sysctl governs that behaviour?
answer
- allocation reserves, touch commits
- the kernel promises more than it has
- failure moves to fault time
- three modes: heuristic, always, strict
- CommitLimit versus Committed_AS
basics
~20 sLinux overcommits: an allocation only creates a mapping, and physical pages are committed on first touch. The kernel may promise more than it can back, so the shortage surfaces later as a fault it cannot satisfy. vm.overcommit_memory selects the policy.
solid answer
~50 sAllocating memory on Linux does not obtain memory; it obtains *address space*. `malloc` ends in `brk` or `mmap`, which records a mapping, and the first write to each page faults so the kernel can attach a physical page. Because most programs reserve far more than they touch — heap arenas, sparse buffers, copy-on-write pages after `fork` — the kernel deliberately promises more than it has. That is **overcommit**, controlled by `vm.overcommit_memory`: mode `0` is the default heuristic, refusing only wildly implausible single requests; mode `1` always says yes; mode `2` is strict accounting, refusing once total commitments would exceed `CommitLimit`, derived from swap plus `vm.overcommit_ratio` percent of RAM. Under the first two modes the failure moves from the allocation, where the program could handle it, to a page fault, where it cannot — so the kernel invokes the OOM killer instead. Mode 2 trades that for honest `ENOMEM` at allocation time, at the cost of provisioning for reservations nobody touches.
code
bash · 3 linescat /proc/sys/vm/overcommit_memory
cat /proc/sys/vm/overcommit_ratio
grep -E '^(MemTotal|SwapTotal|CommitLimit|Committed_AS):' /proc/meminfogo deeper
Know that allocating memory on Linux reserves address space and that real pages arrive when the memory is first written, so an allocation can succeed even when the machine cannot back it.
Explain the three overcommit modes and where failure surfaces in each: heuristic and always-overcommit push it to fault time and the out-of-memory killer, strict accounting returns an error at allocation time.
Be ready to justify a setting for a real workload — what strict accounting costs a fork-heavy or arena-heavy service, and why bounding the actual footprint usually beats changing the sysctl.
Own the fleet-wide stance: whether hosts run heuristic overcommit with headroom and clear victim policy, or strict accounting with capacity provisioned for reservations, and what each choice implies for how services must handle allocation failure.
## Allocation is a promise, not a delivery When a program allocates memory, the C library either extends the heap with `brk` or, for larger requests, asks for a fresh mapping with `mmap`. Either way the kernel records a *virtual memory area*: a range of the process's address space and a note about what backs it. No physical page is allocated. The first time the process reads or writes an address in that range, the CPU raises a page fault; the kernel then finds a free physical page, zeroes it (for anonymous memory), wires it into the page tables and resumes the instruction. This is **demand paging**, and it is why a successful allocation guarantees only that the address range is yours. The design pays off because real programs are sparse. A hash table sized for the worst case, a runtime that reserves its maximum heap up front, a thread pool with 8 MB stack reservations per thread, a large sparse array — all reserve heavily and touch lightly. Above all, `fork` duplicates the parent's entire address space as copy-on-write: if the kernel insisted on backing every promise, a 10 GB process could not fork to run a one-line helper. ## The overcommit policies `vm.overcommit_memory` takes three values: - **0 — heuristic (the default).** Each request is checked on its own against a rough estimate of available memory (free pages plus reclaimable cache plus swap). Obviously absurd single requests are refused; the aggregate is not tracked, so many modest requests can still overcommit the machine badly. - **1 — always overcommit.** Every request succeeds regardless of memory. Intended for workloads with legitimately huge sparse mappings, such as certain scientific and database workloads. - **2 — strict accounting.** The kernel maintains a running total of committed address space (`Committed_AS` in `/proc/meminfo`) and refuses any request that would push it past `CommitLimit`, computed as total swap plus `vm.overcommit_ratio` percent of RAM (default 50, or an absolute value via `vm.overcommit_kbytes`). ```bash cat /proc/sys/vm/overcommit_memory grep -E '^(CommitLimit|Committed_AS):' /proc/meminfo ``` Only mode 2 makes allocation failure honest. In modes 0 and 1, the shortage is discovered at fault time, deep inside a `memcpy` or a store instruction, where there is no error path to return: the kernel cannot fail a page fault for anonymous memory, so it reclaims what it can and, failing that, kills a process. ## Why this design is contentious but rational Critics point out that a program which carefully checks the return value of every allocation still dies without warning, and that the process killed may not be the one that overcommitted. Defenders point out that strict accounting is provisioned against *reservations*, not use, so mode 2 typically wastes a large fraction of RAM or breaks fork-heavy and reservation-heavy software outright — you must size `CommitLimit` for the sum of all promises. A few refinements matter in practice: - `mmap` with **`MAP_NORESERVE`** asks that a mapping not be charged against the commit limit even under strict accounting; `MAP_POPULATE` does the opposite, faulting the pages in immediately so failure surfaces at mapping time. - Shared file-backed mappings are not charged the same way as private anonymous ones; the file provides the backing store. - Some databases and language runtimes document a required overcommit setting, precisely because they reserve huge sparse regions and would otherwise fail under mode 2. ## Diagnosing the symptom The telltale sign is a process that dies with no exception, no stack trace and no message of its own, during a write rather than an allocation. `Committed_AS` climbing towards `CommitLimit` shows how much the machine has promised; the kernel log carries the record of what was killed. The real fix is almost never flipping the sysctl — it is bounding the workload's actual footprint, giving the machine reclaimable headroom, and making the important service a less attractive victim. Mode 2 is worth considering only when you can afford to provision for reservations and genuinely need allocation-time failure, for example so an application can shed load gracefully on `ENOMEM`.
- Why does overcommit make `fork` on a large process cheap?`fork` gives the child a copy of the parent's address space marked copy-on-write: the mappings are duplicated but the physical pages are shared until one side writes. Without overcommit, the kernel would have to guarantee backing for the whole duplicate up front, so a 10 GB process could not fork at all under strict accounting unless the machine had 20 GB of commit headroom.
- What does an application actually gain by running under strict overcommit accounting?Allocation failure becomes an honest `ENOMEM` return the program can handle — shed load, refuse a request, free caches — instead of a fatal signal arriving during a write. The price is provisioning: `CommitLimit` is measured against reservations, so fork-heavy and arena-heavy software may be refused memory while plenty of RAM sits unused.
- Does mode 1, always overcommit, remove the risk of running out of memory?No — it removes the *check*, not the shortage. Every request succeeds, so the machine reaches physical exhaustion at fault time and the out-of-memory killer runs. It is appropriate only for workloads with legitimately huge sparse mappings, and it should be paired with real limits elsewhere so one workload cannot take the host down.
saying these in an interview costs you the question
- Thinks a successful allocation guarantees physical memory
- Calls overcommit a bug rather than a design choice
- Believes mode 1 prevents out-of-memory conditions
- Confuses committed address space with resident memory
- Recommends strict mode without sizing CommitLimit