When a Linux host exhausts memory, how does the kernel's out-of-memory killer choose which process to kill, and what does writing to /proc/<pid>/oom_score_adj change about that choice?
answer
- last resort after reclaim fails
- score is roughly what killing recovers
- resident plus swap, relative to total
- an adjustment in tenths of a percent... in percent
- -1000 exempts entirely
basics
~20 sThe kernel scores every process by how much memory freeing it would recover — mainly resident plus swap usage, relative to total memory — and kills the highest scorer. oom_score_adj, from -1000 to 1000, biases that score, with -1000 exempting a process entirely.
solid answer
~50 sThe out-of-memory killer runs only after reclaim has failed to produce a free page. It then walks eligible tasks and computes a **badness** score for each, dominated by how much memory the kernel would get back by killing it: resident pages plus swap usage plus page-table overhead, expressed relative to the memory available in the failing context. The highest scorer is killed with an uncatchable `SIGKILL`, and the whole thread group dies with it. The intent is pragmatic — kill one thing, recover a lot, and prefer the process most responsible for the shortage. Each process carries `/proc/<pid>/oom_score_adj`, an integer from -1000 to 1000 that is added to the score as a percentage of total memory; the special value -1000 makes a process ineligible, and the resulting score is visible in `/proc/<pid>/oom_score`. That is how you tell the kernel to prefer a batch worker over the database, and it is why the biggest process is usually — but not necessarily — the victim.
code
bash · 3 linesfor p in /proc/[0-9]*; do
printf '%s %s %s\n' "${p#/proc/}" "$(cat $p/oom_score 2>/dev/null)" "$(cat $p/comm 2>/dev/null)"
done | sort -k2 -n -r | headgo deeper
Know that when Linux runs out of memory the kernel kills a process outright with an uncatchable signal, that it generally picks a large one, and that the reason appears in the kernel log rather than the application's.
Explain the mechanism: reclaim runs first, the kill happens only when it fails, and the victim is scored by how much memory killing it would recover, adjustable per process through oom_score_adj.
Show that you can run the postmortem and set policy: read the kernel out-of-memory report, distinguish a host-wide kill from a scoped one, and decide which processes should be preferred victims instead of blanket-exempting the important one.
Own the failure-mode design: whether nodes should degrade, kill a chosen sacrificial workload, or panic and fail over, and how that stance combines with headroom, limits and admission control so a memory shortage never becomes a random outage.
## When the killer runs at all The out-of-memory killer is the last step of a failing allocation, not a background policy. The sequence is: an allocation cannot be satisfied from the free lists, so the kernel enters reclaim — dropping clean page cache, writing back dirty pages, swapping anonymous pages if swap exists, shrinking reclaimable slab. Only when reclaim cannot free a page does the kernel declare an out-of-memory condition and select a victim. Two consequences follow, and candidates often miss both: a machine can be under severe pressure for a long time without any kill (it thrashes instead), and a kill is evidence that reclaim was *exhausted*, not merely that memory looked tight. ## The badness heuristic For every eligible task the kernel computes a score whose dominant term is how much memory the kill would recover: the process's resident pages, plus its swap usage, plus the memory consumed by its page tables. That total is normalised against the memory available in the failing context, so the score is essentially "what fraction of the memory in play does this process hold". Key properties: - **Bigger is likelier.** The heuristic deliberately targets the process holding the most memory, because killing something small rarely resolves the shortage and would mean killing again in a moment. - **It is not "whoever asked last".** The allocation that failed may belong to an innocent bystander; the victim is chosen on holdings, not on blame. (The `vm.oom_kill_allocating_task` sysctl switches to the cruder "kill the task that triggered it" policy, which avoids the expensive scan.) - **The unit is the thread group.** The whole process dies; the kernel does not kill one thread and leave the rest. - **Kernel threads and the init process are not eligible.** The computed value is readable per process at `/proc/<pid>/oom_score`. ## Steering the choice with oom_score_adj `/proc/<pid>/oom_score_adj` holds an integer in `[-1000, 1000]`, added to the badness as a percentage of total memory: - **+1000** — effectively volunteer this process first, whatever its size. - **0** — the default, pure size-based scoring. - **negative values** — make the process progressively harder to pick. - **-1000** — exempt: the process is never selected. ```bash pid=$(pgrep -o sshd) cat /proc/$pid/oom_score /proc/$pid/oom_score_adj # An unprivileged process may only raise its own value (volunteer more); # lowering it below the previous setting requires CAP_SYS_RESOURCE. echo 1000 > /proc/self/oom_score_adj ``` The value is inherited across `fork` and preserved across `exec`, which is how a supervisor can set the policy for everything it starts. Use it as a *policy statement about relative importance*: mark the batch importer and the log shipper as preferred victims rather than trying to shield the database with a deep negative score. Exempting a large process with -1000 does not remove the shortage — it guarantees the kernel kills something else, and if the only remaining candidates are small the system can be left in a worse state, killing repeatedly and still failing. ## Reading the aftermath A killed process leaves nothing in its own logs — it received `SIGKILL` and had no chance to write anything. The record is in the kernel ring buffer: an `Out of memory: Killed process <pid> (<comm>)` line with the victim's virtual size and its anonymous and file-backed resident sizes, preceded by a dump of the memory state and, on modern kernels, an `oom-kill:` summary line naming the constraint and the scoring inputs. That dump is the single most useful artefact for the postmortem, because it shows what every task held at the moment of death, not just who died. Two related sysctls belong in the same answer: `vm.panic_on_oom` makes the kernel panic instead of killing (chosen when a partial system is worse than a reboot, for example in a clustered failover design), and `vm.oom_kill_allocating_task` trades victim quality for speed. Finally, note the scope. A kill can also be triggered inside a memory-limited control group rather than for the whole machine; the scoring is then confined to the tasks in that group, so a host with plenty of free RAM can still record an out-of-memory kill. The mechanism above is the global case.
- Why does the killed process leave no error in its own log?Because it is terminated with `SIGKILL`, which cannot be caught, blocked or ignored. There is no handler, no unwinding and no chance to flush a log line. The evidence lives in the kernel ring buffer instead, as an out-of-memory report naming the victim and dumping the memory state of every task at that moment — which is why you go to the kernel log, not the application log.
- Is setting oom_score_adj to -1000 on your critical service a good idea?Rarely. It guarantees the kernel will kill something else, which on a dedicated host may mean killing whatever is left until the machine is unusable, while the real shortage is untouched. It is better to mark clearly expendable processes as preferred victims, bound the critical service's own footprint, and treat any kill as a capacity or leak problem rather than a targeting problem.
- Can an out-of-memory kill happen while the host has gigabytes of free memory?Yes, when the failing allocation is scoped rather than global — most commonly a memory-limited control group, where the accounting and the victim selection are confined to the tasks inside that group. The kernel log entry names the constraint, which is how you tell a group-scoped kill from a host-wide one. Certain constrained allocations, such as those restricted to a specific NUMA node or memory zone, behave similarly.
- What does vm.panic_on_oom change, and when would you set it?It makes the kernel panic instead of selecting a victim. You choose it when a half-functional node is worse than a dead one — clustered systems where another node takes over cleanly on failure, or appliances where a watchdog reboot is the defined recovery path. On a general-purpose host it converts a survivable event into an outage, so it is a deliberate availability trade, not a default.
saying these in an interview costs you the question
- Says the killer picks the process that requested memory
- Believes the target can catch the signal and clean up
- Thinks the kill happens as soon as memory looks tight
- Shields every important process with a negative score
- Expects the crash reason in the application's own log