A Linux service vanished overnight with nothing useful in its own log. How do you confirm that the kernel's out-of-memory killer took it, and what does the kernel's log line tell you about the process it killed?
answer
- the app never got to log it
- the kernel keeps its own log
- the ring buffer can wrap
- one line names the victim
- anon-rss is the honest figure
basics
~20 sCheck the kernel log with dmesg -T or journalctl -k, searching for "Out of memory" or "oom-kill". The kill line names the PID and command and reports total-vm, anon-rss, file-rss, shmem-rss and oom_score_adj — anon-rss being what the process actually held in RAM.
solid answer
~50 sAn out-of-memory kill leaves no trace in the application's own log, because the process is terminated without warning and never gets to write anything. The evidence is in the kernel ring buffer: `dmesg -T | grep -i -E 'out of memory|oom-kill'`, or `journalctl -k --since yesterday` when the ring buffer may have wrapped or the box has since rebooted. You are looking for three things. First, a line saying which process *invoked* the OOM killer, which is the one whose allocation failed and is frequently not the victim. Second, a dump of every task with its memory footprint at that moment. Third, the verdict: `Out of memory: Killed process 4142 (java) total-vm:...kB, anon-rss:...kB, file-rss:...kB, shmem-rss:...kB, UID:... pgtables:...kB oom_score_adj:...`. The `anon-rss` figure is the one to quote — it is the private, unreclaimable memory that process was holding when it died.
code
bash · 3 linesdmesg -T | grep -i -E 'out of memory|oom-kill|killed process'
journalctl -k --since '3 days ago' --grep 'Out of memory'
sysctl kernel.dmesg_restrictgo deeper
Make dmesg -T or journalctl -k your first move whenever a process disappears without logging anything, and know that the kernel prints an explicit "Out of memory: Killed process" line naming the PID and command.
Be able to read the whole block: the invoking process is not the victim, the task dump shows where memory actually was, and anon-rss rather than total-vm is the meaningful footprint.
Show you know the evidence can vanish — a wrapped ring buffer, a reboot, dmesg_restrict — and that the constraint= field separates a host-wide exhaustion from a scoped limit, which are different incidents with different fixes.
Own the standard: kernel logs shipped and retained so this forensic trail outlives the box, and a runbook that stops teams from restarting the victim and declaring the incident understood.
## Why the application log is empty An OOM kill is delivered as an unstoppable termination. The process gets no chance to flush a buffer, run a shutdown hook, or log a farewell. So "the log just stops" is the *expected* appearance of an OOM kill, not evidence against it, and the investigation has to move to the kernel's own log. ## Where to look ```bash dmesg -T | grep -i -E 'out of memory|oom-kill|killed process' journalctl -k --since '2 days ago' | grep -i -E 'out of memory|oom' ``` Two practical notes. `dmesg` reads a fixed-size ring buffer, so a chatty kernel can overwrite the evidence within hours, and a reboot clears it entirely — `journalctl -k` (equivalently `journalctl --dmesg`) reads the same messages out of the journal, which persists across boots if `Storage=persistent` is configured, and accepts `--since`/`--until` for a specific window. Also, on hosts with `kernel.dmesg_restrict=1`, an unprivileged user gets an empty `dmesg`; you need root or `CAP_SYSLOG`. An empty output is therefore not proof of innocence until you have checked which of these you are hitting. If the service runs under systemd, `systemctl status` and `journalctl -u <unit>` corroborate it from the supervisor's side: the unit will report the main process exiting killed with signal 9, and the restart that followed. ## Reading the report A global OOM event writes a block with three distinct parts. **The invocation line** names the allocation that could not be satisfied: ``` nginx invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0, oom_score_adj=0 ``` The crucial reading skill: *this process is usually not the one that gets killed*. It is merely the unlucky one that asked for a page when there were none left. Blaming it is the single most common misreading of an OOM report. **The task dump** is a table of every process with columns including `pid`, `total_vm`, `rss`, `pgtables_bytes`, `swapents`, `oom_score_adj` and `name`. Sorting this by `rss` tells you where the memory actually was at the moment of the kill — which is the real forensic value, because it survives the process that consumed it. **The verdict line**: ``` Out of memory: Killed process 4142 (java) total-vm:9437184kB, anon-rss:7340032kB, file-rss:2048kB, shmem-rss:0kB, UID:1000 pgtables:15360kB oom_score_adj:0 ``` Field by field: - `total-vm` — address space mapped. Often huge and largely meaningless on its own; a runtime can reserve far more than it uses. - `anon-rss` — resident *private* memory. This is the number that matters: memory nothing could reclaim, which is why the kernel had no alternative. - `file-rss` — resident file-backed pages, which were reclaimable and so are not the cause. - `shmem-rss` — resident shared-memory pages. - `oom_score_adj` — the bias in effect for this process at the time. Alongside it, modern kernels emit a structured summary line: ``` oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0, global_oom,task_memcg=/system.slice/app.service,task=java,pid=4142,uid=1000 ``` `constraint=CONSTRAINT_NONE` with `global_oom` means the whole machine ran out. `constraint=CONSTRAINT_MEMCG` means a control group hit *its* limit while the host still had memory free — a completely different incident with a completely different remedy, and the field that distinguishes them is right there in the line. ## What you do with the finding The report tells you the victim, its private footprint, and whether the shortage was host-wide or scoped. Combine it with the task dump to see whether one process ballooned or whether the machine was collectively over-committed, and check whether the timestamp lines up with a batch job, a deploy, or a traffic peak. `dmesg -T` renders human timestamps, though they can drift after a suspend; the journal's own timestamps are more trustworthy for correlation. ## The habit worth building When any process disappears unexplained, `journalctl -k` is a reflex, not a last resort. The two-minute check either produces a definitive answer or rules out the most common cause of silent death on a Linux server.
- `dmesg` shows nothing, but you still suspect an OOM kill from last week. Where else do you look?The journal. `journalctl -k --since '8 days ago'` reads the same kernel messages from persistent storage, which survives both ring-buffer wrap and reboots when the journal is configured with `Storage=persistent`. Also confirm you are not being blocked by `kernel.dmesg_restrict=1`, which returns an empty `dmesg` to unprivileged users; and check `journalctl -u <unit>` for the supervisor's own record of the process dying on signal 9.
- The report says nginx invoked the oom-killer but the killed process was a Java service. What does that tell you?Only that nginx happened to request a page when none was available. The invoking process is the one whose allocation failed, not the one held responsible — the kernel then chooses a victim separately. Read the task dump in the same block, sorted by rss, to see where the memory actually was; nginx is usually an innocent bystander in that report.
- How would you tell a whole-machine OOM from a limit being hit by one scoped workload?The `oom-kill:` summary line carries `constraint=`. `CONSTRAINT_NONE` together with `global_oom` means the host itself was exhausted. `CONSTRAINT_MEMCG`, with a `task_memcg=` path pointing at a specific control group, means that group reached its own limit while the machine still had free memory — and `free -h` at the time would have looked perfectly healthy.
saying these in an interview costs you the question
- Concludes no OOM because the app log is silent
- Blames the process that invoked the oom-killer
- Quotes total-vm as the memory the process used
- Assumes dmesg output survives a reboot
- Misses that a scoped limit, not the host, was exhausted