skip to content

Process & performance monitoring

The shell-level tools for seeing what a machine is doing: process states, CPU, memory, I/O and load average, and where the pressure actually is. Interviewers ask because the server is slow is the most common prompt in the format, and a good answer is a deliberate sequence of these commands.

on this pageshow

questions

19

A Linux service vanished overnight with nothing useful in its own log. How do you confirm that the kernel's out-of-memory killer took it, and what does the kernel's log line tell you about the process it killed?

level: juniorimportance: must knowfreq 68%

answer

  1. the app never got to log it
  2. the kernel keeps its own log
  3. the ring buffer can wrap
  4. one line names the victim
  5. anon-rss is the honest figure

basics

~20 s

Check the kernel log with dmesg -T or journalctl -k, searching for "Out of memory" or "oom-kill". The kill line names the PID and command and reports total-vm, anon-rss, file-rss, shmem-rss and oom_score_adj — anon-rss being what the process actually held in RAM.

solid answer

~50 s

An out-of-memory kill leaves no trace in the application's own log, because the process is terminated without warning and never gets to write anything. The evidence is in the kernel ring buffer: `dmesg -T | grep -i -E 'out of memory|oom-kill'`, or `journalctl -k --since yesterday` when the ring buffer may have wrapped or the box has since rebooted. You are looking for three things. First, a line saying which process *invoked* the OOM killer, which is the one whose allocation failed and is frequently not the victim. Second, a dump of every task with its memory footprint at that moment. Third, the verdict: `Out of memory: Killed process 4142 (java) total-vm:...kB, anon-rss:...kB, file-rss:...kB, shmem-rss:...kB, UID:... pgtables:...kB oom_score_adj:...`. The `anon-rss` figure is the one to quote — it is the private, unreclaimable memory that process was holding when it died.

code

bash · 3 lines
bash
dmesg -T | grep -i -E 'out of memory|oom-kill|killed process'
journalctl -k --since '3 days ago' --grep 'Out of memory'
sysctl kernel.dmesg_restrict

go deeper

for a junior

Make dmesg -T or journalctl -k your first move whenever a process disappears without logging anything, and know that the kernel prints an explicit "Out of memory: Killed process" line naming the PID and command.

for a middle

Be able to read the whole block: the invoking process is not the victim, the task dump shows where memory actually was, and anon-rss rather than total-vm is the meaningful footprint.

for a senior

Show you know the evidence can vanish — a wrapped ring buffer, a reboot, dmesg_restrict — and that the constraint= field separates a host-wide exhaustion from a scoped limit, which are different incidents with different fixes.

for a principal

Own the standard: kernel logs shipped and retained so this forensic trail outlives the box, and a runbook that stops teams from restarting the victim and declaring the incident understood.

## Why the application log is empty An OOM kill is delivered as an unstoppable termination. The process gets no chance to flush a buffer, run a shutdown hook, or log a farewell. So "the log just stops" is the *expected* appearance of an OOM kill, not evidence against it, and the investigation has to move to the kernel's own log. ## Where to look ```bash dmesg -T | grep -i -E 'out of memory|oom-kill|killed process' journalctl -k --since '2 days ago' | grep -i -E 'out of memory|oom' ``` Two practical notes. `dmesg` reads a fixed-size ring buffer, so a chatty kernel can overwrite the evidence within hours, and a reboot clears it entirely — `journalctl -k` (equivalently `journalctl --dmesg`) reads the same messages out of the journal, which persists across boots if `Storage=persistent` is configured, and accepts `--since`/`--until` for a specific window. Also, on hosts with `kernel.dmesg_restrict=1`, an unprivileged user gets an empty `dmesg`; you need root or `CAP_SYSLOG`. An empty output is therefore not proof of innocence until you have checked which of these you are hitting. If the service runs under systemd, `systemctl status` and `journalctl -u <unit>` corroborate it from the supervisor's side: the unit will report the main process exiting killed with signal 9, and the restart that followed. ## Reading the report A global OOM event writes a block with three distinct parts. **The invocation line** names the allocation that could not be satisfied: ``` nginx invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0, oom_score_adj=0 ``` The crucial reading skill: *this process is usually not the one that gets killed*. It is merely the unlucky one that asked for a page when there were none left. Blaming it is the single most common misreading of an OOM report. **The task dump** is a table of every process with columns including `pid`, `total_vm`, `rss`, `pgtables_bytes`, `swapents`, `oom_score_adj` and `name`. Sorting this by `rss` tells you where the memory actually was at the moment of the kill — which is the real forensic value, because it survives the process that consumed it. **The verdict line**: ``` Out of memory: Killed process 4142 (java) total-vm:9437184kB, anon-rss:7340032kB, file-rss:2048kB, shmem-rss:0kB, UID:1000 pgtables:15360kB oom_score_adj:0 ``` Field by field: - `total-vm` — address space mapped. Often huge and largely meaningless on its own; a runtime can reserve far more than it uses. - `anon-rss` — resident *private* memory. This is the number that matters: memory nothing could reclaim, which is why the kernel had no alternative. - `file-rss` — resident file-backed pages, which were reclaimable and so are not the cause. - `shmem-rss` — resident shared-memory pages. - `oom_score_adj` — the bias in effect for this process at the time. Alongside it, modern kernels emit a structured summary line: ``` oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0, global_oom,task_memcg=/system.slice/app.service,task=java,pid=4142,uid=1000 ``` `constraint=CONSTRAINT_NONE` with `global_oom` means the whole machine ran out. `constraint=CONSTRAINT_MEMCG` means a control group hit *its* limit while the host still had memory free — a completely different incident with a completely different remedy, and the field that distinguishes them is right there in the line. ## What you do with the finding The report tells you the victim, its private footprint, and whether the shortage was host-wide or scoped. Combine it with the task dump to see whether one process ballooned or whether the machine was collectively over-committed, and check whether the timestamp lines up with a batch job, a deploy, or a traffic peak. `dmesg -T` renders human timestamps, though they can drift after a suspend; the journal's own timestamps are more trustworthy for correlation. ## The habit worth building When any process disappears unexplained, `journalctl -k` is a reflex, not a last resort. The two-minute check either produces a definitive answer or rules out the most common cause of silent death on a Linux server.

  • `dmesg` shows nothing, but you still suspect an OOM kill from last week. Where else do you look?
    The journal. `journalctl -k --since '8 days ago'` reads the same kernel messages from persistent storage, which survives both ring-buffer wrap and reboots when the journal is configured with `Storage=persistent`. Also confirm you are not being blocked by `kernel.dmesg_restrict=1`, which returns an empty `dmesg` to unprivileged users; and check `journalctl -u <unit>` for the supervisor's own record of the process dying on signal 9.
  • The report says nginx invoked the oom-killer but the killed process was a Java service. What does that tell you?
    Only that nginx happened to request a page when none was available. The invoking process is the one whose allocation failed, not the one held responsible — the kernel then chooses a victim separately. Read the task dump in the same block, sorted by rss, to see where the memory actually was; nginx is usually an innocent bystander in that report.
  • How would you tell a whole-machine OOM from a limit being hit by one scoped workload?
    The `oom-kill:` summary line carries `constraint=`. `CONSTRAINT_NONE` together with `global_oom` means the host itself was exhausted. `CONSTRAINT_MEMCG`, with a `task_memcg=` path pointing at a specific control group, means that group reached its own limit while the machine still had free memory — and `free -h` at the time would have looked perfectly healthy.

saying these in an interview costs you the question

  • Concludes no OOM because the app log is silent
  • Blames the process that invoked the oom-killer
  • Quotes total-vm as the memory the process used
  • Assumes dmesg output survives a reboot
  • Misses that a scoped limit, not the host, was exhausted

context

open as a page

You have a shell on an unfamiliar Linux server and need a full process listing. What is the difference between `ps aux` and `ps -ef`, and when would you reach for `ps -eo` instead of either?

level: juniorimportance: must knowfreq 74%

basics

~20 s

ps aux and ps -ef are BSD-style and UNIX-style invocations of the same tool, differing in default columns: aux prints %CPU, %MEM, VSZ and RSS; -ef prints PPID and start time. ps -eo lets you choose columns and sort order.

open as a page

A Linux application server feels slow and `vmstat 1` shows the `wa` column sitting around 40%. What does iowait actually measure, and why is a high value on its own not proof that the disk is the problem?

level: middleimportance: must knowfreq 62%

basics

~20 s

Iowait is idle CPU time that happened while at least one I/O request was outstanding. It is a subset of idle, so a high value means the CPUs had nothing else to run — it locates spare capacity, not a slow disk. Confirm with iostat -x latency before blaming storage.

open as a page

Overnight, a database host's disk latency doubled. In `iostat -x` output, which fields tell you whether the device itself got slower or the queue in front of it got deeper, and why can `%util` at 100% be meaningless on an SSD or a RAID array?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Read the await columns for per-request latency, aqu-sz for queue depth, and r/s/w/s for offered load. Latency up with load and queue flat means the device slowed; latency up with a deeper queue means you sent more work. %util only counts time with at least one request in flight, so on parallel devices 100% is not saturation.

open as a page

You run `vmstat 1` on a Linux server and the first line shows an almost idle machine, while every line after it shows 90% system time. Why does the first line disagree with the rest, and which lines should you actually read?

level: juniorimportance: should knowfreq 45%

basics

~10 s

The first report from vmstat (and from iostat, mpstat and pidstat) is an average since system boot, not a sample of the last second. Ignore it and read the interval lines that follow.

open as a page

You are triaging a busy Linux server that has both top and htop installed. What does htop's interactive display give you that a flat, non-interactive process list does not, and what do its per-core CPU meters and their coloured segments actually show?

level: juniorimportance: should knowfreq 62%

basics

~20 s

htop adds an interactive view a static list cannot: one meter per CPU core, a process tree, full command lines, and search, filter and signal keys. Each core bar splits its usage by colour — green user time, red kernel time, blue niced work.

open as a page

In `vmstat 1` output on Linux, what do the `r` and `b` columns under `procs` count, and how do you use them together with the machine's CPU count to decide whether a box is CPU-bound or I/O-bound?

level: middleimportance: should knowfreq 55%

basics

~20 s

r counts tasks that are runnable — running or waiting for a CPU — at the moment the line is printed; b counts tasks in uninterruptible sleep, in practice waiting on I/O. Sustained r well above the CPU count means CPU saturation; a sustained non-zero b points at storage.

open as a page

Inside htop you have found a misbehaving process and want to signal it without leaving the interface. Explain what F9 does, how tagging rows with Space changes the target set, and what can go wrong when htop's thread display is switched on.

level: middleimportance: should knowfreq 45%

basics

~20 s

F9 (or k) opens htop's signal menu with SIGTERM preselected and sends your choice to the highlighted row — or to every row tagged with Space. With thread display enabled, signalling a thread row still takes down the whole process.

open as a page

On a Linux host, the `available` column of `free -h` comes from MemAvailable in /proc/meminfo, and it is always smaller than MemFree plus Cached. What does the kernel refuse to count as available, and why?

level: middleimportance: should knowfreq 50%

basics

~20 s

MemAvailable estimates how much memory a new workload could get without swapping. The kernel holds back its free-page watermark and part of the page cache, and excludes cache it cannot simply drop: tmpfs/shmem pages, mlocked pages, and dirty pages awaiting writeback.

open as a page

A Linux host is reported to be pegged at high CPU, but `ps -eo pid,pcpu,args --sort=-pcpu | head` and `top -b -n1` both show no process above a few percent. Why do these two tools disagree with what the machine is doing right now?

level: middleimportance: should knowfreq 45%

basics

~20 s

Neither reading covers the last few seconds. The %CPU that ps prints is CPU time divided by the process's entire lifetime, and top's first batch iteration has no earlier sample to subtract from, so both show lifetime averages. Take a second top sample.

open as a page

On a Linux host you run `pgrep postgres-exporter` and it prints nothing, even though `ps` clearly shows the process running. What is `pgrep` matching against, and what changes when you add `-f`?

level: middleimportance: should knowfreq 52%

basics

~20 s

By default pgrep matches the kernel's short process name, which is capped at 15 characters, so a longer name such as postgres-exporter can never match. The -f flag matches the full command line instead, which finds it — and matches much more besides.

open as a page

An incident happened at 03:00 on a Linux server and nobody was logged in at the time. If sysstat is installed, how do you go back and look at CPU and disk behaviour for that window, and what determines whether the data exists at all?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Read the sysstat archive with sar -f /var/log/sa/saDD -s 03:00 -e 04:00, adding -u for CPU, -d -p for devices and -q for run-queue history. Whether it exists depends on collection being enabled, the collection interval, and the retention setting — Debian and Ubuntu ship it disabled.

open as a page

A Linux server shows 3 GB of swap used in `free -h`, but users report no problem and the box feels responsive. Which measurements tell you whether swap is actually hurting this machine, and where do you read them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Swap used is a stock, not a rate: it can be cold residue paged out weeks ago. Judge swap by paging traffic instead — the pswpin/pswpout counters in /proc/vmstat over an interval — and by memory stall time in /proc/pressure/memory.

open as a page

A long-running Linux daemon starts logging "Too many open files" after a few days of uptime. Using `ps`, `lsof` and `/proc`, how do you confirm it is leaking file descriptors and find what it is holding open?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Find the PID, count entries in /proc/<pid>/fd, and compare that against the limit in /proc/<pid>/limits. Sample the count repeatedly: a steadily rising number is a leak. Then group lsof -nP -p <pid> output by type to see what is accumulating.

open as a page

You are setting the memory policy for a fleet of Linux application servers and a colleague proposes running them all with swap disabled. What does the vm.swappiness setting actually control, and how would you decide?

level: principalimportance: should knowfreq 38%

basics

~20 s

vm.swappiness is a relative weighting between reclaiming anonymous pages and file-backed pages, not an on/off switch or a percentage of RAM. Disabling swap does not remove memory pressure; it leaves the kernel only file pages to evict, so thrash reappears as executable-text refaults.

open as a page

A 32-CPU Linux host reports roughly 4% CPU utilisation in aggregate, yet one service on it is visibly slow. What can `mpstat -P ALL 1` show you that the aggregate figure hides, and what do its `%soft` and `%steal` columns mean?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

The aggregate is a mean across all CPUs, so a single core pinned at 100% contributes about 3% on a 32-CPU host. mpstat -P ALL 1 breaks utilisation out per CPU, exposing that imbalance; %soft is time in softirq handlers and %steal is time a vCPU was runnable while the hypervisor ran something else.

open as a page

A production Linux host is badly overloaded and you are on a laggy SSH session. A colleague suggests you run htop and read the numbers out to the incident channel. Why is an interactive process viewer a poor instrument in that situation, and what would you reach for instead?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

htop has no batch mode, keeps no history, and rescans /proc on every refresh, so it produces nothing you can paste into a ticket and adds work to a host that has none to spare. Capture a one-shot snapshot instead, then look interactively.

open as a page

Adding up the RSS that `ps` prints for every process on a Linux box gives a total far larger than the machine's physical memory. Why does that happen, and which per-process figure can you meaningfully add up instead?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

RSS charges every shared page in full to each process mapping it, so shared libraries, executable text, forked copy-on-write pages and shared memory are counted many times. PSS, in /proc/<pid>/smaps_rollup, splits each shared page across its mappers and can be summed.

open as a page

On a Linux box running a pre-forking web server with 40 worker processes, adding up the RSS column from `ps aux` gives about 60 GB on a machine with 16 GB of RAM, and the machine is not swapping. Why can that sum exceed physical memory, and what would you measure instead?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

RSS counts every resident page a process maps, including pages shared with other processes, so forked workers each report the same shared libraries and copy-on-write parent pages. Summing RSS double-counts them. Use PSS from /proc/<pid>/smaps_rollup, which divides each shared page among its users.

open as a page