skip to content

Process inspection: ps and top

Reading ps and top output properly — process states, memory columns, load average — is how you turn "the server is slow" into a named process. It is the first thing most Linux troubleshooting interviews ask you to walk through.

on this pageshow

questions

5

You have a shell on an unfamiliar Linux server and need a full process listing. What is the difference between `ps aux` and `ps -ef`, and when would you reach for `ps -eo` instead of either?

level: juniorimportance: must knowfreq 74%

answer

  1. two option ancestries, one program
  2. undashed options are the BSD grammar
  3. aux gives memory, -ef gives parentage
  4. -o names the columns you want
  5. --sort=-rss beats piping through sort

basics

~20 s

ps aux and ps -ef are BSD-style and UNIX-style invocations of the same tool, differing in default columns: aux prints %CPU, %MEM, VSZ and RSS; -ef prints PPID and start time. ps -eo lets you choose columns and sort order.

solid answer

~40 s

They are the same program with two option ancestries. `ps aux` is the BSD form — options with no leading dash — and gives me the resource view: USER, PID, %CPU, %MEM, VSZ, RSS, STAT, START, TIME and COMMAND. `ps -ef` is the UNIX/POSIX form and gives me the relationship view: UID, PID, PPID, C, STIME, TTY, TIME and CMD, so it is what I use when I care about parentage. Both list every process on the box, so the choice is about columns, not coverage. In practice I reach for `ps -eo` almost immediately, because it lets me name exactly the fields I want and sort on one — for example `ps -eo pid,ppid,user,stat,%cpu,rss,etimes,args --sort=-rss`. That gives a stable, greppable, script-friendly line instead of whatever the default happened to include.

code

bash · 1 line
bash
ps -eo pid,ppid,user,stat,%cpu,rss,etimes,args --sort=-rss | head -n 15

go deeper

for a junior

Know that both forms list every process and that the difference is which columns print: aux carries the memory numbers, -ef carries PPID. Be able to say out loud that BSD options take no dash.

for a middle

Be ready to build a custom listing on the spot with -eo and --sort, and to explain why the COMMAND column is truncated on a terminal but not in a pipeline.

for a senior

Show that you drive ps from the question you are answering: name the columns, sort on one, add etimes or nlwp when the answer depends on age or thread count, and know that the output is a single snapshot.

for a principal

Talk about making process listings reproducible across a fleet — fixed -o field sets in runbooks and diagnostic collectors so output can be diffed between hosts and over time rather than eyeballed.

## Two ancestries, one binary On Linux, `ps` comes from the procps-ng package and deliberately accepts three option styles: BSD options (no leading dash, e.g. `aux`), UNIX/POSIX options (single dash, e.g. `-ef`), and GNU long options (double dash, e.g. `--sort`). This is a compatibility decision, not a functional split. `ps aux` and `ps -ef` both select every process on the system; they differ almost entirely in which columns they print by default. ## What each one actually prints `ps aux` decomposes as `a` (processes belonging to any user, not just yours), `u` (user-oriented output format) and `x` (include processes with no controlling terminal — daemons). Its columns are: ``` USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND ``` That is the resource-oriented view. `RSS` is resident set size in kibibytes, `VSZ` is the size of the virtual address space, `STAT` is the process state letter plus modifier flags, and `TIME` is accumulated CPU time. `ps -ef` decomposes as `-e` (every process) and `-f` (full format). Its columns are: ``` UID PID PPID C STIME TTY TIME CMD ``` That is the relationship-oriented view: it gives you `PPID`, the parent process ID, which `ps aux` does not print by default. `C` is an integer CPU-utilisation figure used for scheduling, and `STIME` is the start time. Notably `-ef` prints no memory columns at all — if someone tells you they read RSS out of `ps -ef`, they have not run it. ## The dash is not decoration `ps -aux` is not a spelling variant of `ps aux`. With a dash, `u` is the UNIX option that takes a username argument, so `-aux` parses as "processes of the user named x". procps-ng notices this common mistake and prints a warning before falling back to BSD behaviour when no user `x` exists, but the rule to remember is that dashed and undashed options are different grammars and should not be mixed casually. ## ps -eo: stop accepting the defaults The form worth internalising is `-o` (or `--format`), which takes a comma-separated list of field names and prints exactly those, in that order: ``` ps -eo pid,ppid,user,stat,%cpu,rss,etimes,args --sort=-rss | head -n 15 ``` Useful fields include `pid`, `ppid`, `user`, `stat`, `%cpu`, `%mem`, `rss`, `vsz`, `etime` (elapsed as `[[dd-]hh:]mm:ss`), `etimes` (elapsed in plain seconds, far easier to compare in a script), `nlwp` (number of threads), `comm` (the short process name) and `args` (the full command line). You can rename a header inline with `pid=PROCESS`, and an empty header (`pid=`) suppresses the header row entirely — handy when the output feeds another command. `--sort` takes the same field names with a leading `-` for descending, so `--sort=-%cpu` puts the hungriest process first and removes the need to pipe through `sort` and count columns. ## Truncation, threads, and trees By default `ps` truncates each line to the terminal width, which silently cuts off long Java or Python command lines exactly where the interesting argument lives. `ps auxww` (or `ps -ef ww`) disables that truncation; a single `w` widens once, a second `w` makes it unlimited. When output is not a terminal, procps-ng does not truncate, which is why the same command can look different in a pipeline than on screen. Two more variants earn their keep: `ps -eLf` (or adding `nlwp`/`lwp` to `-o`) shows threads rather than only processes, which matters for a JVM or any thread-pool service; and `ps auxf` draws an ASCII parentage tree, as does `pstree -p`, when you want to see which supervisor spawned what. ## What ps is and is not `ps` walks `/proc` once and prints a snapshot. It is not a live feed, and it does not average over an interval — the `%CPU` it reports is CPU time divided by the process's whole lifetime, which is why a process that started spiking ten seconds ago can still show a low number. For a moving picture you need a tool that takes two samples. But as the first command on an unfamiliar box, `ps -eo` with the columns you care about and a sort will name the suspect faster than anything else.

  • Which of the two default formats would you pick if you suspected a service was being restarted by the wrong parent, and why?
    `ps -ef`, because PPID is one of its default columns. Reading the parent PID tells you immediately whether the process was started by systemd, by a shell, or by some supervisor wrapper. With `ps aux` you would have to add the column yourself, so in practice I would just run `ps -eo pid,ppid,user,args` or `pstree -p` and read the parentage directly.
  • Why does `ps aux | grep java` sometimes show a match even after the Java process has exited?
    Because the `grep` itself is a process whose command line contains the pattern, and `ps` captured it. The usual fixes are to bracket a character in the pattern (`grep [j]ava`), or to skip the pipeline entirely and use `pgrep -a java`, which matches process names rather than a text dump and never reports itself.
  • What does the `etimes` field give you that `TIME` does not?
    `TIME` is accumulated CPU time — how much processor the process has consumed. `etimes` is elapsed wall-clock time since it started, in seconds. Comparing them is informative: a process with hours of elapsed time and seconds of CPU time is idle or blocked, while one where the two are close has been burning a core continuously.

Same camera, two preset modes: one exposes for resource usage, the other for family relationships. -eo is switching to manual.

saying these in an interview costs you the question

  • Says ps aux and ps -ef list different sets of processes
  • Claims ps -ef prints RSS or %MEM by default
  • Treats ps -aux as identical to ps aux
  • Assumes ps output is a live feed rather than a snapshot
  • Reads a truncated COMMAND column as the full command line

context

open as a page

A Linux host is reported to be pegged at high CPU, but `ps -eo pid,pcpu,args --sort=-pcpu | head` and `top -b -n1` both show no process above a few percent. Why do these two tools disagree with what the machine is doing right now?

level: middleimportance: should knowfreq 45%

basics

~20 s

Neither reading covers the last few seconds. The %CPU that ps prints is CPU time divided by the process's entire lifetime, and top's first batch iteration has no earlier sample to subtract from, so both show lifetime averages. Take a second top sample.

open as a page

On a Linux host you run `pgrep postgres-exporter` and it prints nothing, even though `ps` clearly shows the process running. What is `pgrep` matching against, and what changes when you add `-f`?

level: middleimportance: should knowfreq 52%

basics

~20 s

By default pgrep matches the kernel's short process name, which is capped at 15 characters, so a longer name such as postgres-exporter can never match. The -f flag matches the full command line instead, which finds it — and matches much more besides.

open as a page

A long-running Linux daemon starts logging "Too many open files" after a few days of uptime. Using `ps`, `lsof` and `/proc`, how do you confirm it is leaking file descriptors and find what it is holding open?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Find the PID, count entries in /proc/<pid>/fd, and compare that against the limit in /proc/<pid>/limits. Sample the count repeatedly: a steadily rising number is a leak. Then group lsof -nP -p <pid> output by type to see what is accumulating.

open as a page

On a Linux box running a pre-forking web server with 40 worker processes, adding up the RSS column from `ps aux` gives about 60 GB on a machine with 16 GB of RAM, and the machine is not swapping. Why can that sum exceed physical memory, and what would you measure instead?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

RSS counts every resident page a process maps, including pages shared with other processes, so forked workers each report the same shared libraries and copy-on-write parent pages. Summing RSS double-counts them. Use PSS from /proc/<pid>/smaps_rollup, which divides each shared page among its users.

open as a page