skip to content

A production Linux host is badly overloaded and you are on a laggy SSH session. A colleague suggests you run htop and read the numbers out to the incident channel. Why is an interactive process viewer a poor instrument in that situation, and what would you reach for instead?

level: seniorimportance: nice to knowfreq 33%

answer

  1. nothing you can paste into a ticket
  2. no batch output, nothing to pipe
  3. every refresh rescans all of /proc
  4. starts at the moment you log in
  5. snapshot first, explore afterwards

basics

~20 s

htop has no batch mode, keeps no history, and rescans /proc on every refresh, so it produces nothing you can paste into a ticket and adds work to a host that has none to spare. Capture a one-shot snapshot instead, then look interactively.

solid answer

~50 s

Because htop is only an interface. It has no batch or non-interactive mode, so there is nothing to pipe, timestamp or paste into the incident channel; it keeps no history, so the minutes before you logged in are gone; and it repaints the whole screen every refresh, which over a laggy link means reading a picture that is already stale. It is not free either — every refresh walks `/proc` and reads several files per task, which on a host with tens of thousands of tasks is real work on a machine that has none to spare. And F9 sits two keystrokes away on production. I would take a one-shot snapshot I can attach to the ticket, then a sampled stream for a timeline, and open htop afterwards only to poke around — with `--readonly` and a longer `-d` delay.

code

bash · 2 lines
bash
# observation only, refresh every 5 seconds (-d is in tenths), one user
htop --readonly -d 50 -u postgres

go deeper

for a junior

Know that htop only draws to a terminal — you cannot redirect or pipe its output — and that it shows the present moment with no record of what came before. That alone rules it out as a way to report numbers to other people.

for a middle

Be ready to explain the cost side: each refresh walks /proc and reads files for every task, and the whole screen repaints every interval. Mention -d for a longer delay and --readonly as the observation-only mode in htop 3.x.

for a senior

Show incident discipline — capture a snapshot you can attach before you start exploring, run a sampled stream for the timeline, and open the interactive viewer last. Say out loud why the artefact matters more than the fastest look.

for a principal

Treat this as a tooling and readiness question rather than a keystroke one. Decide what is collected before an incident so nobody is reading bars aloud, and set the norm that destructive interfaces are not the default way engineers touch production.

## The instrument has to match the job htop is genuinely good at one thing: letting a human orient quickly on a machine they are sitting in front of. Everything else people ask of it during an incident, it does badly — not because it is poorly written, but because a full-screen interactive terminal program is structurally the wrong shape for capture, for automation, and for a host that is already in trouble. ## No batch mode, so no artefact htop is TUI-only. There is no non-interactive one-shot output mode, nothing to redirect into a file, nothing to pipe into a filter, nothing to timestamp. Whatever you learn from the screen leaves the incident as a sentence you typed by hand into a chat window, and by the time anyone reads it the numbers have changed. An incident needs artefacts: a captured snapshot with a timestamp, attached to the ticket, that the next person can read without trusting your transcription. This is where a batch-capable snapshot tool wins outright — anything you can run once, redirect to a file, and attach. Take the snapshot first, before you start exploring; the state you most want is the state that existed when you arrived, and interactive poking is exactly what destroys your chance to record it. ## No history A live viewer starts at the moment you run it. The interesting question in almost every incident is what changed, and change is a comparison against a time you were not logged in. htop cannot scroll back and has nothing to diff against. For "was this happening at 03:10?" you need collected data — a historical collector on the host, or a metrics pipeline off it — and no amount of staring at a live screen substitutes. ## The observer is not free Each refresh, htop enumerates `/proc` and reads several files for every task on the system, then re-renders the whole screen. On a normal host this is negligible. On a host with tens of thousands of tasks, or one whose CPUs are already fully committed, it is measurable work — and it is competing for exactly the resource you are trying to free. It is not unusual for an observer tool to appear near the top of its own list on a saturated box. Optional per-process columns that require reading a task's memory-map details make each pass markedly more expensive again; leave them off unless you specifically need them. The default refresh interval is around a second and a half. `-d` sets it in tenths of a second, so raising it is the cheapest single mitigation if you do want to sit and watch: ``` htop --readonly -d 50 -u postgres ``` ## The laggy link makes it worse A full-screen repaint every interval over a slow or lossy SSH session is a bad trade in both directions: the terminal spends the link's capacity redrawing, and what you see is whatever managed to arrive, which may be seconds behind the machine. A single command that prints a page of text and exits crosses a bad link far better than a program that repaints forever. ## It is a live-fire interface F9 sends signals, and F7/F8 change scheduling priority. On a production host, in the dark, in a hurry, with a cursor whose position you did not verify, that is an unnecessary hazard. htop 3.x accepts `--readonly`, which disables the system- and process-changing features outright; it is a sensible default for anyone using htop as an observation tool rather than a control panel. ## What to run instead A workable order for an overloaded box: 1. **Capture a snapshot** you can attach — a single non-interactive process listing, redirected to a file with a timestamp. 2. **Start a sampled stream** in another window for a short timeline: the system-statistics tools that print one line per interval give you a sequence you can keep, which is exactly what the live screen cannot. 3. **Then, if you still want to explore**, open htop read-only with a longer delay, narrowed with `-u` or `-p`, and use it for what it is good at: finding the specific row that explains the numbers you already captured. ## The general rule Interactive viewers are for orientation; recorded output is for evidence; collected time series are for history. Reaching for the interactive tool first is a habit that costs you the artefact, and the artefact is the thing the post-incident review will actually ask for.

  • You run htop inside a container. What do its header meters actually describe?
    The host, not the container. The process list is limited by the PID namespace, so you see only the container's tasks, but the meters come from /proc/stat and /proc/meminfo, which are the host's unless something is faking them. CPU and memory headroom must be judged from the cgroup's own limits and usage, not from htop's bars.
  • If you do want htop on a production host, how would you make it safer and cheaper?
    Run it with `--readonly`, which disables the killing and renicing actions in htop 3.x, and raise `-d` — it is in tenths of a second, so `-d 50` refreshes every five seconds instead of roughly every one and a half. Narrow the view with `-u` or `-p`, and keep destructive work in explicit commands you can read back before pressing Enter.
  • Someone asks whether a CPU spike occurred at 03:10 last night. Why can no live viewer answer that?
    Because a live viewer has no memory: it begins sampling when you start it and keeps nothing. Answering questions about the past needs data that was already being collected — a historical collector running on the host, or a metrics pipeline off it. That is a decision made before the incident, not during it.

saying these in an interview costs you the question

  • Says htop has a batch flag that prints one snapshot
  • Treats a live reading as evidence for the incident timeline
  • Assumes an observer tool is free on a saturated host
  • Believes htop keeps history you can scroll back through
  • Leaves htop open on production with the kill key armed

context