skip to content

You run `vmstat 1` on a Linux server and the first line shows an almost idle machine, while every line after it shows 90% system time. Why does the first line disagree with the rest, and which lines should you actually read?

level: juniorimportance: should knowfreq 45%

answer

  1. one of these lines is not like the others
  2. the tool needs two readings to make a rate
  3. there was nothing before the first sample
  4. uptime is the only earlier reference point
  5. first report equals since boot

basics

~10 s

The first report from vmstat (and from iostat, mpstat and pidstat) is an average since system boot, not a sample of the last second. Ignore it and read the interval lines that follow.

solid answer

~50 s

These tools print their rate columns as counters divided by elapsed time, and for the very first report the elapsed time is the whole uptime. So on a box that has been up for 40 days, a CPU storm that started five minutes ago is averaged into 40 days of idleness and effectively vanishes. Every subsequent line covers only the interval you asked for, which is what you want. In practice I always run something like `vmstat 1 10` or `iostat -xz 1` and discard the first block — `iostat -y` will suppress it for you. One nuance for `vmstat`: the `procs` and memory columns in that first line are read live from the kernel, so they are current; it is the rate columns (`si`/`so`, `bi`/`bo`, `in`, `cs` and the CPU percentages) that are since-boot averages.

code

bash · 6 lines
bash
# Discard the since-boot first report explicitly
vmstat 1 10 | tail -n +4

# iostat can suppress it for you: -y drops the first report,
# -x adds extended device stats, -z hides idle devices
iostat -xyz 1 10

go deeper

for a junior

Know that the first report covers time since boot and the rest cover your interval, and say plainly that you skip line one. Always pass an interval, as in vmstat 1 or iostat -x 1.

for a middle

Explain the mechanism: these tools compute rates from cumulative counter deltas, and the first report has no earlier reading to subtract, so it divides by uptime. Point out that vmstat's procs and memory fields are live even on that first line.

for a senior

Show the capture habit. When you attach evidence to an incident, you run a bounded interval series such as vmstat 1 30 or iostat -xyz 1, discard or suppress the first block, and note the uptime so a reader can judge how badly a since-boot average would have flattened the signal.

for a principal

Frame it as a data-hygiene rule for whoever reads the output later. Standardise the invocation your runbooks and diagnostic collection scripts use so pasted output is unambiguous, and treat single-report snapshots in postmortems as evidence that has to be re-gathered rather than argued over.

## The behaviour All of the sysstat tools (`iostat`, `mpstat`, `pidstat`, `sar` in live mode) and `vmstat` from procps-ng work the same way: they read cumulative kernel counters, subtract the previous reading, and divide by the elapsed time to produce a rate. That requires a previous reading. For the very first report there isn't one, so the tools fall back to the only earlier reference point that exists — the moment the counters started, i.e. boot. The consequence is that the first line answers a different question from all the others. Later lines answer "what happened in the last second"; the first line answers "what has this machine averaged since it was switched on". ``` $ vmstat 1 3 procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu----- r b swpd free buff cache si so bi bo in cs us sy id wa st 2 0 0 1204312 84332 2419884 0 0 3 17 62 109 1 0 98 0 0 9 0 0 1203044 84332 2419884 0 0 0 0 4812 9931 4 90 6 0 0 11 0 0 1202880 84332 2419884 0 0 0 0 4903 9877 3 91 6 0 0 ``` The first line's `1 0 98 0 0` is 40 days of mostly-idle history. The next two are the incident. ## Why it matters in an incident This is the single most common way a beginner misreads these tools. Two failure modes: 1. **Running the tool with no interval at all.** `iostat -x` on its own prints exactly one report — the since-boot one — and then exits. People screenshot that and conclude the disk is fine. It says nothing about right now. The same is true of a bare `mpstat` or `vmstat` with no delay argument. 2. **Reading only the first line of an interval run.** The tool is doing the right thing, the reader is one line off. The fix is a habit, not knowledge: never run these tools without an interval, and never trust the first block. ## The exceptions inside vmstat's own first line `vmstat` mixes two kinds of column, and only one kind is affected: - **Instantaneous columns** — the `procs` columns `r` and `b`, and the memory columns `swpd`, `free`, `buff`, `cache` — are sampled from the kernel when the line is printed. They are current even on line one. - **Rate columns** — `si`/`so`, `bi`/`bo`, `in` (interrupts/s), `cs` (context switches/s) and the CPU percentage block — are counter deltas, and on line one the delta is taken against boot. So the first line is not useless; it is *mixed*, which is worse, because it looks internally consistent. A first line showing `r 11` next to `id 98` is not a contradiction — the run queue is now, the idle percentage is history. ## Suppressing it `iostat` has a dedicated flag for this: `iostat -y 1` omits the first report entirely. Combine it with `-x` (extended device statistics) and `-z` (skip devices with no activity in the interval) and you get a clean rolling view: ``` $ iostat -xyz 1 ``` Other tools have no equivalent, so you discard the block yourself. When capturing evidence for a ticket, capture several intervals — `vmstat 1 30` — rather than a single snapshot, so a reviewer can see whether the condition is sustained or a one-second blip. ## The same trap in written-down history The since-boot line is also why comparing a first-report number against a monitoring graph never lines up: the graph is a rate over a short window, the first line is a rate over uptime. If you need history rather than a live sample, that is what the sysstat archive under `/var/log/sa` is for, read back with `sar -f`.

  • If the first line is since boot, what does a bare `iostat -x` with no interval argument tell you?
    Almost nothing about the present. It prints one since-boot report and exits, so a device that has been pathological for ten minutes on a box with 40 days of uptime still looks healthy. Always give it an interval — `iostat -x 1` — and read the second report onward, or add `-y` to drop the first one.
  • Are any columns in vmstat's first line still trustworthy?
    Yes. The `procs` columns `r` and `b` and the memory columns are sampled live when the line prints, so they reflect the moment you ran the command. Only the counter-derived rates — `si`/`so`, `bi`/`bo`, `in`, `cs` and the CPU percentages — are averaged over uptime. That mix is what makes the line look plausible and mislead people.

It is a car's lifetime average fuel consumption on the same display as the instantaneous readout. The lifetime figure is real, but it will not tell you that you are flooring it right now.

saying these in an interview costs you the question

  • Reading the first vmstat line as the current second
  • Running `iostat -x` with no interval and calling the disk healthy
  • Assuming every column in the first line is a since-boot average
  • Thinking the tool is broken because line one contradicts line two
  • Screenshotting one report instead of capturing several intervals

context