skip to content

In `vmstat 1` output on Linux, what do the `r` and `b` columns under `procs` count, and how do you use them together with the machine's CPU count to decide whether a box is CPU-bound or I/O-bound?

level: middleimportance: should knowfreq 55%

answer

  1. two queues, not one
  2. one waits for a CPU, one waits for a device
  3. the number is meaningless until you divide
  4. compare against nproc
  5. instantaneous sample, so read many lines

basics

~20 s

r counts tasks that are runnable — running or waiting for a CPU — at the moment the line is printed; b counts tasks in uninterruptible sleep, in practice waiting on I/O. Sustained r well above the CPU count means CPU saturation; a sustained non-zero b points at storage.

solid answer

~40 s

Both are instantaneous samples taken when vmstat prints the line, not averages. I compare `r` against `nproc`: if an 8-CPU box shows `r` around 30 across many consecutive lines, roughly 22 tasks are queued for a CPU they cannot get, and the machine is CPU-saturated regardless of what the utilisation percentages say. `b` is the other queue — tasks parked in uninterruptible sleep, which on a normal server means block I/O in flight. A steady `b` of several tasks alongside a high `wa` column sends me to `iostat -x` next. The important discipline is to read a run of lines, not one: a single sample of `r` catches whatever happened to be scheduled in that microsecond and is easy to over-read.

go deeper

for a junior

Recall that r is the run queue and b is the blocked-on-I/O queue, and that you compare r against the number of CPUs from nproc before calling it high.

for a middle

Explain that both columns are instantaneous samples rather than rates, that r includes tasks currently on a CPU, and walk the four combinations of high or low r and b to the tool you would reach for next.

for a senior

Demonstrate the triage discipline: capture a bounded run rather than a snapshot, state the CPU count alongside the numbers, and know when the columns are telling you the host is fine and the problem is in the application or a dependency.

for a principal

Own the judgement about how much a host-level queue signal is worth on your platform. On virtualised or quota-constrained fleets the visible CPU count is not the available CPU time, and you should be clear with your teams about when run-queue depth is a capacity argument and when it is an application-concurrency argument.

## What the two columns are The `procs` block at the left of `vmstat` output has exactly two columns, and they are the cheapest saturation signal on the box. ``` procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu----- r b swpd free buff cache si so bi bo in cs us sy id wa st 34 0 0 1204312 84332 2419884 0 0 0 8 6021 14887 92 7 1 0 0 ``` - **`r`** — the number of tasks in the runnable state: currently executing on a CPU plus those ready to execute and waiting their turn. On modern procps-ng the currently-running tasks are included, so on a fully busy 8-CPU box the floor for `r` is about 8. - **`b`** — the number of tasks in uninterruptible sleep. These are not waiting for a CPU; they are waiting for something the kernel will not let them be interrupted out of, which on a typical server means a block-device read or write, and sometimes a stalled network filesystem. Both are sampled at print time. They are not rates and not averages, which makes them noisy sample-by-sample and meaningful only across a run of lines. ## Reading `r` against the CPU count The number alone means nothing until you divide it by the number of CPUs, which you get from `nproc` or the CPU count in `lscpu`. - `r` consistently below the CPU count: there is CPU headroom. Whatever is slow, it is not "no CPU available". - `r` hovering at roughly the CPU count: fully utilised but not queueing much. - `r` several times the CPU count, sustained: tasks are spending real time waiting for a CPU. Every runnable task is accumulating scheduling delay, and latency at the application level rises even though throughput may look fine. That ratio is the useful form: `r/nproc`. An `r` of 30 is catastrophic on a 4-CPU VM and unremarkable on a 96-core host. ## Reading `b` `b` is usually 0 or flickering at 1 on a healthy machine. A sustained value means a set of tasks are permanently parked waiting for I/O to complete. Three shapes come up in practice: - **`b` high, `wa` high, `r` low.** Classic storage bottleneck. The CPUs have nothing to run because the work is all blocked. Next stop is `iostat -x` to look at the per-device latency and queue depth, and `pidstat -d 1` to see which process is issuing the I/O. - **`b` high, `wa` low, `r` high.** The CPUs are busy with other work, so blocked tasks are not showing up as idle-waiting time. Storage may still be the problem for those particular tasks; `wa` being low does not clear the disk. - **`b` high and pinned, with no disk activity in `bi`/`bo` at all.** Suspect something that is not a local disk: a hung NFS or other network mount, a stuck device. The tasks will be unkillable in that state, which is its own diagnostic signal. ## Using them together as a first triage The two columns split the world in one glance: ``` r >> nproc, b ~ 0 -> CPU saturation; profile what is burning it r low, b > 0 -> I/O bound; go to iostat -x r >> nproc, b > 0 -> both; fix the bigger one first and re-measure r low, b = 0, slow -> not a host resource problem at all ``` That last row matters more than people expect. If the run queue is short and nothing is blocked and the box still feels slow, the bottleneck is somewhere else entirely — a downstream dependency, a lock inside the application, a single-threaded hot path that cannot use the CPUs you have. Host counters have told you what they can, and continuing to stare at them is wasted time. ## The caveats worth stating out loud - **Sample noise.** One line is one instant. Always read at least ten intervals before drawing a conclusion, and prefer a bounded run such as `vmstat 1 30` when capturing evidence. - **The CPU count must be the one the workload actually has.** On a virtual machine or under a CPU quota, the number of CPUs visible to `nproc` may not be the amount of CPU time the workload can consume, and `r` will look worse than the box's own utilisation suggests. - **`r` is a count of tasks, not of processes.** Threads are scheduled independently, so a single multi-threaded process can put dozens of entries into `r` on its own. - **`b` is a count, not a latency.** Two tasks blocked on a device that answers in 200 microseconds is fine; two blocked on a device answering in 200 milliseconds is not. The column tells you where to look, `iostat -x` tells you how bad it is.

  • An 8-CPU host shows `r` around 30 but the CPU columns report roughly 40% idle. How do you reconcile that?
    Read more lines before reconciling anything — `r` is one instant while the CPU percentages are averaged over the whole interval, so a bursty workload can queue deeply for a few milliseconds each second and still leave idle time overall. If the pattern holds across many samples, look for a short, intense periodic burst: a scheduled job, a garbage-collection pause, a batched flush.
  • You see `b` stuck at 4 but `bi` and `bo` are both zero. What does that combination suggest?
    Tasks are blocked in uninterruptible sleep without any local block-device traffic to explain it. The usual cause is a network filesystem whose server has stopped answering, or a device that has stopped completing requests. Check what those tasks have open and which mounts are involved before assuming a disk problem — a local disk that is merely slow would still show non-zero `bi`/`bo`.

saying these in an interview costs you the question

  • Calling any r value high without knowing the CPU count
  • Reading a single vmstat line as proof of saturation
  • Assuming b counts processes waiting for the CPU
  • Treating b as zero-means-storage-is-fine
  • Forgetting that threads, not processes, fill the run queue

context