A Linux application server feels slow and `vmstat 1` shows the `wa` column sitting around 40%. What does iowait actually measure, and why is a high value on its own not proof that the disk is the problem?
answer
- it is carved out of idle, not out of busy
- adding CPU work makes it fall
- a good pointer, a bad verdict
- the sum of the CPU columns is still 100
- measure the device, not the CPU's boredom
basics
~20 sIowait is idle CPU time that happened while at least one I/O request was outstanding. It is a subset of idle, so a high value means the CPUs had nothing else to run — it locates spare capacity, not a slow disk. Confirm with iostat -x latency before blaming storage.
solid answer
~50 sThe `wa` column is not a fifth kind of busy. The kernel accounts a tick as iowait when the CPU was idle **and** there was at least one I/O in flight for it, so iowait time is carved out of idle time, not out of user or system time. Two consequences fall straight out. First, high iowait can simply mean the machine has little else to do — start a CPU-heavy batch job on the same host and `wa` will collapse toward zero while the storage behaves exactly as before. Second, and worse, a genuinely I/O-starved service on a busy box shows *low* iowait, because there is no idle time left to attribute. So I treat `wa` as a hint to go look, never as a verdict: the verdict comes from `iostat -x` device latency and queue depth, and from `pidstat -d 1` to see who is issuing the I/O.
go deeper
Be ready to say that iowait is idle CPU time recorded while an I/O was outstanding, and that you would check iostat -x before blaming a disk.
Explain the accounting: the bucket comes out of idle, so us + sy + id + wa + st still totals 100. Show both failure modes — an idle host inflating it and a busy host hiding a real bottleneck.
Demonstrate the investigation. Route from wa to per-device latency and queue depth, to pidstat -d for attribution, and be explicit that a device answering in microseconds with a shallow queue is healthy regardless of what the CPU bucket says.
Push back on iowait as an alerting or capacity signal. It moves with unrelated CPU load and averages away on wide hosts, so anything that pages a human or justifies spend should be built on device latency and queue depth instead, with iowait kept as context at most.
## The definition, precisely Linux accounts CPU time into buckets — user, nice, system, idle, iowait, irq, softirq, steal, guest — and `vmstat`'s `wa`, `iostat`'s `%iowait`, `mpstat`'s `%iowait` and `sar -u`'s `%iowait` all read the same bucket. The rule for that bucket is: **the CPU was idle, and at least one I/O request issued from that CPU had not yet completed.** The first half of that sentence is the half everyone forgets. Iowait is *idle time with a label on it*. The CPU is doing nothing in either case; the label only records that something was outstanding while it did nothing. Add up `us + sy + id + wa + st` and you still get 100% — `wa` is not extra work, it is a slice taken out of what would otherwise be reported as idle. ## The two consequences that break naive readings **High iowait does not mean slow storage.** It means the CPUs were free while I/O was pending. A single-threaded batch job reading a file on an otherwise empty 16-CPU host can drive `%iowait` high while the device is answering in a healthy 200 microseconds — the CPUs are idle because nothing else is queued, and every idle tick gets the iowait label. Nothing is wrong. **Low iowait does not clear storage.** Run something CPU-hungry on that same host and `%iowait` falls toward zero, because the ticks are now being spent in user time instead of idle time. The disk did not get faster; the accounting simply had nowhere left to record the wait. This is the failure mode that misleads people on busy application servers: the box that is genuinely being killed by disk latency often shows an unremarkable `wa`. ``` # same storage, same latency, wildly different wa $ vmstat 1 # idle host, one reader r b ... us sy id wa st 0 1 ... 1 2 57 40 0 $ vmstat 1 # add a CPU-bound job; wa collapses r b ... us sy id wa st 9 1 ... 88 8 2 2 0 ``` ## The other distortions - **It is per-CPU, then averaged.** The aggregate `%iowait` is the mean across all CPUs. On a 32-CPU host, one CPU spending its entire second in iowait contributes about 3% to the aggregate. `mpstat -P ALL 1` shows the per-CPU breakdown and is where you look when the aggregate number seems too small to explain the symptom. - **Attribution is approximate.** The accounting is tied to the CPU that had the outstanding request. Tasks migrate between CPUs, so the wait may be labelled somewhere other than where you would expect. Treat the figure as a fleet-of-CPUs statistic, not as a per-task measurement. - **It says nothing about which device.** Iowait covers block I/O generally. A stalled network filesystem, a slow root disk and a saturated data volume all land in the same bucket. - **On a virtual machine, check `st` alongside it.** Steal time is a separate bucket, and a host that is being starved by its hypervisor is a different problem with a different fix. ## What to do instead Use `wa` as a routing signal and then get a real measurement: 1. **`iostat -xz 1`** — per-device extended statistics. The latency columns (`r_await`/`w_await`, or a single `await` on sysstat 11 and earlier) tell you how long requests are actually taking, and `aqu-sz` tells you how deep the queue in front of the device is. That is a measurement of storage; `%iowait` is a measurement of CPU idleness. 2. **`pidstat -d 1`** — per-process read and write throughput, which answers "who is doing this" in a way no aggregate counter can. 3. **`vmstat`'s `b` column** — how many tasks are actually parked waiting, which is a count you can compare against the number of workers your service runs. If `iostat` shows single-digit-millisecond service with a shallow queue, the storage is fine no matter what `wa` says, and you should be looking at concurrency, locking or a downstream dependency instead. ## How to say it in an interview The compact formulation that lands: *"Iowait is idle time, not busy time. It tells me the CPUs had nothing to run while I/O was outstanding. It is a good pointer and a terrible verdict — for a verdict I want device latency and queue depth from `iostat -x`."* Follow it with the falsifying observation — that adding CPU load makes `wa` disappear without touching the disk — and you have demonstrated you understand the accounting rather than the folklore.
- A host shows `%iowait` near zero, yet the application's disk reads are visibly slow. Is storage cleared?No. Iowait needs idle time to be recorded in, and a busy host has none, so a real storage bottleneck can show a `%iowait` of zero. Go straight to `iostat -x` and read `r_await`/`w_await` and `aqu-sz` per device, plus `pidstat -d 1` for the issuing process. Device latency is the measurement; iowait is only an idleness label.
- On a 64-CPU host, `%iowait` reads 2% while one service is clearly stalling on I/O. Where does that 2% come from?It is a mean across all 64 CPUs, so a couple of CPUs spending most of their time in iowait barely moves the aggregate. `mpstat -P ALL 1` breaks it out per CPU and will show the imbalance directly. This is why the aggregate figure is nearly useless for diagnosis on wide machines.
- Does high `%iowait` ever justify buying faster storage on its own?Not on its own. It tells you there was spare CPU while I/O was pending, which is also what a perfectly healthy sequential read looks like on an idle box. The purchase case needs device latency that is high relative to what the media should deliver, a queue that is genuinely deep, and evidence that application latency tracks it — all of which come from `iostat -x` over time, not from the CPU bucket.
Iowait is a cashier standing idle while a customer's card terminal thinks. It records that the till was free during the wait; it does not measure how slow the terminal is, and hiring a second cashier who is always busy makes the idle-time number vanish without speeding anything up.
saying these in an interview costs you the question
- Calling iowait a fifth kind of CPU busy time
- Concluding the disk is the bottleneck from wa alone
- Assuming low iowait proves storage is healthy
- Reading the aggregate figure on a many-CPU host
- Ignoring steal time when the host is a virtual machine