Overnight, a database host's disk latency doubled. In `iostat -x` output, which fields tell you whether the device itself got slower or the queue in front of it got deeper, and why can `%util` at 100% be meaningless on an SSD or a RAID array?
answer
- one number contains the other
- latency includes the wait before service
- queue depth separates the two causes
- the busy percentage predates parallel devices
- never idle is not the same as full
basics
~20 sRead the await columns for per-request latency, aqu-sz for queue depth, and r/s/w/s for offered load. Latency up with load and queue flat means the device slowed; latency up with a deeper queue means you sent more work. %util only counts time with at least one request in flight, so on parallel devices 100% is not saturation.
solid answer
~50 sI compare four things across the before-and-after window. `r_await`/`w_await` are the per-request times *including* time spent queued, so they are what the application feels. `aqu-sz` is the average number of requests outstanding. `r/s` and `w/s` are the offered load, and `rareq-sz`/`wareq-sz` the average request size. If await doubled while IOPS and `aqu-sz` are unchanged, the device or the storage behind it got slower — a degraded array rebuilding, a cloud volume out of burst credit, a noisy neighbour. If await doubled *and* `aqu-sz` rose with IOPS, you are queueing your own extra load, which is Little's law rather than a fault. `%util` I largely ignore: it measures the fraction of wall time with at least one request in flight, which on a device that services many requests concurrently — NVMe, an SSD, a RAID set, a virtual volume — reaches 100% long before the device is out of capacity.
code
bash · 8 lines# live per-device view, skipping the boot average and idle devices
iostat -xyz 1 10
# who is actually issuing the I/O
pidstat -d 1 10
# the same devices, replayed from this month's sysstat archive
sar -d -p -f /var/log/sa/sa15 -s 02:30 -e 04:30go deeper
Know that iostat -x reports per-device latency in the await columns and that %util is not a capacity figure. Be able to run iostat -xz 1 and name the busiest device.
Explain that await includes queue waiting time and that aqu-sz is the average number of outstanding requests, then use the two together to say whether rising latency came from more load or a slower device.
Show the full separation: apply the relationship between throughput, queue depth and latency, check request size for a workload-shape change, compare the device-mapper and physical layers, and explain concretely why %util is unusable on parallel storage.
Own what gets alerted on and what gets bought. Latency percentiles and queue depth are the durable signals across hardware generations; a %util threshold inherited from single-spindle days will misfire on every modern volume, and capacity arguments need offered load and latency together rather than a utilisation number.
## The columns that matter `iostat -xz 1` gives per-device extended statistics. The ones you actually read during an incident: - **`r/s`, `w/s`** — completed read and write requests per second. The offered load. - **`rkB/s`, `wkB/s`** — throughput in kilobytes per second. - **`r_await`, `w_await`** — average time in milliseconds for a request to be served, measured from when it was inserted into the queue to when it completed. **This includes queue waiting time**, which is the single most important property of the number: it is what the application experiences, not what the device did. - **`aqu-sz`** — average queue length: the mean number of requests outstanding over the interval. (Called `avgqu-sz` in sysstat 11 and earlier.) - **`rareq-sz`, `wareq-sz`** — average request size in kilobytes, which tells you whether the workload shape changed. - **`%util`** — the percentage of wall-clock time during which at least one request was in flight. On sysstat 12 and later there is no single combined `await` column with `-x`; reads, writes, discards and flushes are broken out separately (`r_await`, `w_await`, `d_await`, `f_await`). On sysstat 11 and earlier you get one `await`, plus `avgqu-sz` and a `svctm` column. `svctm` was removed in sysstat 12 because the way it was derived made it unreliable — do not build an argument on it even where you still see it. ## Separating a slower device from a deeper queue Because await contains queue time, a rise in await has exactly two sources: each request takes longer, or each request waits behind more requests. Little's law relates them — roughly, `aqu-sz ≈ (r/s + w/s) × await` in seconds — so you can tell the two apart by looking at which side of the equation moved. ``` latency up, IOPS flat, aqu-sz flat -> the device (or what backs it) got slower latency up, IOPS up, aqu-sz up -> you are offering more load; queueing is the effect latency up, IOPS down, aqu-sz up -> the device is failing to keep up; capacity has fallen latency flat, IOPS up -> healthy scaling, nothing to fix ``` The third row is the alarming one: fewer completions with a growing queue means the device's effective capacity dropped underneath you while the workload stayed the same. Also check `rareq-sz`/`wareq-sz`. A workload that shifts from large sequential requests to small random ones will show higher per-request latency on spinning media, lower throughput, and no fault anywhere — the cause is a query plan or an application change, not the storage. ## Why %util lies on modern storage `%util` is computed as the fraction of the interval during which the device had at least one request outstanding. For a single-actuator hard disk that could only ever service one request at a time, that was a decent proxy for saturation, and a generation of engineers learned to alert on it. On anything that services requests **concurrently** the proxy breaks completely: - An **NVMe SSD** has many deep hardware queues and happily runs dozens of requests in parallel. One steady request in flight at all times gives `%util` 100% while the device is at a small fraction of its capacity. - A **RAID array** spreads requests over many physical devices. The array is never idle long before any individual member is busy. - A **virtualised or network-attached volume** presents a queue whose real capacity is elsewhere entirely. So `%util` at 100% means only "never completely idle". It is not a capacity metric, it cannot be aggregated meaningfully, and an alert on it will fire on healthy systems and stay silent on degraded ones. Latency and queue depth are the metrics that survive. ## The overnight-doubling workflow 1. **Get the history, not just now.** `sar -d -p -f /var/log/sa/saNN` for the overnight window if sysstat collection is on, so you can see when the change started rather than guessing. 2. **Watch live.** `iostat -xz 1` for a minute, and note the device names — with LVM or device-mapper the interesting device may be `dm-N` rather than the physical one, and both are reported. Compare the two layers: latency added between them is the mapping layer. 3. **Attribute it.** `pidstat -d 1` shows per-process read and write throughput, which tells you whether a new workload arrived — a backup, a rebuild, a migration, a runaway query. 4. **Look underneath the block layer.** A cloud volume out of burst credit, a RAID member resyncing, a failing device retrying, or a hypervisor neighbour all show the same shape from inside the guest: unchanged offered load, higher await. The guest counters cannot distinguish them, so at that point you go to the storage layer's own evidence. ## What to say out loud "Await includes queue time, so it is the number the application feels; `aqu-sz` tells me whether that time is queueing or service. If await moved without IOPS or queue depth moving, the device got slower and I go look underneath it. I do not use `%util` on anything with parallelism — it saturates at 100% while the device is half idle."
- Why did sysstat stop reporting `svctm`, and what did people use it for?People used it as "device service time excluding queueing", subtracting it from await to size the queue delay. It was derived rather than measured, and the derivation stopped being meaningful once devices serviced many requests concurrently, so sysstat 12 removed it. Use `aqu-sz` alongside await instead: the queue depth tells you directly how much of the latency is waiting.
- On a host using LVM, `iostat -x` shows both `dm-3` and `nvme0n1`. Which do you read?Both, and the comparison is the point. The `dm-` device is what the filesystem issues against and carries the latency the application sees; the physical device shows what the hardware delivered. Similar numbers mean the mapping layer is transparent; latency present on the mapper and absent on the device points at something in between, such as a snapshot's copy-on-write or a thin pool.
- An `iostat -x` capture shows w_await unchanged but application write latency has clearly risen. Where do you look next?At everything above the block device. Filesystem journal commits, `fsync` frequency, page-cache writeback behaviour, or an application that changed its durability settings can all add latency without any single request getting slower. `pidstat -d 1` for the process's actual write rate and a look at whether the workload's fsync pattern changed are the useful next steps.
saying these in an interview costs you the question
- Treating %util at 100% as device saturation
- Reading await as pure device service time
- Alerting on %util for SSD or RAID volumes
- Ignoring aqu-sz when latency rises
- Quoting svctm as if it were a measured value