A 32-CPU Linux host reports roughly 4% CPU utilisation in aggregate, yet one service on it is visibly slow. What can `mpstat -P ALL 1` show you that the aggregate figure hides, and what do its `%soft` and `%steal` columns mean?
answer
- an average across many CPUs hides one
- the summary line answers a different question
- which column is hot names the culprit
- one core can be full while the box looks idle
- some of that time was never yours
basics
~20 sThe aggregate is a mean across all CPUs, so a single core pinned at 100% contributes about 3% on a 32-CPU host. mpstat -P ALL 1 breaks utilisation out per CPU, exposing that imbalance; %soft is time in softirq handlers and %steal is time a vCPU was runnable while the hypervisor ran something else.
solid answer
~60 sAn aggregate of 4% on 32 CPUs is consistent with one CPU flat out and thirty-one idle, and that is a very common shape: a single-threaded hot path, or all of a NIC's receive processing landing on one core. `mpstat -P ALL 1` prints a line per CPU per interval, so the imbalance is immediately visible, and the column it lands in tells you what kind of imbalance it is. `%usr` and `%sys` on one CPU point at a task; `%soft` concentrated on one CPU points at interrupt or softirq handling — network receive processing is the usual culprit — and is fixed by spreading interrupts across CPUs rather than by adding capacity. `%steal` is different in kind: it is time this vCPU was ready to run but the hypervisor gave the physical CPU to someone else, so it is a signal about the host you are on, not about your workload. I also check whether the hot CPU is the same one every second, which distinguishes a pinned or affinity-bound workload from one that migrates.
go deeper
Know that mpstat -P ALL 1 shows one line per CPU while other tools show only an average, and that a single busy core barely moves the average on a wide machine.
Explain what the individual columns separate — user, system, hard and soft interrupt time, steal, idle — and why the column that is hot changes the diagnosis rather than just the magnitude.
Demonstrate the follow-through: check whether the hot CPU index is stable across intervals, take %soft hotspots to interrupt distribution rather than capacity, and treat sustained steal as evidence about the hypervisor that you escalate rather than tune.
Decide where per-CPU visibility is worth keeping. Aggregate host utilisation drives capacity planning badly on wide or virtualised fleets, and a platform that only records machine averages will systematically miss single-core saturation and hypervisor contention until a service is already degraded.
## Why the aggregate hides this Every aggregate CPU figure — `vmstat`'s CPU block, `iostat`'s CPU report, `mpstat`'s `all` line, `sar -u` — is a mean across CPUs. Divide by 32 and a fully saturated core becomes 3.1%. Divide by 96 and it becomes 1%. On wide machines the aggregate is close to useless for finding a bottleneck; it only tells you about capacity of the machine as a whole. `mpstat -P ALL 1` prints one row per CPU per interval, plus the `all` summary: ``` $ mpstat -P ALL 1 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle all 3.21 0.00 0.94 0.03 0.00 0.31 0.00 0.00 0.00 95.51 0 1.00 0.00 2.00 0.00 0.00 96.00 0.00 0.00 0.00 1.00 1 99.00 0.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 100.00 ... ``` Two saturated CPUs out of 32, and the summary line says 95% idle. The summary is not wrong; it is answering a question you did not ask. ## Reading which column is hot The column carries the diagnosis: - **`%usr` pinned on one CPU** — a single-threaded piece of work. Nothing about the host will fix it; the ceiling is one core's throughput. `pidstat -t 1` breaks usage down per thread and identifies the culprit. - **`%sys` pinned on one CPU** — heavy kernel-side work: syscall-bound loops, contention in the kernel, sometimes a driver. - **`%soft` pinned on one CPU** — softirq processing. On servers this is overwhelmingly network receive handling. A single-queue NIC, or a multi-queue NIC whose interrupts have all been steered to one CPU, funnels every packet's processing through one core. The symptom is packet-rate-dependent latency with the machine appearing idle. The remedy is distributing interrupt handling across CPUs, not buying a bigger box. - **`%irq` pinned on one CPU** — hard interrupt handling, same neighbourhood as above but the top half of the handler. - **`%iowait` concentrated on a few CPUs** — I/O waits accrue to the CPUs that had the requests outstanding, so the aggregate can look tiny while a couple of CPUs sit in it constantly. - **`%steal` non-zero anywhere** — you are on a virtual machine and the hypervisor is not giving you the CPU when you want it. Sustained double-digit steal is a host-level problem: a noisy neighbour, an oversubscribed host, or an exhausted CPU credit allowance. Your own code cannot fix it and the CPU columns for your own work will look deceptively light because the time was taken from you rather than spent by you. - **`%guest` / `%gnice`** — only meaningful when this machine is itself running virtual machines; that time is being spent running guest code. ## Same CPU or a different one? Watch several intervals and check whether the hot CPU index moves. - **Same index every second.** Something is bound there: a workload with a fixed CPU affinity, an interrupt steered to that CPU, or a single-queue device. Also the classic CPU 0 case, where a default interrupt configuration sends everything to the first core. - **A different index each second, one hot at a time.** A single-threaded workload the scheduler is free to migrate. The concurrency ceiling is one core wherever it runs. That one observation separates "a configuration is pinning work" from "the work itself cannot parallelise", and they have completely different fixes. ## Where to go next 1. `pidstat -t 1` — per-thread CPU, to name the thread burning the core. 2. `mpstat -P ALL 1` over a longer window, or `sar -P ALL` from the archive, to see whether the imbalance is chronic or new. 3. For a `%soft` hotspot, look at how the device's interrupts are distributed across CPUs and whether the NIC is presenting multiple receive queues. 4. For sustained `%steal`, take it to whoever owns the hypervisor or the instance sizing; there is nothing to tune inside the guest. ## The interview-sized answer "A machine-wide percentage is an average, and averages hide exactly the failure mode I care about — one saturated core. `mpstat -P ALL 1` shows the per-CPU picture, and which column is hot tells me whether it is my code, kernel work, interrupt handling, or the hypervisor taking time away from me."
- What does sustained `%steal` around 20% mean for your tuning options inside the guest?It means a fifth of the time your vCPU wanted to run, the hypervisor scheduled something else, and no change inside the guest recovers it. Your own utilisation figures will look light because that time was taken rather than spent. The fix is external: a less contended host, a larger or dedicated instance type, or an allowance that is not being exhausted.
- How would you tell a pinned workload from a single-threaded one that simply migrates?Watch the per-CPU rows over many intervals. If the same CPU index is hot every second, something is binding the work there — an affinity setting or an interrupt steered to that core. If the hot index wanders while only ever one is busy, the work is single-threaded and free to move. The first is a configuration fix, the second is an application-concurrency limit.
saying these in an interview costs you the question
- Concluding a wide host is idle from the aggregate figure
- Treating %soft as ordinary application CPU time
- Believing steal time can be tuned away inside the guest
- Assuming a busy CPU 0 is just coincidence
- Adding cores to fix a single-threaded hot path