You run `perf stat` against a CPU-bound Linux process and see roughly 0.3 instructions per cycle together with a high cache-miss count. What is that telling you about the workload, and what would an IPC near 3 mean instead?
answer
- how well the core is working, not where
- instructions divided by cycles
- stalled versus retiring
- low means waiting on memory
- hardware counters need a real PMU
basics
~20 sLow instructions-per-cycle with heavy cache misses means the CPU is stalling on memory rather than doing work — the fix is data layout and access patterns. A high IPC means the core is genuinely retiring instructions, so the code is doing too much work, and you optimise the algorithm.
solid answer
~50 sIPC is `instructions` divided by `cycles`, and `perf stat` prints it directly as "insn per cycle". At around 0.3 the core is spending most of its cycles stalled — it is issued work it cannot complete because operands are not there — and the accompanying `cache-misses` count points at memory as the reason. That is a data problem: the working set does not fit, or the access pattern defeats prefetching, so the answer is better locality rather than fewer instructions. An IPC near 3 says the opposite: the pipeline is well fed and the core is retiring close to its issue width, so the CPU is genuinely executing the work you asked for. Then the profile is telling you to remove work — a better algorithm, less redundant computation — because there is no stall to reclaim.
code
bash · 8 lines# attach to a running process and count for ten seconds
perf stat -p 4242 sleep 10
# detailed run of a command, including cache breakdown
perf stat -d ./workload
# only the counters needed for the IPC and miss-ratio verdict
perf stat -e cycles,instructions,cache-references,cache-misses,branch-misses ./workloadgo deeper
Know that perf stat counts events for a whole run instead of sampling, and that instructions per cycle is the headline number it derives for you.
Explain the mechanics: IPC is instructions over cycles, below roughly 1 means the core is stalling, and cache-misses only means something as a fraction of cache-references.
Demonstrate that the verdict changes the fix — memory-bound work needs better data locality while compute-bound work needs less work — and mention that a cloud VM without an exposed PMU cannot produce these counters at all.
Own the question of whether microarchitectural tuning is worth anyone's time on this workload versus cheaper wins in architecture or capacity, and what host types your fleet would need for these counters to be available when it is.
## What perf stat actually gives you `perf record` answers "where". `perf stat` answers "how well". It does not sample and it does not attribute anything to a function; it programs the CPU's performance monitoring unit (PMU) to *count* events for the duration of a command or an attached process, then prints the totals. ```bash perf stat -p 4242 sleep 10 perf stat -d ./workload perf stat -e cycles,instructions,cache-references,cache-misses,branch-misses ./workload ``` A typical run prints `task-clock`, `context-switches`, `cpu-migrations`, `page-faults`, then the hardware counters `cycles`, `instructions`, `branches` and `branch-misses`, with derived figures in the comment column — including the one that matters most here, `insn per cycle`. ## Why IPC is the first number to read A modern superscalar core can retire several instructions per cycle. IPC is the ratio of instructions retired to cycles elapsed, so it measures how much of the core's capability the workload is actually using. - **Low IPC (roughly below 1)** — the core is *stalled* for most cycles. It has work queued but cannot complete it, almost always because it is waiting for data. Memory is the usual culprit: a last-level cache miss costs on the order of a couple of hundred cycles, during which the core retires nothing. - **High IPC (roughly 2 and above)** — the pipeline is full and the core is retiring near its issue width. The machine is not the bottleneck; the amount of work is. This single split changes the entire optimisation strategy, which is why interviewers like the question. Two processes can both show 100% CPU, and one of them is doing useful work while the other is waiting on RAM with the lights on. ## Confirming the memory hypothesis `cache-misses` alone means nothing — a big program does a lot of everything. What you want is the **ratio** perf prints for you: `cache-misses` as a percentage of `cache-references`. A few percent is unremarkable; tens of percent alongside sub-1 IPC is a strong memory-bound signal. `perf stat -d` adds the level-1 data cache and last-level cache breakdown, which localises the miss. On CPUs that expose them, `stalled-cycles-frontend` and `stalled-cycles-backend` split the stall further: frontend stalls point at instruction supply (icache misses, branch mispredicts), backend stalls at data supply and execution resources. Not every CPU model exposes these, and perf prints `<not supported>` when it cannot count them. `branch-misses` as a share of `branches` is the other classic stall source. A misprediction throws away the speculative work in flight; a data-dependent branch inside a hot loop can dominate a profile while every individual instruction looks cheap. ## The trap: no PMU in a virtual machine This is the practical gotcha, and it catches people the first time they try this on cloud infrastructure. Hardware counters come from the physical CPU's PMU, and a hypervisor must explicitly expose it to the guest. On most cloud VMs it does not. The symptom is unmistakable: `cycles`, `instructions`, `cache-misses` and friends all print `<not supported>`, and there is no IPC line at all. Software events — `task-clock`, `context-switches`, `page-faults`, `cpu-clock` — still work, because the kernel counts those itself. So an honest answer includes the caveat: this analysis needs bare metal, a dedicated host, or a hypervisor with PMU virtualisation enabled. On an ordinary VM you fall back to sampling with the software clock event, which still tells you *where* time goes even though it cannot tell you *why* the core was slow. ## Multiplexing and scaling A PMU has a limited number of physical counters. Ask `perf stat` for more events than there are counters and the kernel time-slices them, scaling the results up to a full-run estimate. perf tells you it did this by printing a percentage next to the event. When you see figures like 33% enabled, treat the counts as estimates — and if the numbers matter, ask for fewer events per run. ## What to do with each verdict **Memory-bound (low IPC, high miss ratio):** shrink or reorganise the data. Structure-of-arrays instead of array-of-structures, contiguous iteration instead of pointer chasing, smaller records so more fit per cache line, batching so a hot structure stays resident. **Compute-bound (high IPC):** the CPU is doing exactly what you asked, so ask for less. Now `perf record` and a flame graph are the right next step, because the fix is finding which function to remove work from. **Neither (low IPC, low miss ratio):** look at `branch-misses`, at lock contention showing up as context switches, or at the possibility that the process is not really CPU-bound at all.
- `perf stat` on your cloud VM prints `<not supported>` for cycles and instructions. What happened, and what can you still measure?Hardware events come from the physical CPU's PMU, and the hypervisor is not exposing it to the guest, so the kernel cannot program those counters. Software events still work — `task-clock`, `context-switches`, `cpu-migrations`, `page-faults` — and sampling still functions via the `cpu-clock` event. You get where time is spent, just not the microarchitectural reason. Bare metal or a PMU-enabled host is the only real fix.
- You asked perf stat for a dozen events and it printed a percentage beside each one. What is that percentage?The share of the run during which that event was actually counted. A PMU has a fixed number of physical counters, so when you request more events than fit, the kernel multiplexes them in time slices and scales the observed counts up to a full-run estimate. Low percentages mean the numbers are estimates and can be noisy — request fewer events per run when precision matters.
saying these in an interview costs you the question
- Treats a raw cache-misses count as bad without the ratio
- Says high IPC means the code is efficient and needs no work
- Assumes hardware counters always work inside a cloud VM
- Confuses perf stat with a profiler that names hot functions
- Ignores the multiplexing percentage and quotes scaled counts as exact