A Linux box is pinned at 100% CPU and you have already identified the guilty process. What does `perf top` show you that a process-level view cannot, and how does it collect that information?
answer
- which function, not which process
- the profiler interrupts, it does not instrument
- overhead is a share of samples
- attach live; record to keep it
basics
~20 sperf top is a live sampling profiler. It interrupts the CPU many times a second, records which instruction was executing, resolves it to a function, and ranks functions by how many samples landed in them — so it names the hot code, not just the hot process.
solid answer
~50 sA process-level view tells you *who* is burning CPU; `perf top` tells you *what code* is burning it. It asks the kernel's perf subsystem for a periodic event — by default the hardware `cycles` counter, falling back to the software `cpu-clock` event when no PMU is exposed — and on each overflow it captures the instruction pointer, the thread, and optionally a call chain with `-g`. It then aggregates those samples live and shows an `Overhead` percentage per symbol, alongside the shared object the address came from. Nothing has to be recompiled or restarted: I attach to the running process with `perf top -p <pid>`, and can narrow further with `-e cycles:u` for user space only or `-F 99` to lower the sample rate. Because it samples, the percentages are shares of samples, not measured milliseconds.
code
bash · 5 lines# live profile of one process, 99 samples per second, with call graphs
perf top -p 4242 -F 99 -g
# user-space cycles only, grouped by shared object then symbol
perf top -e cycles:u --sort dso,symbolgo deeper
Know that perf top is a live sampling profiler that names hot functions rather than hot processes, and that you attach it to a running PID without restarting anything.
Be ready to explain the sampling mechanism — a periodic hardware or timer event, one instruction pointer captured per overflow — and why the Overhead column is a share of samples rather than measured time.
Show triage judgement: user versus kernel symbols as a first split, -g only once a symbol dominates, a lower -F on a saturated box, and the recognition that an empty CPU profile on a slow service redirects the whole investigation.
Own the tradeoff between always-on and on-demand profiling: what sampling overhead you are willing to pay on production hosts, and what has to be true about symbols and build flags fleet-wide before a profile taken at 3am is actually readable.
## Two different questions about a busy CPU Process-level tools answer an accounting question: over the last interval, how much CPU time did each *process* consume? That is enough to find the offender and nowhere near enough to fix it. `perf top` answers the next question down: inside that CPU time, which *functions* were actually executing? ## How sampling works `perf top` is a sampling profiler. It opens a performance event through the kernel's perf subsystem and asks to be notified periodically. The default event is the hardware `cycles` counter from the CPU's performance monitoring unit (PMU); on hosts where no PMU is exposed — many virtual machines — perf falls back to the software `cpu-clock` event, which is driven by a timer instead. Each time the counter overflows, the kernel takes a **sample**: the instruction pointer at the moment of the interrupt, the PID and TID, the CPU, and — if you passed `-g` — a call chain walked back from that point. Crucially this is *not* instrumentation. The target is not recompiled, not restarted, not wrapped in anything. You attach to a process that is already misbehaving in production, which is exactly the situation the question describes. Because it samples, every number you read is statistical. A symbol credited with 40% of the profile is one in which 40% of the *collected samples* landed. With a few thousand samples that is a reliable estimate of where CPU time goes; with fifty samples it is noise. This is the single most common misreading of profiler output. ## Reading the output The live table has three columns that matter: - **Overhead** — the share of samples attributed to this symbol. - **Shared Object** — the DSO the sampled address belongs to: the executable itself, a shared library such as `libc.so.6`, or `[kernel.kallsyms]`. - **Symbol** — the function name, prefixed with `[.]` for a user-space address and `[k]` for a kernel one. That prefix is a fast triage signal on its own. A profile dominated by `[k]` frames says the time is going into kernel work — syscalls, memory management, network processing — while a profile dominated by `[.]` frames in your own binary points at application code. ## Narrowing the view ```bash perf top -p 4242 # one process only perf top -e cycles:u # user-space cycles only perf top -F 99 -g # 99 Hz, with call graphs perf top --sort dso,symbol # group by library first ``` The `:u` and `:k` event modifiers restrict sampling to user or kernel mode. Lowering `-F` reduces the profiler's own cost on an already-saturated box. Pressing `a` on a selected symbol drops into annotation, attributing samples to individual instructions inside that function — useful when a single loop is the problem. ## Where it falls down 1. **It is ephemeral.** `perf top` is a live view with no artifact at the end. For anything you want to keep, compare, or hand to someone else, use `perf record` to write a `perf.data` file and analyse it with `perf report`. 2. **Symbols may be missing.** Stripped binaries and packages without debug information produce rows of raw hex addresses or `[unknown]`. The fix is installing the matching debug symbols (`-dbgsym` packages on Debian/Ubuntu, `dnf debuginfo-install` on Fedora/RHEL), not changing anything about the sampling. 3. **Kernel symbols can be hidden.** If `kernel.perf_event_paranoid` or `kernel.kptr_restrict` are restrictive, kernel addresses resolve to zeros and you lose the `[k]` side of the picture. 4. **It only sees on-CPU time.** A process that is slow because it is blocked — waiting on disk, on a lock, on a socket — is not running, so it generates no CPU samples. A quiet profile on a slow service is itself a finding: the problem is not CPU. 5. **The profiler costs something.** At high sample rates on a busy machine, the sampling interrupts themselves consume CPU; the kernel also throttles beyond `kernel.perf_event_max_sample_rate`. ## What a good answer sounds like "I'd attach `perf top -p <pid>` for thirty seconds, look at whether the overhead is concentrated in user or kernel symbols, and if one symbol dominates I'd add `-g` to see who is calling it. Then I'd switch to `perf record` to capture something I can keep." That sequence — live triage, then capture — is what the interviewer is listening for.
- When would you reach for `perf record` instead of `perf top`?Whenever the artifact matters. `perf top` is a live view that leaves nothing behind, so it suits thirty seconds of triage. `perf record` writes a `perf.data` file you can analyse with `perf report`, diff against a later capture, turn into a flame graph, or send to someone else. Bursty problems also need recording, because the burst is over before you finish reading the live table.
- You attach perf top to a service that users call slow, and the profile is almost empty. What does that tell you?That the service is not CPU-bound. Sampling only fires while a thread is on-CPU, so time spent blocked on disk, a lock, a database round trip or a socket read produces no samples at all. That is a useful negative result: it redirects the investigation towards I/O latency, contention or an upstream dependency rather than towards the code.
A process-level view is the electricity bill for the building; perf top is walking the floors with a thermal camera to see which room is actually running the heaters.
saying these in an interview costs you the question
- Says perf top measures exact time spent per function
- Thinks the target must be restarted or rebuilt first
- Reads Overhead as a percentage of wall-clock duration
- Assumes a quiet profile means the service is healthy
- Reports hex addresses without noticing symbols are missing