skip to content

A Linux daemon is already running and burning most of its time in the kernel. What does `strace -c -p <pid>` give you that a full trace does not, and how do you read its output?

level: middleimportance: should knowfreq 45%

answer

  1. aggregate instead of transcript
  2. one row per system call
  3. calls, errors, time columns
  4. system time by default, latency with -w
  5. ordering is what you trade away

basics

~20 s

strace -c aggregates instead of printing every call: one row per system call with call count, error count and time. It answers which syscalls dominate and which fail, at the cost of losing the sequence in which they happened.

solid answer

~50 s

`-p <pid>` attaches to a process that is already running, and `-c` suppresses the line-by-line output in favour of a summary table printed when you detach with Ctrl-C. Each row is one system call with the number of calls, the number that returned an error, and time attributed to it — by default the system CPU time spent inside the call, or wall-clock latency if you pass `-w` instead. That distinction matters: `-c` finds a process making millions of cheap calls, `-w` finds one blocked in a handful of slow ones. I read it top-down by time share, then check the errors column, because a large error count often turns out to be benign `EAGAIN` on non-blocking sockets rather than a fault. What you give up is ordering — a summary cannot show you the sequence that led to a failure, so `-C` prints both when you want each.

code

bash · 3 lines
bash
strace -f -c -p 4242
# ... let it run a few seconds, then press Ctrl-C to detach and print the table
strace -f -w -c -p 4242

go deeper

for a junior

Know that -p attaches to a process that is already running and that -c prints a per-syscall summary instead of every line, and that Ctrl-C detaches without killing the target.

for a middle

Be ready to walk the summary columns and explain what a huge call count or a big error count each suggest. Say clearly that -c totals system time while -w totals wall-clock latency, and pick the right one for a given symptom.

for a senior

Demonstrate the workflow: summarise first to choose a target, then take a short filtered trace for detail. Note that attaching costs the same as full tracing, that mean usecs/call hides outliers, and that slowing a watchdogged service can restart it out from under you.

for a principal

Own the judgement about touching a live production process at all: what a short attach can cost a latency-sensitive service, when a low-overhead in-kernel tracer is the responsible instrument instead, and what should be instrumented so nobody needs to attach next time.

## Two different questions A raw trace answers "what did this process do, in order?" A summary answers "where is its kernel time going?" They are different investigations, and picking the wrong one wastes the window you have on a live process. When a daemon is already misbehaving and you cannot restart it, `strace -c -p <pid>` is the low-commitment first look: attach, let it run for a few seconds, press Ctrl-C, read the table. ## Attaching and detaching `-p <pid>` attaches to a running process; the tracee keeps running and you can attach to several at once by repeating `-p` or passing a comma-separated list. Pressing Ctrl-C detaches cleanly and leaves the process running — that is the intended way to stop, and it is also when the `-c` table is printed. Add `-f` and strace follows the process's threads and children, which for a multi-threaded server is usually what you want. One hard constraint: a process can have exactly one tracer. If gdb, another strace, or a supervisor is already attached, your attach fails with `Operation not permitted` even as root, and no amount of privilege changes that. ## Reading the table The summary has a row per syscall and these columns: ``` % time seconds usecs/call calls errors syscall ------ ----------- ----------- --------- --------- ---------------- 58.31 0.482113 4 120528 futex 21.02 0.173784 1 173112 86556 read ``` - **calls** — how many times the syscall was made. A very large number of very cheap calls is its own finding: a chatty loop reading one byte at a time, or a poll spinning with a zero timeout. - **errors** — how many returned an error. Read this sceptically. `EAGAIN` on a non-blocking socket, `EINTR` on an interrupted call, and `ENOENT` from cache-miss lookups are all routine. A high count is a signal to look, not proof of a fault. - **seconds / % time** — by default the *system CPU time* attributed to the call. This is what you want for a CPU-bound process. - **usecs/call** — the mean, which hides a bimodal distribution. Two calls at 500 ms among ten thousand at 2 µs will not stand out here. ## -c versus -w, the distinction people miss `-c` summarises system time; `-w` summarises the wall-clock difference between the start and end of each call. A process blocked for seconds in `read` on a slow network filesystem consumes almost no system time, so it barely registers under `-c` and dominates under `-w`. The rule of thumb: `-c` for "who is burning CPU in the kernel", `-w` for "what is this process waiting on". `-C` (capital) prints the normal per-call output *and* the summary, for when you need both the shape and the sequence. ## What a summary cannot tell you Ordering, arguments, and causality are all gone. If the question is "which file did it fail to open before it exited", the summary tells you only that some `openat` calls returned ENOENT. Switch to a filtered raw trace — `-e trace=file`, `-e trace=network`, or an explicit list — once the summary has told you where to aim. The productive pattern is summary first to pick the target, filtered trace second to see the detail, because the filtered trace is also dramatically cheaper than tracing everything. ## Cost and the safety of attaching Attaching is not free even in summary mode: strace is still stopping the process at every syscall to count it, so the overhead is the same order as a raw trace even though the output is small. Keep the window short, and prefer narrowing with `-e trace=` when you already have a hypothesis. If the process is behind a health check or a watchdog, remember that slowing it down can itself trigger a restart — which is the fastest way to destroy the state you were trying to inspect. ## Companion flags worth knowing `-o file` keeps output away from your terminal, `-T` prints the time spent in each individual call in raw mode, and `-tt` gives absolute timestamps you can line up against application logs. For a first look at a suspect process, `strace -f -c -p <pid>` for five seconds, then Ctrl-C, is a complete diagnostic step on its own.

  • The summary shows 86,000 errors on read. Should that alarm you?
    Not on its own. An event-driven server on non-blocking sockets gets `EAGAIN` every time it reads a descriptor with no data ready, and `EINTR` appears whenever a signal interrupts a blocking call. Both count as errors in the table and are entirely normal. Confirm by taking a short raw trace and looking at which errno actually dominates before treating it as a fault.
  • Your attach fails with "Operation not permitted" although you are root. What is the most likely reason?
    Something else is already tracing the process — a debugger, a supervisor, or a forgotten strace session — and the kernel allows only one tracer per process. Check `TracerPid` in `/proc/<pid>/status`; a non-zero value names the current tracer. The other possibilities are a restrictive `ptrace_scope` under Yama, or running inside a container without CAP_SYS_PTRACE.

saying these in an interview costs you the question

  • Reading the errors column as proof something is broken
  • Believing -c is cheap because its output is small
  • Expecting a summary to show the order calls happened in
  • Using -c to find a process blocked on slow I/O
  • Assuming Ctrl-C kills the process you attached to

context