Why can attaching strace to a busy production process slow it down by an order of magnitude, and what would you do instead when you cannot afford that?
answer
- cost scales with call count
- two stops per system call
- tracer must be scheduled each time
- filter before you attach
- watchdogs restart what you slow down
basics
~20 sstrace uses ptrace, which stops the traced process twice for every system call and wakes strace to inspect it. Each traced call becomes several context switches, so a syscall-heavy process can slow by ten to a hundred times.
solid answer
~50 sThe cost is structural, not a matter of output volume. ptrace stops the tracee on entry to every traced syscall and again on return; each stop wakes strace, which reads the registers and memory out of the target and resumes it. That turns a call costing under a microsecond into several context switches and a handful of extra ptrace operations, so a process making a million syscalls a second effectively cannot make them any more. Two orders of magnitude is normal in the worst case. On a live service that alone can trip health checks or a `WatchdogSec` timeout, and the timing change can hide the race you were chasing. So I narrow hard before attaching — `-e trace=` on the two or three calls I actually care about, and on strace 5.3 and later `--seccomp-bpf`, which uses an in-kernel filter so untraced syscalls never stop at all. I keep the window to seconds, and for a genuinely hot path I use a low-overhead in-kernel tracer instead.
code
bash · 2 linesstrace --seccomp-bpf -e trace=openat,connect -o /tmp/t.log /usr/local/bin/svc
strace -f -e trace=connect -o /tmp/t.log -p 4242 # attach briefly, Ctrl-C to detachgo deeper
Know that strace makes the traced program much slower and that this is inherent to how it works, not a configuration mistake — so you use it on a test instance or for a very short window, never casually on production.
Be able to explain the mechanism: ptrace stops the process on syscall entry and exit, and each stop schedules the tracer, so cost scales with the number of calls rather than their size. Name -e trace= filtering as the main mitigation.
Demonstrate operational judgement — the health-check and watchdog restarts, timing-dependent bugs that stop reproducing, the single-tracer limit — and describe a concrete procedure: summarise, filter with --seccomp-bpf where available, attach for seconds, detach cleanly, and prefer a drained instance.
Own the standing policy: which classes of problem justify a ptrace session on production at all, where the organisation should invest in low-overhead in-kernel tracing instead, and how to make that capability available so engineers are not tempted to reach for the dangerous tool under pressure.
## Where the time goes When strace attaches, the kernel is told to stop the tracee at syscall boundaries. For each traced system call the sequence is roughly: the tracee traps on entry and is put to sleep; the tracer is woken and scheduled; the tracer reads the tracee's registers and copies string arguments out of its address space; the tracer resumes the tracee; the tracee executes the actual call; it traps again on return; the tracer is woken again to read the return value; the tracer resumes it once more. That is two stops, at least two tracer wakeups, and several context switches per syscall — on top of an operation that for something like `clock_gettime` or a small `read` from page cache would have cost well under a microsecond. The overhead is therefore proportional to the *number* of syscalls, not to how much work each one does. A process doing a few large disk reads per second notices almost nothing. An event loop doing a million small `recvmsg`/`epoll_wait` cycles per second is crushed. A second, smaller cost is strace's own work: decoding arguments, formatting, and writing output. Writing to a terminal is much worse than writing to a file with `-o`, and printing long strings with a big `-s` value adds up. But even with output redirected to `/dev/null` the stop-and-schedule cost remains, which is why "just redirect the output" is not the answer. ## What that means on a live service The practical consequences are worth stating explicitly in an interview, because they are what separates someone who has done this on a production box from someone who has read about it: - **Health checks and watchdogs fire.** A service that suddenly takes 50x longer to answer may be marked unhealthy and restarted, destroying the state you were inspecting. If it is a systemd unit with `WatchdogSec=` set, the supervisor will restart it for missing its keepalive. - **Timeouts cascade.** Upstream callers time out, retry, and add load, so the trace changes the very behaviour under investigation. - **Heisenbugs vanish.** Races, lock contention and timing-dependent bugs frequently stop reproducing under tracing, because the syscall boundaries are now enormously wider than everything else. - **Only one tracer.** A process can have a single tracer, so if a debugger is attached you cannot attach at all, and while you are attached nobody else can. ## Reducing the cost properly **Narrow the syscall set.** `-e trace=openat,connect` or a class such as `-e trace=%file` is the single biggest lever, because unfiltered tracing pays the stop cost for every call the process makes, including the hot ones you do not care about. **Use in-kernel filtering.** Historically the filtering was done in strace *after* the stop, so a filtered trace still paid the stop cost for every syscall. strace 5.3 added `--seccomp-bpf`, which installs a seccomp filter so that only the syscalls matching `-e trace=` trap to the tracer at all; the rest run at full speed. On a process whose hot path is a syscall you are not interested in, this is the difference between unusable and routine. It applies to newly launched processes, so it is a `strace <cmd>` technique rather than an `-p` one. **Shorten the window.** Attach, count to five, detach with Ctrl-C. A five-second sample of a repeating problem is nearly always enough, and the blast radius is bounded. **Summarise before you transcribe.** `-c` or `-w` tells you which call to aim at; only then take a filtered raw trace. Note that summary mode is *not* cheaper — it still stops on every call — so the saving comes from the filtering you do next, not from `-c` itself. **Trace a canary, not the fleet.** If the workload is behind a load balancer, drain one instance and trace that. ## When to use a different instrument When the hot path itself is what you need to observe, ptrace is the wrong mechanism no matter how you tune it. The alternatives run the observation logic inside the kernel so the traced process is never stopped and never context-switched: in-kernel tracing frameworks attach to tracepoints and aggregate in kernel memory, with overhead measured in nanoseconds per event rather than microseconds. That is the honest answer to "we need this on a production path, continuously" — and knowing where the boundary lies between an ad-hoc ptrace session and a continuous low-overhead tracer is the judgement the question is testing. The reasonable default: strace is a scalpel for a broken or idle-ish process where you need exact arguments and errno values, and a liability on a hot path where you need aggregate behaviour.
- Does redirecting strace output to a file with -o remove most of the overhead?It removes a real but secondary cost — terminal writes and formatting — and is always worth doing. It does not touch the dominant cost, which is stopping the tracee twice per syscall and scheduling the tracer. A process making a million calls a second is still crippled with output going to /dev/null. Narrowing the traced syscall set is the lever that matters.
- Why does adding -e trace= alone not always help as much as you expect?Classically the filtering happened in strace after the stop, so every syscall still trapped to the tracer and only the printing was suppressed. strace 5.3 introduced `--seccomp-bpf`, which installs an in-kernel seccomp filter so non-matching syscalls never stop the process. With that, filtering finally reduces the stop count rather than just the output volume.
- You attach to a systemd-managed service and it gets restarted a few seconds later. What happened?Most likely the unit sets `WatchdogSec=` and the service missed its keepalive because tracing slowed it past the deadline, so systemd killed and restarted it. A failing health check upstream produces the same outcome. Check the unit's configuration and the journal for the unit before attaching, and if a watchdog is configured, trace a drained instance instead.
saying these in an interview costs you the question
- Believing the overhead comes mainly from printing output
- Attaching strace to a hot production path for minutes
- Expecting a timing-dependent race to still reproduce under tracing
- Thinking -c is cheaper because it prints less
- Assuming two people can attach tracers to one process at once