You want to time a function inside a running user-space application using bpftrace's `uprobe` and `uretprobe` probes. How does a uprobe actually attach to the process, roughly what does each hit cost compared with a kernel tracepoint, and which kinds of binary make this approach unreliable?
answer
- a breakpoint in a private copy of the page
- the trap is what you pay for
- the code is in the library, not the binary
- return probes rewrite the stack
- markers the author put there on purpose
basics
~20 sA uprobe patches a breakpoint instruction into a private copy of the target's executable page, so every thread that reaches that address traps into the kernel and runs the BPF program. Each hit costs roughly a microsecond, far more than a tracepoint. Stripped, statically linked, Go and JIT-compiled binaries make it unreliable.
solid answer
~60 sYou give bpftrace a path and a symbol — `uprobe:/lib/x86_64-linux-gnu/libc.so.6:malloc` — and the kernel resolves that to a file offset and installs a breakpoint there, in a copy-on-write copy of the mapped page, so other users of the same binary are unaffected. Every thread that executes that instruction traps into the kernel, runs the BPF program, and is single-stepped back onto the original instruction. That trap is why a uprobe costs on the order of a microsecond per hit, an order of magnitude more than a kernel tracepoint, so on a hot function the overhead is real. Reliability depends on the binary: you need a symbol, so a stripped binary means finding the offset yourself; you must target the object that actually contains the code, which for a dynamically linked program means the shared library, not the executable. `uretprobe` is the fragile part — it works by hijacking the return address on the stack, so runtimes that move or rewrite stacks, and languages whose calling convention isn't the C one, give wrong or missing results. Where the application ships USDT markers, use those instead.
code
bash · 2 lines# Attach to the object that actually contains the code, and scope to one pid
bpftrace -e 'uprobe:/lib/x86_64-linux-gnu/libc.so.6:getaddrinfo /pid == 4242/ { @[str(arg0)] = count(); }'go deeper
Know that uprobes trace user-space functions without changing the program, that you must name the file and the symbol, and that for a dynamically linked program the function usually lives in a shared library.
Explain the breakpoint-in-a-private-page mechanism and why each hit traps into the kernel, making a uprobe roughly an order of magnitude more expensive than a tracepoint. Know that stripped binaries need offsets.
Show judgement about when not to use them: hot functions where the trap cost dominates, Go and JIT runtimes where return probes and positional arguments are unreliable, and cases where USDT markers give a stable interface instead.
Own the tradeoff between asking application teams to ship USDT markers or first-class instrumentation versus relying on symbol-level probing, and what that commits both sides to as the software evolves.
## How the attachment works A uprobe is dynamic instrumentation of user space, driven from the kernel. You name a **file and a symbol**, and the kernel resolves the symbol to an offset within that file. It then arranges that, wherever that file is mapped executable, the instruction at that offset is replaced by a breakpoint. The replacement happens in a **private, copy-on-write copy of the page**, not in the file on disk and not in the page cache copy other processes share. So instrumenting `malloc` in libc does not corrupt the shared library or affect processes you did not target when you have scoped the probe. When a thread executes the patched instruction it traps into the kernel, the kernel runs your BPF program, then executes the displaced original instruction out of line and returns control. Because the attachment is keyed to the *file*, a uprobe on a library applies to every process that maps it unless you filter — hence the near-universal `/pid == $1/` predicate in real scripts. ## What it costs Every hit is a trap: a full transition into the kernel and back, plus the out-of-line single step. That puts a uprobe in the region of a microsecond per hit, versus a tracepoint or an optimised kernel probe measured in tens or low hundreds of nanoseconds. Concretely, instrumenting a function called ten thousand times a second is unremarkable; instrumenting one called ten million times a second will visibly slow the application. The cost is per *hit*, so scoping to one process reduces the event rate but not the per-event price. The practical discipline is the same as for kernel tracing: filter in the predicate so the expensive work is skipped, and aggregate rather than print. ## Getting a symbol to attach to The symbol must be findable in the object file. Three situations recur: **Dynamically linked programs.** The code for a library function lives in the library, so you attach to the `.so`, not the executable: ``` bpftrace -e 'uprobe:/lib/x86_64-linux-gnu/libc.so.6:getaddrinfo { printf("%s %s\n", comm, str(arg0)); }' ``` Attaching to the executable for a libc function finds nothing — a very common first mistake. `ldd` on the binary, or the maps of the running process, tell you which object is actually loaded. **Statically linked programs.** Everything is in the one binary, so you attach there — but only if the symbols survived. **Stripped binaries.** With no symbol table there is nothing to resolve. You can still attach if you know the *offset*, which you obtain from a copy of the binary that has symbols or from separate debug information. bpftrace accepts an address form for exactly this case. Distribution debuginfo packages are the usual source. You can enumerate what is attachable with a listing, e.g. `bpftrace -l 'uprobe:/lib/x86_64-linux-gnu/libc.so.6:*'`, and the classic binutils tools (`nm -D`, `readelf -s`, `objdump`) show the symbol table directly. ## Where uretprobes go wrong A `uretprobe` cannot simply patch the function's return instruction — a function may have many exits and may be entered from anywhere. Instead the kernel patches the *entry*, saves the caller's return address, and substitutes an address of its own so that when the function returns, control lands back in the kernel, which then fires the probe and jumps to the real return address. This is a stack manipulation, and it breaks when the runtime has its own opinions about the stack: - **Go.** Goroutine stacks start small and are moved when they grow, and Go's calling convention has not matched the C one that bpftrace's positional `arg0`, `arg1` built-ins assume. Return probes on Go binaries are a well-known source of wrong results and crashes; entry probes with hand-decoded arguments are more workable, but Go tracing is genuinely specialist work. - **Runtimes that do their own unwinding or non-local exits.** Anything that longjmps past the instrumented return, or unwinds through it for exception handling, can leave the substituted return address unused. - **JIT-compiled code.** Java, Node and similar runtimes generate machine code at runtime. There is no file offset and no symbol in an ELF object to attach to at all, so uprobes are the wrong instrument entirely. ## The stable alternative: USDT USDT (userland statically defined tracing) probes are markers the application's own authors compiled in — a name, a provider, and a defined set of arguments, recorded in an ELF note. Attaching to one costs about the same as a uprobe, but you get an interface the application maintainer intends to keep, at a semantically meaningful point, rather than a symbol that the next release may inline away. glibc ships USDT markers, as do several language runtimes when built with the option enabled. bpftrace attaches to them with `usdt:/path/to/binary:provider:name` and lists them with `bpftrace -l 'usdt:/path/to/binary:*'`. If the application you are tracing has them, prefer them. ## The practical checklist Confirm which object holds the code; confirm the symbol exists there; filter by pid; prefer entry probes and aggregation over return probes and per-event printing; and when the target is Go or a JIT runtime, reach for the runtime's own instrumentation before reaching for uprobes.
- You attach a uprobe on a libc function to one process, but worry about the rest of the fleet on that host. Can the patch affect other processes?Not their execution. The breakpoint is installed in a private copy-on-write copy of the mapped page, so the on-disk library and the shared page-cache copy are untouched. What does apply broadly is scope: a uprobe keyed to a file fires for every process mapping it, so without a pid filter your program runs for all of them — costing them the trap even if you discard the event.
- Why are uretprobes specifically problematic on Go binaries?Because a uretprobe works by substituting the return address on the stack, and Go moves goroutine stacks when they grow, copying and rewriting them. The substituted address can be relocated or invalidated. Go's calling convention also doesn't match the C one that positional argument built-ins assume, so even entry-probe arguments need hand decoding. Entry probes are workable with care; return probes are not.
- What would you use instead of a uprobe on a Java or Node application?Not uprobes — JIT-compiled code has no ELF symbol or file offset to attach to. Use the runtime's own instrumentation: USDT markers where the build provides them, or the runtime's profiling and tracing interfaces. Uprobes can still be useful on the runtime's native layer, such as libc calls it makes, but not on application-level methods.
saying these in an interview costs you the question
- Thinks a uprobe modifies the binary or library on disk
- Attaches to the executable for a function that lives in libc
- Assumes uprobes cost the same as a kernel tracepoint
- Uses uretprobes on Go binaries and trusts the numbers
- Expects uprobes to reach Java or Node application methods