skip to content

You are handed a Linux box with bpftrace installed and asked to watch which files a service opens. Walk through what the one-liner `bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s\n", comm, str(args->filename)); }'` actually does, and explain how you would have discovered that probe and its argument fields in the first place.

level: juniorimportance: should knowfreq 40%

answer

  1. probe, optional filter, action block
  2. where the filename pointer really lives
  3. one flag lists what exists, another its fields
  4. a helper copies bytes into the program
  5. system-wide unless you say otherwise

basics

~20 s

The one-liner attaches a small BPF program to the openat syscall-entry tracepoint and prints the calling process name and the filename for every openat on the whole machine. bpftrace -l lists available probes and bpftrace -lv shows a tracepoint's argument fields.

solid answer

~50 s

A bpftrace program is a list of `probe { action }` blocks. Here the probe is `tracepoint:syscalls:sys_enter_openat`, a static instrumentation point the kernel exposes at the entry of the `openat` system call; the action runs in the kernel every time any process on the box calls it. `comm` is a bpftrace built-in holding the current task's command name, and `args` is the tracepoint's argument struct, so `args->filename` is the path pointer the caller passed. That pointer lives in the caller's memory, so `str()` is needed to copy the bytes into the program and turn them into a printable string. To find the probe I'd run `bpftrace -l 'tracepoint:syscalls:*open*'` to see what exists, then `bpftrace -lv tracepoint:syscalls:sys_enter_openat` to see the field names before writing anything. As written it traces the whole system, so in practice I'd add a filter such as `/comm == "nginx"/` or `/pid == 1234/`.

code

bash · 3 lines
bash
# List the probes that exist, then their argument fields
bpftrace -l 'tracepoint:syscalls:*open*'
bpftrace -lv tracepoint:syscalls:sys_enter_openat

go deeper

for a junior

Be able to read a one-liner out loud: this is the probe, this is the action, comm is the process name. Know that bpftrace -l is how you find out what can be traced, and that it needs root.

for a middle

Explain why a string field needs str() — the pointer is in the caller's memory and BPF must copy it through a helper with a fixed-size buffer. Show that you would add a filter so the probe fires only for the process under investigation.

for a senior

Show judgement about cost: per-event printf is bounded by the ring buffer and the terminal, so on a busy host you aggregate in the kernel and print once. Mention that syscall entry and exit are separate probes and how you correlate them.

for a principal

Frame when an ad-hoc one-liner is the right instrument at all versus adding permanent instrumentation, and what you would standardise for on-call engineers so a paste-in trace on a production host is a routine, bounded action rather than a risk.

## What bpftrace is bpftrace is a front end for eBPF tracing: you write a short, high-level script, bpftrace compiles it to BPF bytecode, the kernel's verifier checks it, and the kernel runs it at the instrumentation points you named. Nothing is patched into the traced application and no process is stopped — the program runs in kernel context on the same thread that triggered the event and then execution continues. That is why one-liners like this can be pasted onto a production box. ## The shape of a program Every bpftrace program is one or more blocks of the form: ``` probe[,probe] /filter/ { action } ``` - **probe** — where to attach. The type comes first: `tracepoint:`, `kprobe:`, `kretprobe:`, `uprobe:`, `usdt:`, `profile:`, `interval:`, plus the special `BEGIN` and `END` blocks. - **filter** (optional, in slashes) — a predicate; the action only runs when it is true. - **action** — the statements to run, in braces. ## Reading this particular probe `tracepoint:syscalls:sys_enter_openat` names a *static tracepoint*: a marker the kernel developers declared in the source and that the kernel exposes with a documented set of fields. The `syscalls:` category holds an entry (`sys_enter_*`) and exit (`sys_exit_*`) tracepoint for each system call. Attaching here means the block runs on entry to `openat()`, before the kernel has done the work, which is why the filename is available but the resulting file descriptor is not — for that you would also attach to `tracepoint:syscalls:sys_exit_openat` and read the built-in `args->ret`. ## The built-ins in the action - `comm` — the command name (up to 16 characters) of the task that triggered the probe. - `pid`, `tid` — the process and thread id of that task. - `args` — a pointer to the tracepoint's argument struct, with one member per declared field. For `sys_enter_openat` those include `dfd`, `filename`, `flags` and `mode`. - `nsecs` — a nanosecond timestamp, the usual basis for latency measurements. `printf()` works like C's, with the format string first. ## Why `str()` is required `args->filename` is a **pointer** into the calling process's address space, not a string. A BPF program may not simply dereference arbitrary memory — the verifier rejects that — so it must copy through a helper. `str()` is bpftrace's wrapper over that copy, and it returns a fixed-size buffer (64 bytes by default, tunable with the `BPFTRACE_MAX_STRLEN` environment variable), so very long paths are truncated. Printing `args->filename` directly gives you a meaningless hex address, which is one of the most common beginner mistakes. ## Discovering what you can trace You never have to guess: ``` bpftrace -l 'tracepoint:syscalls:*open*' # which probes exist bpftrace -lv tracepoint:syscalls:sys_enter_openat # and their fields ``` `-l` lists matching probes and accepts shell-style wildcards; `-lv` adds the argument list for tracepoints. Listing is also a cheap existence check before you ship a script to a fleet: if `-l` prints nothing on a host, the probe is not there and the script will fail to attach. ## Scope and cost As written the one-liner is system-wide: every `openat` by every process, including the terminal you are typing in, produces a line. Two things follow. First, **filter early**. A filter block is evaluated in the kernel, so `/comm == "nginx"/` or `/pid == 1234/` discards unwanted events before anything is copied to user space: ``` bpftrace -e 'tracepoint:syscalls:sys_enter_openat /comm == "nginx"/ { printf("%s\n", str(args->filename)); }' ``` Second, **per-event printing has a ceiling**. Each `printf` pushes a record through a per-CPU ring buffer that a user-space thread drains; if events arrive faster than they are drained, bpftrace warns that it lost events, and the terminal itself becomes the bottleneck. For high-rate events the idiomatic answer is to aggregate into a map in the kernel (`@[comm] = count();`) and print once at exit, rather than streaming a line per event. Finally, this needs root (or the BPF-related capabilities) — an unprivileged user cannot load a tracing program.

  • If you also wanted the file descriptor that each open returned, what would you change?
    Add a second block on `tracepoint:syscalls:sys_exit_openat` and read `args->ret`, which carries the return value — a non-negative fd or a negative errno. Because the filename is only available at entry, the usual pattern is to stash it in a map keyed by `tid` in the entry block and look it up in the exit block, deleting the entry afterwards so the map does not grow.
  • Why does printing `args->filename` without `str()` show a hex number instead of a path?
    Because the tracepoint field is a `char *` pointing into the caller's address space, and `%s` on a raw pointer is not how BPF reads memory. A BPF program cannot dereference arbitrary addresses — it must go through a copy helper, which is what `str()` wraps. Without it you print the pointer value itself.
  • The one-liner floods your terminal on a busy host. What is the idiomatic fix?
    Stop printing per event and aggregate in the kernel: `@[comm, str(args->filename)] = count();` accumulates into a map and bpftrace prints the totals when you hit Ctrl-C. Add a filter to narrow to one process as well. This keeps the data in the kernel instead of pushing a record per event through the ring buffer.

saying these in an interview costs you the question

  • Thinks bpftrace stops or slows the traced process like a debugger
  • Prints args->filename directly and expects a path
  • Assumes the one-liner only traces the service you care about
  • Believes you must recompile or restart the application to trace it
  • Says tracepoint arguments have to be guessed from kernel source

context