skip to content

Tracing and Observability

Using eBPF for observability: kprobes and uprobes, tracepoints, BTF and CO-RE for portability, bpftrace one-liners, the bcc tool collection, and reading kernel structures safely across kernel versions. Interviewers ask because a bpftrace one-liner can answer a latency question in seconds that logging cannot answer at all.

on this pageshow

questions

5

You need a latency distribution for block I/O on a busy production server, over ten minutes at tens of thousands of I/Os per second. Explain how an eBPF tracing tool such as bpftrace or bcc's biolatency builds that histogram without shipping one record per I/O to user space, and why that design is what makes it safe to run on a live box.

level: middleimportance: must knowfreq 45%

answer

  1. the maths happens where the event happens
  2. two probes and a timestamp in between
  3. buckets, not samples, leave the kernel
  4. output size independent of event rate
  5. the start map is the thing that leaks

basics

~20 s

The BPF program does the maths in the kernel. It stores a start timestamp per in-flight request, computes the delta on completion, and increments a bucket in a kernel-resident map. User space reads that map once at the end, so the cost per event is a few lookups rather than a record pushed through a ring buffer.

solid answer

~60 s

The tool attaches two probes: one where the request is issued, which stores `nsecs` in a map keyed by the request, and one where it completes, which subtracts to get the latency. The crucial part is what happens next — instead of emitting the sample, the program computes a log2 bucket and increments a counter in a BPF map that lives in the kernel. Per event you pay a couple of map operations, on the order of a microsecond, with no copy to user space and no context switch. The user-space process is idle the whole time; it only reads the map and formats the histogram when you stop the tool. That is the whole reason eBPF tracing is production-viable at high event rates: the data volume leaving the kernel is proportional to the number of *buckets*, not the number of events. Per-event output is still possible — bcc's biosnoop prints a line per I/O — but it goes through a ring buffer that can drop records under load, so you use it for low-rate events or short windows.

code

bash · 6 lines
bash
# Time an operation across two probes and bucket the result in the kernel
bpftrace -e 'kprobe:vfs_read { @start[tid] = nsecs; }
kretprobe:vfs_read /@start[tid]/ {
  @us = hist((nsecs - @start[tid]) / 1000);
  delete(@start[tid]);
}'

go deeper

for a junior

Know that the BPF program runs inside the kernel at the moment of the event, and that tools like biolatency summarise there rather than printing every event. Say plainly that this is why they can run on a live server.

for a middle

Explain the two-probe timestamp pattern and the map that accumulates buckets, and be able to state that the data leaving the kernel scales with bucket count rather than event rate. Mention deleting the start entry.

for a senior

Reason about overhead as per-event cost times event rate, and show the aggregate-then-stream workflow: histogram to prove a tail exists, filtered per-event trace to attribute it. Call out map exhaustion as a silent under-count.

for a principal

Own the question of what may be attached to a hot path in production at all — who is allowed to run it, for how long, with what blast radius — and where a hand-run trace ends and permanent instrumentation should begin.

## The design that does not work The obvious way to measure latency is to emit a record per event and do the arithmetic later: timestamp, device, duration, one line each. At tens of thousands of events per second this fails on three counts. Each record must be copied out of the kernel; a user-space reader must be scheduled to drain it; and something must store or parse the stream. The measurement starts to cost more than the work being measured, and worse, the cost lands on the very machine that is already in trouble. This is the structural reason older per-event tracing approaches were reserved for a lab. ## What eBPF does instead The defining property of eBPF observability is that **the reduction happens in the kernel, at the event, before anything is copied out**. A histogram tool is built from two probe blocks plus a map: 1. **Start probe.** On the issue path, record the current time: `@start[key] = nsecs;`. The key identifies the in-flight operation — a thread id for a syscall, a device plus sector for a block request. 2. **End probe.** On the completion path, look the key up; if present, compute `nsecs - @start[key]`, feed it into the histogram, and delete the entry so the map does not grow without bound. In bpftrace that is roughly: ``` kprobe:vfs_read { @start[tid] = nsecs; } kretprobe:vfs_read /@start[tid]/ { @us = hist((nsecs - @start[tid]) / 1000); delete(@start[tid]); } ``` `hist()` is not a function that returns data to you — it is an *aggregation* that lives in a kernel map. Each call computes which power-of-two bucket the value falls into and increments that bucket's counter. bcc's `biolatency` does exactly the same thing against the block layer. ## Why this is cheap Per event you pay: one map insert, one map lookup, one delete, a subtraction, and a bucket increment. That is a small number of hash operations, typically well under a microsecond, executed inline on the thread that was going to run anyway. Nothing is written to a pipe, nothing wakes a user-space process, no context switch occurs. The amount of data that ever crosses into user space is bounded by the *shape* of the aggregation, not by the event rate. A log2 latency histogram has a few dozen buckets whether it saw a thousand events or a billion. Ten minutes at 30,000 events per second is 18 million samples and still one small map read at the end. A second reason it is safe: the program was verified before it loaded. It cannot loop unboundedly, cannot dereference an unchecked pointer, and cannot be attached in a way that lets it panic the kernel. The failure mode of a bad tracing program is a load-time rejection, not a crashed production host. ## The counters that must be maintained The start map is the part that bites people. If the completion probe never fires for some requests — the operation errors out on a path you did not instrument, or the tool starts mid-flight — entries accumulate. BPF maps have a fixed maximum number of entries fixed at load time, so a leaking start map eventually stops accepting inserts and you silently under-count. Always delete on the completion path, and be suspicious of a histogram whose sample count is far below the event rate you expect. ## When you genuinely need per-event output Aggregation answers "what is the distribution"; it cannot answer "which request was the slow one, and what was running at the time". For that you do stream events — bcc's `biosnoop` prints one line per I/O with the process, device, sector and latency, and bpftrace's `printf()` does the same. These push records through a per-CPU ring buffer that user space drains. The buffer is finite, so under load records are dropped and the tool tells you it lost events. That is an acceptable trade for a short, targeted window or a low-rate event (process execs, for instance), and a bad one for every block I/O on a busy database host. The practical workflow is therefore two-stage: aggregate first to find out *whether* there is a tail and where it sits, then stream a narrow, filtered slice — one device, one process, one latency threshold — to find out *what* is in it. Filtering in the probe's predicate keeps the second stage cheap, because the discard also happens in the kernel. ## What to say about overhead Be concrete rather than absolute. Overhead is roughly the per-event cost times the event rate, so the question is always "how hot is this probe". Attaching to something that fires a million times a second is expensive even at a microsecond each; attaching to process exec is free in practice. Tracing tools that let you filter early and aggregate in the kernel let you keep that product small, which is the actual skill.

  • What goes wrong if the completion probe never fires for some of the operations you timed?
    The entries keyed on those operations stay in the start map forever. BPF maps have a fixed maximum entry count set at load time, so a leaking map eventually refuses new inserts and the tool quietly stops recording new samples. The symptoms are a histogram with far fewer samples than the known event rate. You avoid it by deleting on completion and by instrumenting every path that ends a request.
  • When would you prefer streaming one record per event over an in-kernel histogram?
    When you need attribution rather than distribution — which process, which file, which sector was slow — or when the event is rare enough that the volume is trivial, such as process execs or TCP connects. It's the right tool for a short, filtered window after aggregation has already told you a tail exists; not for a permanently attached probe on a hot path.
  • Why is a log2 histogram used rather than storing the raw samples and computing percentiles exactly?
    Because raw samples grow with the event rate and buckets do not. A power-of-two histogram gives constant memory and constant readout size for any volume, at the price of resolution: you learn a latency was between 8 and 16 milliseconds, not that it was 11. For an incident that resolution is almost always enough, and bpftrace's linear histogram is available when you need finer bands over a known range.

saying these in an interview costs you the question

  • Thinks eBPF is cheap because the programs are small
  • Assumes every event is copied to user space and filtered there
  • Claims eBPF tracing has zero overhead on any probe
  • Ignores deleting start-map entries and calls the map unbounded
  • Treats a per-event streaming tool as safe on a hot kernel path

context

open as a page

You are handed a Linux box with bpftrace installed and asked to watch which files a service opens. Walk through what the one-liner `bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s\n", comm, str(args->filename)); }'` actually does, and explain how you would have discovered that probe and its argument fields in the first place.

level: juniorimportance: should knowfreq 40%

basics

~20 s

The one-liner attaches a small BPF program to the openat syscall-entry tracepoint and prints the calling process name and the filename for every openat on the whole machine. bpftrace -l lists available probes and bpftrace -lv shows a tracepoint's argument fields.

open as a page

A bpftrace script that attaches with a `kprobe:` on a kernel function stops attaching after your fleet is upgraded to a newer kernel, while a colleague's script using a `tracepoint:` probe for roughly the same event keeps working. Why do the two behave differently across kernel versions, and what does that mean for a tracing tool you have to ship to many machines?

level: middleimportance: should knowfreq 36%

basics

~20 s

A kprobe attaches to an internal kernel function name with no stability promise, so a rename, an inline or a signature change breaks it on the next kernel. A tracepoint is a declared instrumentation point with named fields that kernel developers try not to break, so it survives upgrades far better.

open as a page

Your bpftrace one-liner runs fine as root on a Linux host but fails to load when you run it inside a container, and when you run it from the host against a containerised workload the PIDs and file paths it prints do not match what you see inside the container. Explain both problems and how you would work around each.

level: seniorimportance: should knowfreq 28%

basics

~20 s

Inside a container the loading fails because the bpf and perf_event_open system calls are usually blocked by the default seccomp profile and the required capabilities are dropped. From the host, the PIDs printed are the host's, and the file paths belong to the container's mount namespace, so neither matches what the container shows.

open as a page

You want to time a function inside a running user-space application using bpftrace's `uprobe` and `uretprobe` probes. How does a uprobe actually attach to the process, roughly what does each hit cost compared with a kernel tracepoint, and which kinds of binary make this approach unreliable?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

A uprobe patches a breakpoint instruction into a private copy of the target's executable page, so every thread that reaches that address traps into the kernel and runs the BPF program. Each hit costs roughly a microsecond, far more than a tracepoint. Stripped, statically linked, Go and JIT-compiled binaries make it unreliable.

open as a page