A bpftrace script that attaches with a `kprobe:` on a kernel function stops attaching after your fleet is upgraded to a newer kernel, while a colleague's script using a `tracepoint:` probe for roughly the same event keeps working. Why do the two behave differently across kernel versions, and what does that mean for a tracing tool you have to ship to many machines?
answer
- one is declared, the other is interposed
- internal symbols carry no promise
- the quiet failure is worse than the loud one
- offsets are a separate problem from names
- check the probe exists before you rely on it
basics
~20 sA kprobe attaches to an internal kernel function name with no stability promise, so a rename, an inline or a signature change breaks it on the next kernel. A tracepoint is a declared instrumentation point with named fields that kernel developers try not to break, so it survives upgrades far better.
solid answer
~60 sA `kprobe` is dynamic instrumentation: it hooks whatever symbol you named, and kernel internals are explicitly not a stable interface. Between releases a function can be renamed, split, made `static` and inlined away, or have its arguments reordered. If the symbol is gone the probe simply fails to attach; worse, if it still exists but the arguments moved, the probe attaches and prints nonsense. A `tracepoint` is a static marker the kernel developers placed deliberately, with a named field layout exported through tracefs, and there is real reluctance to break the widely used ones. That is the first half of portability. The second half is struct fields: a program that reads members out of kernel structures depends on offsets that change between builds, which is what BTF and CO-RE solve — libbpf relocates the offsets at load time against the running kernel's own type information. So the rule for a fleet tool is: prefer a tracepoint where one exists, treat kprobes as version-specific, verify with `bpftrace -l` before attaching, and provide a fallback.
go deeper
Know that kprobes hook internal kernel functions while tracepoints are deliberate, named instrumentation points, and that internal function names can change between kernel releases.
Explain all three kprobe failure modes — symbol gone, function inlined, arguments reordered — and say which one is silent. Be able to describe what BTF and CO-RE fix and what they do not.
Show how you would ship such a tool to a fleet: probe-existence checks before attach, fallback chains, validation as part of kernel upgrades, and sanity checks on argument values so a signature change cannot mislead an incident.
Own the policy question — whether kernel-version-coupled tooling belongs in your standard incident toolkit at all, how tracing coverage is validated ahead of a fleet-wide kernel rollout, and who owns that regression surface.
## Two very different kinds of attachment point **A kprobe is dynamic.** The kernel lets you interpose on essentially any function that appears in the symbol table, by patching the instruction at that address. Nobody agreed in advance that you would do this. The set of functions and their signatures is an *internal implementation detail* of the kernel, and Linux offers no stability guarantee for internal interfaces — only for the user-space ABI (system calls, and the shape of the things they return). **A tracepoint is static.** A kernel developer wrote a `TRACE_EVENT` declaration at that spot, naming the event and the fields it publishes. The field layout is exported at runtime under tracefs, which is how `bpftrace -lv` can tell you the names. Tracepoints are not a formal ABI either, but they are used widely enough — by `perf`, by `ftrace`, by every tracing tool in existence — that changing a popular one draws immediate objection, so in practice they are far more durable. ## The three ways a kprobe breaks 1. **The symbol disappears.** A refactor renames or removes the function. `bpftrace` reports that there were no probes to attach and exits. This is the *good* failure: loud and immediate. 2. **The function is inlined.** A `static` function that the compiler chose to inline has no independent entry point and therefore no address to patch. Whether this happens can vary with compiler version and kernel config, so the same source can be probeable on one distribution's build and not on another's — the same kernel version, different result. 3. **The signature changes.** The name survives but arguments are added, reordered or retyped. The probe attaches happily and `arg0` now means something else. This is the dangerous failure: no error, wrong data, and a conclusion drawn from it. Real tools have lived through all three. Block-layer tooling in particular has had to change which functions it attaches to as the kernel's I/O accounting was reworked and functions were renamed or inlined — one of the standard cautionary examples for kprobe-based tools. ## Where tracepoints stop helping Tracepoints are not everywhere. There is no tracepoint at every interesting function, and where none exists a kprobe is the only option — which is exactly why kprobes are indispensable rather than merely legacy. Tracepoints can also be added over time, so a script that uses one may fail on *older* kernels rather than newer ones, and the `syscalls:` tracepoints depend on a kernel config option being enabled. ## The second half of portability: struct offsets Attaching is one problem; reading data is another. A program that walks a kernel structure — say, reaching through the current task to pull out a field — depends on the byte offset of that member. Offsets change between kernel versions and even between two builds of the same version with different configs. The old bcc answer was to compile the program *on the target machine* at runtime, against that machine's kernel headers. That works but demands a compiler and matching headers on every production host, and costs seconds and hundreds of megabytes of memory at each tool start. The modern answer is **BTF plus CO-RE** (Compile Once – Run Everywhere). The kernel ships a description of its own types, and a program compiled once carries relocation records saying "this access is to field X of struct Y". At load time the loader looks up the running kernel's BTF and rewrites each access to the correct offset for *this* kernel. A single binary then runs across many kernel versions with no compiler on the host. It fixes offsets; it does not conjure a field that no longer exists, and it does not make a removed function reappear. ## What this means for a fleet tool - **Prefer a tracepoint when one covers the event.** Accept slightly less precise placement in exchange for surviving upgrades. - **Treat kprobe scripts as version-pinned artefacts.** Record which kernels they were validated against, and re-validate as part of the kernel upgrade process rather than at 3 a.m. - **Check before you attach.** `bpftrace -l 'kprobe:some_func'` is a cheap existence test; empty output means the probe is not on this host. Bake that check into whatever wraps the tool. - **Provide a fallback chain.** Attach to the preferred point, and if it is absent fall back to an older name or a coarser tracepoint, rather than failing outright. - **Do not silently trust arguments.** If a kprobe's arguments matter, sanity-check the values — an implausible pointer or a nonsense size is the signature of a changed signature. Worth knowing as an aside: on modern kernels there is also a lower-overhead attachment for kernel functions built on BPF trampolines and BTF (`fentry`/`fexit`), which is type-aware rather than positional. It removes the argument-marshalling guesswork but does not remove the underlying issue: it still targets an internal function that the kernel may rename or inline.
- Which of the two failure modes — probe fails to attach, or probe attaches with changed arguments — is worse, and why?The silent one. A failed attach stops the tool and you fix it. A changed signature leaves the probe working and reporting plausible-looking numbers that mean something else, and those numbers then drive an incident decision. It's why kprobe-based tools should sanity-check argument values and why type-aware attachment is an improvement over positional arguments.
- Does BTF and CO-RE make a kprobe-based tool portable across kernels?Only partly. CO-RE fixes struct field offsets, so reading members of kernel structures works across builds without on-host compilation. It does nothing about the attachment point: if the function was renamed, inlined or removed, the probe still fails to attach, and if its arguments were reordered you still read the wrong ones.
- Why can a function be probeable on one distribution's kernel and not on another build of the same version?Because inlining is a compiler decision influenced by the kernel configuration, compiler version and optimisation flags. A `static` function that the compiler folded into its caller has no independent entry point in the symbol table, so there is nothing to patch. Same source, different build, different set of attachable symbols.
saying these in an interview costs you the question
- Believes kernel function names are a stable interface
- Assumes a probe that attaches is therefore reading the right arguments
- Thinks tracepoints exist for every interesting kernel function
- Says CO-RE makes any tracing script kernel-independent
- Blames the tracing tool rather than the missing symbol on upgrade