In eBPF, fentry and fexit programs (BPF_PROG_TYPE_TRACING) can target many of the same kernel functions as kprobe and kretprobe programs. How do they attach differently, and what does an fexit program see that a kretprobe cannot?
answer
- trampoline at the ftrace patch site
- BTF gives the verifier the prototype
- exit program sees args and return together
- kretprobe needs a map and a key
- preallocated return instances can run out
basics
~20 sfentry and fexit programs attach through a BPF trampoline at the compiler-inserted patch site and read typed arguments derived from BTF. An fexit program sees the function's arguments and its return value together; a kretprobe sees only the return.
solid answer
~50 sfentry/fexit are `BPF_PROG_TYPE_TRACING` programs attached via a BPF trampoline installed at the function's compiler-inserted `fentry` patch site, rather than by planting a breakpoint the way a classic kprobe does — so the call path is a jump into generated code and the overhead is lower. Because the kernel has BTF, the verifier knows the target's real prototype, so your program takes typed arguments instead of decoding `struct pt_regs`. The big practical win is `fexit`: it runs after the function returns and has both the original arguments and the return value in one program. With a `kretprobe` you only get the return value, so you must stash the arguments in a map from a paired kprobe and correlate them by thread — extra map operations, stale entries when the return never happens, and missed events once the number of in-flight calls exceeds the return probe's preallocated instances.
code
c · 12 lines#include "vmlinux.h"
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_tracing.h>
SEC("fexit/do_unlinkat")
int BPF_PROG(unlink_exit, int dfd, struct filename *name, long ret)
{
bpf_printk("do_unlinkat returned %ld", ret);
return 0;
}
char LICENSE[] SEC("license") = "GPL";go deeper
Know that fentry and fexit are newer, cheaper attach points for kernel functions, and that fexit can see the return value alongside the arguments. Naming that difference is enough at this level.
Explain the mechanism: a BPF trampoline at the compiler-inserted patch site rather than a breakpoint, and typed arguments made possible by kernel BTF. Contrast with decoding struct pt_regs by hand.
Demonstrate that you have written the kprobe/kretprobe correlation pattern and felt its costs — the per-thread map, the leaked entries, the misses from an exhausted return-instance pool — and can state the kernel and BTF conditions under which you would still fall back to it.
Own the fleet-level decision: which tracing attach mechanism your agents standardise on given the oldest kernel you support, whether you carry both code paths, and what you accept losing when a host has no BTF.
## The same targets, a different mechanism `fentry` and `fexit` programs are loaded as `BPF_PROG_TYPE_TRACING` with an expected attach type of `BPF_TRACE_FENTRY` or `BPF_TRACE_FEXIT`. They aim at kernel functions, just as kprobes do, but almost everything about how they get there is different. ## The BPF trampoline The kernel is compiled so that most functions begin with a patchable nop sequence at the very top — the site the ftrace machinery has always used. Attaching an fentry program installs a **BPF trampoline** at that site: generated machine code that saves the arguments in a known layout, calls your program, and continues into the real function. For `fexit`, the same trampoline also captures the return value and calls your program on the way out. Contrast that with a classic kprobe, whose baseline mechanism is a breakpoint instruction and a trap. Kprobes can be optimised into jumps, and the gap narrows when they are, but the trampoline path is designed for this from the start and is the cheaper of the two in practice. ## Typed arguments instead of registers A kprobe program's context is `struct pt_regs`, and you pull arguments out of it with architecture-specific macros. If you get the target's signature wrong, nothing complains — you silently read the wrong register. fentry/fexit require the kernel to carry BTF (built with `CONFIG_DEBUG_INFO_BTF`), which means the verifier knows the target function's prototype. Your program declares the arguments by name and type, and the verifier checks you against the real signature: attach to a function whose prototype does not match what you declared and the load fails rather than producing garbage. ```c SEC("fexit/do_unlinkat") int BPF_PROG(unlink_exit, int dfd, struct filename *name, long ret) { bpf_printk("do_unlinkat -> %ld", ret); return 0; } ``` The trailing `ret` parameter is the whole point of `fexit`. ## Why the entry/exit correlation problem disappears The classic tracing pattern is "time this function" or "log the arguments of calls that failed". With kprobes you need two programs: - a **kprobe** at entry, which sees the arguments and writes them into a hash map keyed by the thread id; - a **kretprobe** at return, which sees only the return value, looks the entry up in the map, emits the pair, and deletes the entry. That design has three real costs. It doubles the probe overhead and adds two map operations per call. It leaks map entries whenever a call never returns down the traced path — the entry sits there until something evicts it, which is why these tools use bounded or LRU maps. And it is lossy: a kretprobe works by replacing the return address, which requires an instance from a preallocated pool sized by `maxactive`. When more calls are in flight than there are instances, return probes are simply missed, and the miss count shows up in the kprobes listing in tracefs rather than anywhere your tool will notice. An `fexit` program collapses all of that into one program with the arguments and the return value in scope at once. No map, no correlation key, no misses from an exhausted instance pool. ## What you give up fentry/fexit are not a strict superset: - **Kernel and architecture support.** They arrived in Linux 5.5 and are supported on the main architectures, but on an old kernel or one built without BTF they are simply unavailable — which is why portable tooling still ships kprobe fallbacks. - **Function granularity only.** A kprobe can attach at a symbol *plus an offset*, anywhere inside a function. fentry attaches at the function's entry patch site, full stop. - **Untouchable functions.** Functions marked `notrace` or `noinstr`, and those the compiler inlined, have no patch site — the same fence kprobes hit, for the same reason. ## The interview framing When asked to choose, the answer is: on a modern kernel with BTF, reach for fentry/fexit first, because the arguments are typed and checked, the overhead is lower, and fexit removes an entire class of correlation bugs. Fall back to kprobe/kretprobe when the kernel is old, when BTF is absent, or when you need to probe at an offset inside a function rather than at its entry. And be able to say *why* the kretprobe pattern is lossy — that detail is the one that separates people who have read about eBPF from people who have shipped it.
- Your fentry attach fails on a kernel where the equivalent kprobe attaches fine. What would you check?Whether the kernel carries BTF at all — `CONFIG_DEBUG_INFO_BTF`, in practice the presence of `/sys/kernel/btf/vmlinux` — because without it the verifier cannot type-check the attach. Then whether the kernel is older than 5.5, and whether the target is `notrace`/`noinstr` or lives in a module whose BTF is not loaded. A declared prototype that does not match the real one also fails at load.
- Why is the kprobe plus kretprobe pattern lossy, and how would you detect the loss?A return probe consumes an instance from a pool sized by `maxactive`; when more calls are in flight than there are instances, the return is not recorded at all. The missed count is exposed in the kprobes listing under tracefs, not by your tool, so you have to go look. Raising `maxactive` reduces it; moving to `fexit` removes the mechanism entirely.
- Is there any tracing job fentry cannot do that a kprobe can?Yes — probing at an offset inside a function. A kprobe can attach at `symbol+offset`, which is how you instrument a specific branch or loop within a large function. fentry attaches only at the function's entry patch site, and fexit at its return, so anything mid-function still needs a kprobe.
saying these in an interview costs you the question
- Claims fexit is just a faster kretprobe with the same data
- Thinks fentry works on any kernel regardless of BTF
- Says kretprobes never miss events
- Believes fentry can attach at a symbol plus offset
- Treats pt_regs decoding as equivalent to typed arguments