skip to content

eBPF

Running verified programs inside the kernel to observe and steer it: hooks and program types, maps, the verifier's safety model, tracing, XDP and tc networking, and the libbpf and bcc toolchains. Interviewers ask because eBPF now sits under modern observability, service meshes, and runtime security, so knowing it means knowing how those products actually work.

on this pageshow

explore

questions

page 1 of 2

In eBPF, what is a map, and how does a user-space program get at the data that a kernel-side eBPF program writes into one?

level: juniorimportance: must knowfreq 78%

answer

  1. shared key/value store living in the kernel
  2. helpers on one side, syscall on the other
  3. the user-space handle is a file descriptor
  4. copies in and out, not shared memory
  5. refcounted — last reference frees it

basics

~20 s

An eBPF map is a kernel-resident key/value store shared by BPF programs and user space. BPF code uses helpers like bpf_map_lookup_elem(); user space copies values in and out via the bpf() syscall on the map's file descriptor.

solid answer

~50 s

A map is the only way an eBPF program can keep state or hand anything back, because the program itself has no memory that outlives one invocation beyond its small stack. It is a kernel object created with the `bpf()` syscall (`BPF_MAP_CREATE`), fixed at creation time in type, `key_size`, `value_size` and `max_entries`, and identified in user space by a file descriptor. Inside the kernel the BPF program touches it only through helpers — `bpf_map_lookup_elem()`, `bpf_map_update_elem()`, `bpf_map_delete_elem()`. User space issues the matching `bpf()` commands on that descriptor, usually through libbpf wrappers, which **copy** the key in and the value out; it is not a pointer into the program's memory. Because the handle is a file descriptor, the map is reference-counted: when the last descriptor and program reference go away the map is freed, unless it has been pinned into the bpffs filesystem to outlive the loader.

code

c · 26 lines
c
#include "vmlinux.h"
#include <bpf/bpf_helpers.h>

struct {
    __uint(type, BPF_MAP_TYPE_HASH);
    __uint(max_entries, 10240);
    __type(key, __u32);
    __type(value, __u64);
} openat_count SEC(".maps");

SEC("tracepoint/syscalls/sys_enter_openat")
int count_openat(void *ctx)
{
    __u32 pid = bpf_get_current_pid_tgid() >> 32;
    __u64 init = 1;

    __u64 *val = bpf_map_lookup_elem(&openat_count, &pid);
    if (!val) {
        bpf_map_update_elem(&openat_count, &pid, &init, BPF_ANY);
        return 0;
    }
    __sync_fetch_and_add(val, 1);
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

go deeper

for a junior

Be able to say plainly that a map is a key/value store in the kernel, that BPF code reaches it with helper calls, and that user space reaches it through a file descriptor.

for a middle

Explain that type, key size, value size and capacity are fixed at BPF_MAP_CREATE, and that user-space lookups copy through the bpf() syscall rather than sharing memory.

for a senior

Show you know the cost model: aggregate in the kernel, poll summaries rather than per-event records, and know that iteration with BPF_MAP_GET_NEXT_KEY is not a consistent snapshot.

for a principal

Own the interface decision — which state lives in maps, which maps are pinned and therefore become a contract between processes, and who is allowed to hold a descriptor to them.

## Why maps exist at all An eBPF program is a short piece of verified bytecode that the kernel runs at a hook point — a kprobe firing, a packet arriving, a tracepoint being hit. It gets a context pointer, a stack of a few hundred bytes, and no ability to write into any process's memory. When it returns, everything local to it is gone. So there has to be somewhere to put results, and somewhere to keep state between invocations. That somewhere is a **map**. A map is a kernel-owned data structure, most commonly a key/value store, that survives across program invocations and is reachable from both the BPF side and user space. Almost every real eBPF tool is structured the same way: a small program in the kernel that filters and aggregates into a map, plus a user-space agent that reads the map and prints, exports, or acts on it. ## The shape is fixed when the map is created Maps are created by the `bpf()` system call with the `BPF_MAP_CREATE` command, which takes: - `map_type` — `BPF_MAP_TYPE_HASH`, `BPF_MAP_TYPE_ARRAY`, `BPF_MAP_TYPE_PERCPU_HASH`, `BPF_MAP_TYPE_LRU_HASH`, `BPF_MAP_TYPE_RINGBUF`, `BPF_MAP_TYPE_PERF_EVENT_ARRAY`, and a few dozen more specialised ones. - `key_size` and `value_size` in bytes. - `max_entries` — the capacity. - flags. All of these are fixed for the life of the map. Keys are exactly `key_size` bytes; there is no variable-length key and no automatic growth when the map fills. In practice you rarely call the syscall by hand — you declare the map in the BPF C source and the loader creates it: ```c struct { __uint(type, BPF_MAP_TYPE_HASH); __uint(max_entries, 10240); __type(key, __u32); __type(value, __u64); } bytes_by_pid SEC(".maps"); ``` The loader creates the map, gets a file descriptor, and patches that descriptor into the program's instructions before loading it, so the verifier knows exactly which map each helper call touches and can check that the key and value sizes the program uses match the map's declaration. ## The kernel side: helpers only BPF code cannot reach into the map's internals. It calls helpers: - `bpf_map_lookup_elem(&map, &key)` returns a pointer to the value **inside the map**, or NULL. The verifier forces a NULL check before you dereference it; because the pointer is into the map itself, writing through it updates the stored value directly. - `bpf_map_update_elem(&map, &key, &value, flags)` inserts or overwrites. Flags are `BPF_ANY`, `BPF_NOEXIST` (fail if present) and `BPF_EXIST` (fail if absent). - `bpf_map_delete_elem(&map, &key)` removes an entry. All of these return an integer status that a careless program ignores — a real source of silent data loss when a map is full. ## The user-space side: it is a file descriptor User space never sees a kernel pointer. It holds a file descriptor and issues `bpf()` commands against it: `BPF_MAP_LOOKUP_ELEM`, `BPF_MAP_UPDATE_ELEM`, `BPF_MAP_DELETE_ELEM`, and `BPF_MAP_GET_NEXT_KEY` for iteration (you walk a hash map by asking for the key after the one you have; passing a NULL/absent key gets the first). Newer kernels add batch commands so you can drain many entries per syscall. libbpf wraps these as `bpf_map_lookup_elem(fd, &key, &value)` and friends, and `bpf_map__fd(skel->maps.bytes_by_pid)` gets the descriptor from a generated skeleton. The important property: **each of these calls copies**. The key goes into the kernel, the value comes back out, one syscall at a time. Polling a large map every second is real overhead, which is why the aggregation belongs in the BPF program and only the summary belongs in the map. The one genuine exception is `BPF_MAP_TYPE_RINGBUF`, whose buffer user space memory-maps and consumes without a syscall per record. ## Lifetime Because the handle is a file descriptor, a map is reference-counted like any other kernel object. References come from open descriptors, from loaded programs that use it, and from bpffs pins. When the last one goes, the map and its contents are freed. That is why a tool that loads a program and exits leaves nothing behind — and why long-lived or shared maps are pinned into the BPF filesystem so a separate process can reopen them later. ## What people get wrong They assume the map is shared memory both sides poke at directly; it is not, outside of the ring buffer. They assume a map grows when it fills; it does not. And they assume the data outlives the loader by default; it does not.

  • If maps are freed when the last reference goes, how would you keep one alive after the loading process exits?
    Pin it into the BPF filesystem, normally mounted at /sys/fs/bpf. A pin is an extra reference held by a filesystem entry, so the map survives the loader's exit and any other privileged process can reopen it by path and keep reading or writing it. Removing the pin drops that reference again.
  • How do you iterate over every entry of a BPF hash map from user space, and what should you be careful about?
    Repeatedly issue BPF_MAP_GET_NEXT_KEY, feeding back the key you just received, and look up each key as you go. It is not a snapshot: the BPF program keeps inserting and deleting while you walk, so you can miss entries or see ones that vanish before the lookup. Batch commands cut the syscall count but not the raciness.
  • Why is it better to aggregate inside the BPF program than to export every event and aggregate in user space?
    Every record you export costs a copy and, for per-key polling, a syscall; a busy hook can fire millions of times a second, so the export path becomes the bottleneck and starts dropping data. Counting or histogramming in the kernel turns millions of events into a handful of map entries that user space reads cheaply.

saying these in an interview costs you the question

  • Thinks a map is shared memory both sides write to directly
  • Says the map grows automatically when max_entries is reached
  • Assumes map data survives the loader process by default
  • Believes BPF programs can write into user-space process memory
  • Thinks keys can be any length rather than exactly key_size

context

open as a page

An eBPF program is loaded into the Linux kernel with a fixed program type such as BPF_PROG_TYPE_KPROBE or BPF_PROG_TYPE_XDP. What does that program type decide about the program, and why can't you attach one program anywhere you like?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An eBPF program's type fixes which kernel hooks it may attach to, the layout of the single context argument it receives, which BPF helper functions it may call, and how the kernel interprets its return value.

open as a page

In Linux eBPF, at what point does the kernel's verifier examine your program, and what happens to a program that fails verification?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The eBPF verifier runs inside the kernel at load time, when the bpf() syscall submits the program — before it is attached to anything and before a single instruction executes. A rejected program never runs: the load call fails and the kernel returns a verifier log saying why.

open as a page

An eBPF tool needs to stream individual event records to a user-space agent. What does BPF_MAP_TYPE_RINGBUF give you that BPF_MAP_TYPE_PERF_EVENT_ARRAY does not, and when would you still reach for the perf buffer?

level: middleimportance: must knowfreq 66%

basics

~20 s

BPF_MAP_TYPE_RINGBUF is one buffer shared by all CPUs: events keep their global order, memory does not scale with CPU count, and bpf_ringbuf_reserve() lets a program fill a record in place. The perf buffer is per-CPU and is the only option below Linux 5.8.

open as a page

In Linux eBPF networking, what does a program attached to the tc (traffic control) hook see that an XDP program does not, and when would you choose tc over XDP?

level: middleimportance: must knowfreq 55%

basics

~20 s

A tc BPF program runs after the kernel has built an sk_buff, so it sees packet metadata, works on every driver, and can be attached on egress as well as ingress. XDP is faster but ingress-only and metadata-free.

open as a page

In Linux XDP (eXpress Data Path), where in the packet receive path does an XDP program run, and what do the return codes XDP_PASS, XDP_DROP, XDP_TX, XDP_REDIRECT and XDP_ABORTED each tell the driver to do?

level: middleimportance: must knowfreq 62%

basics

~20 s

An XDP program runs inside the NIC driver's receive routine, before the kernel allocates an sk_buff. Its return value is a verdict: PASS to the normal stack, DROP the packet, TX back out the same NIC, REDIRECT elsewhere, ABORTED for error drops.

open as a page

How does the Linux kernel physically attach an eBPF program at a kprobe compared with at a tracepoint, and what does that mechanical difference mean for what you can attach to?

level: middleimportance: must knowfreq 62%

basics

~20 s

A kprobe patches the live kernel, replacing the instruction at a symbol's address with a trap or a jump, so it can go almost anywhere. A tracepoint is static instrumentation compiled into the source and exists only where a kernel developer placed one.

open as a page

Two ways to build an eBPF tool on Linux are the bcc framework and libbpf with CO-RE. Where does each one compile the BPF C program, and what does that difference cost you when you ship the tool to a fleet of servers?

level: middleimportance: must knowfreq 62%

basics

~20 s

bcc compiles the BPF C source with an embedded Clang on every host at run time, so LLVM and matching kernel headers must be installed there. libbpf with CO-RE compiles once into a portable binary and relocates struct field offsets at load time using the kernel's own BTF.

open as a page

You need a latency distribution for block I/O on a busy production server, over ten minutes at tens of thousands of I/Os per second. Explain how an eBPF tracing tool such as bpftrace or bcc's biolatency builds that histogram without shipping one record per I/O to user space, and why that design is what makes it safe to run on a live box.

level: middleimportance: must knowfreq 45%

basics

~20 s

The BPF program does the maths in the kernel. It stores a start timestamp per in-flight request, computes the delta on completion, and increments a bucket in a kernel-resident map. User space reads that map once at the end, so the cost per event is a few lookups rather than a record pushed through a ring buffer.

open as a page

An eBPF program calls bpf_map_lookup_elem() and dereferences the returned pointer straight away, and the kernel refuses to load it with an invalid memory access mentioning map_value_or_null. What is the verifier objecting to, and what does it want the code to do?

level: middleimportance: must knowfreq 62%

basics

~20 s

bpf_map_lookup_elem() returns NULL when the key is absent, so the verifier types its result as a maybe-NULL pointer and refuses any dereference of it. An explicit NULL check on the returned pointer refines the type on the surviving branch, after which the dereference is allowed.

open as a page

A libbpf-based tool loads and attaches an eBPF program on a Linux host, then the user-space process exits and the program vanishes from `bpftool prog show`. What owns a loaded BPF object's lifetime, and how do you make one outlive the process that loaded it?

level: seniorimportance: must knowfreq 44%

basics

~20 s

Loaded BPF programs, maps and links are refcounted kernel objects with no name in any namespace; the loader's file descriptors are usually the only reference, so closing them at process exit frees everything. Pinning to the bpffs filesystem at /sys/fs/bpf takes an extra reference that survives the process.

open as a page

You are handed a Linux box with bpftrace installed and asked to watch which files a service opens. Walk through what the one-liner `bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s\n", comm, str(args->filename)); }'` actually does, and explain how you would have discovered that probe and its argument fields in the first place.

level: juniorimportance: should knowfreq 40%

basics

~20 s

The one-liner attaches a small BPF program to the openat syscall-entry tracepoint and prints the calling process name and the filename for every openat on the whole machine. bpftrace -l lists available probes and bpftrace -lv shows a tracepoint's argument fields.

open as a page

In eBPF, what is a helper function, and why can a BPF program call bpf_map_lookup_elem() but not an arbitrary kernel function such as kmalloc()?

level: middleimportance: should knowfreq 55%

basics

~20 s

BPF helpers are a fixed set of kernel functions, each with a numbered ID and a prototype the verifier knows, that BPF programs are allowed to call. Arbitrary kernel functions like kmalloc() have no such contract and cannot be proven safe.

open as a page

You have SSH'd into a Linux host you did not set up and want to know which eBPF programs are loaded and what they are attached to. Which bpftool subcommands answer that, and what do their outputs tell you?

level: middleimportance: should knowfreq 45%

basics

~20 s

Start with bpftool prog show for loaded programs and bpftool map show for their maps. Attachment points come from bpftool link show, bpftool net show for interface hooks and bpftool cgroup tree for cgroup hooks, because prog show alone does not say where a program is attached.

open as a page

A bpftrace script that attaches with a `kprobe:` on a kernel function stops attaching after your fleet is upgraded to a newer kernel, while a colleague's script using a `tracepoint:` probe for roughly the same event keeps working. Why do the two behave differently across kernel versions, and what does that mean for a tracing tool you have to ship to many machines?

level: middleimportance: should knowfreq 36%

basics

~20 s

A kprobe attaches to an internal kernel function name with no stability promise, so a rename, an inline or a signature change breaks it on the next kernel. A tracepoint is a declared instrumentation point with named fields that kernel developers try not to break, so it survives upgrades far better.

open as a page

Older Linux kernels rejected any eBPF program containing a backward jump. What does the verifier require of a loop today, and what problem does the bpf_loop() helper solve?

level: middleimportance: should knowfreq 45%

basics

~20 s

The verifier must prove every eBPF program terminates. Originally it demanded a loop-free control-flow graph, so loops had to be unrolled; since Linux 5.3 it accepts loops whose iteration count it can bound, and bpf_loop() moves the iteration to runtime so verification cost stops scaling with the trip count.

open as a page

The verifier rejects an eBPF program because its stack usage is too large. How much stack does an eBPF program get on Linux, why is the limit so small, and where should a large scratch buffer live instead?

level: middleimportance: should knowfreq 40%

basics

~20 s

An eBPF program gets 512 bytes of stack, and with subprograms it is the combined depth of the call chain that must fit. The limit is small because the program runs on the ordinary kernel stack in any context. Large buffers belong in a per-CPU array map used as scratch space.

open as a page

An eBPF program counts events by looking a key up in a BPF_MAP_TYPE_HASH and doing (*val)++, and under load the totals come out too low. What is wrong, and what do __sync_fetch_and_add() and BPF_MAP_TYPE_PERCPU_HASH each change about it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Incrementing a value fetched from a shared map is a non-atomic read-modify-write, so CPUs running the hook concurrently overwrite each other's counts. __sync_fetch_and_add() makes the update atomic; a per-CPU map gives each CPU its own value, and user space sums across CPUs.

open as a page

In an XDP-based L4 load balancer, what is the difference between returning XDP_TX and returning XDP_REDIRECT for a forwarded packet, and why do such designs typically encapsulate the packet and use direct server return rather than plain NAT?

level: seniorimportance: should knowfreq 33%

basics

~20 s

XDP_TX resends the packet out the interface it arrived on; XDP_REDIRECT hands it to a different target through a redirect map. Encapsulation with direct server return lets backends reply straight to the client, so only the request half crosses the load balancer.

open as a page

You attach an XDP drop program to a 25 GbE Linux server's interface. The filter works, but throughput and CPU use are no better than the iptables rule it replaced, and `ip link show` reports the attachment as xdpgeneric. What are XDP's three attach modes, and what has happened here?

level: seniorimportance: should knowfreq 40%

basics

~20 s

XDP attaches in native mode inside the driver, generic mode in the shared kernel receive path after the sk_buff is allocated, or offloaded onto a capable NIC. Generic mode is a correctness fallback with none of the performance benefit — the driver did not support native XDP.

open as a page

A colleague loaded an XDP program onto a production interface from a terminal that has since been closed, and traffic is still being affected. Why do some eBPF attachments survive the process that created them while others vanish with it, and how would you find and remove this one?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Attachments made through a bpf_link end when the last file descriptor to that link closes, while netlink XDP attachments, tc filters and cgroup attachments hold their own kernel reference and outlive the loader. Those have to be detached explicitly.

open as a page

In eBPF, fentry and fexit programs (BPF_PROG_TYPE_TRACING) can target many of the same kernel functions as kprobe and kretprobe programs. How do they attach differently, and what does an fexit program see that a kretprobe cannot?

level: seniorimportance: should knowfreq 40%

basics

~20 s

fentry and fexit programs attach through a BPF trampoline at the compiler-inserted patch site and read typed arguments derived from BTF. An fexit program sees the function's arguments and its return value together; a kretprobe sees only the return.

open as a page

A libbpf CO-RE binary that works on a current Ubuntu server fails to load on an older Linux host, and `/sys/kernel/btf/vmlinux` does not exist there. What is missing, and what are your options for that host?

level: seniorimportance: should knowfreq 38%

basics

~20 s

That kernel was built without CONFIG_DEBUG_INFO_BTF, so it exposes no BTF and libbpf has nothing to resolve the binary's CO-RE relocations against. Options are a kernel built with BTF enabled, supplying an externally generated BTF for that exact kernel, or falling back to a runtime-compiled bcc tool.

open as a page

Your bpftrace one-liner runs fine as root on a Linux host but fails to load when you run it inside a container, and when you run it from the host against a containerised workload the PIDs and file paths it prints do not match what you see inside the container. Explain both problems and how you would work around each.

level: seniorimportance: should knowfreq 28%

basics

~20 s

Inside a container the loading fails because the bpf and perf_event_open system calls are usually blocked by the default seccomp profile and the required capabilities are dropped. From the host, the PIDs printed are the host's, and the file paths belong to the container's mount namespace, so neither matches what the container shows.

open as a page

A large eBPF program is rejected with the verifier reporting that it processed too many instructions, even though the program itself is far shorter than the instruction limit. What is the verifier counting, and how do you restructure the program so it loads?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The verifier counts instructions it simulates across every reachable path, not instructions in the program, so branches and loops multiply the work. Reduce path explosion: cut branches, help its bounds tracking, move iteration into bpf_loop(), and split logic into independently verified global subprograms or tail calls.

open as a page

You need packet filtering and load balancing at close to line rate on commodity Linux servers. How would you decide between an XDP/eBPF datapath and a kernel-bypass framework such as DPDK, and what does each approach cost you?

level: principalimportance: should knowfreq 42%

basics

~20 s

DPDK takes the NIC away from the kernel for maximum throughput, at the cost of the whole kernel network stack, its tooling, and dedicated busy-polling cores. XDP stays inside the kernel, slightly slower but selective, interrupt-driven and operable with normal tools. Choose by how much of the traffic is exceptional.

open as a page

You own the eBPF-based tracing tools for a fleet of Linux servers spanning several kernel versions and distributions. How would you decide between shipping bcc-based tools and libbpf CO-RE binaries, and what would you standardise in the build and distribution pipeline?

level: principalimportance: should knowfreq 26%

basics

~20 s

Standardise on libbpf CO-RE binaries for anything shipped, because they need no compiler or headers on production hosts, and keep bcc for ad-hoc work on lab machines. Then fix a minimum kernel and BTF baseline, a pinned build toolchain, and a load-test matrix covering every kernel in the fleet.

open as a page

Your XDP program on a Linux host is dropping flood packets with XDP_DROP, but `tcpdump -i eth0` on that host shows none of those packets at all. Why can tcpdump not see them, and where does its own filter actually run?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

tcpdump reads packets from an AF_PACKET socket that is fed later in the receive path, after the driver has built an sk_buff. Native XDP runs before that, so packets it drops never reach the tap. Count drops in a BPF map instead.

open as a page

In a libbpf project on Linux, what does `bpftool gen skeleton prog.bpf.o > prog.skel.h` produce, and how does the generated skeleton change the user-space loader code?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

It generates a C header that embeds the compiled BPF object and declares a struct named after it, with open, load, attach and destroy functions plus typed members for every map, program, link and global variable — replacing hand-written libbpf boilerplate and file loading at runtime.

open as a page

A long-running eBPF tool keys a BPF_MAP_TYPE_HASH by connection and never observes some of those connections close, so entries accumulate. What happens once the map reaches max_entries, and what does switching to BPF_MAP_TYPE_LRU_HASH change?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

A BPF hash map has a fixed max_entries; once full, inserting a new key fails with -E2BIG and, if the program ignores the return value, the event is silently dropped. BPF_MAP_TYPE_LRU_HASH instead evicts an approximately least-recently-used entry so inserts keep succeeding.

open as a page

showing 1–30 of 33