In eBPF, what is a map, and how does a user-space program get at the data that a kernel-side eBPF program writes into one?
answer
- shared key/value store living in the kernel
- helpers on one side, syscall on the other
- the user-space handle is a file descriptor
- copies in and out, not shared memory
- refcounted — last reference frees it
basics
~20 sAn eBPF map is a kernel-resident key/value store shared by BPF programs and user space. BPF code uses helpers like bpf_map_lookup_elem(); user space copies values in and out via the bpf() syscall on the map's file descriptor.
solid answer
~50 sA map is the only way an eBPF program can keep state or hand anything back, because the program itself has no memory that outlives one invocation beyond its small stack. It is a kernel object created with the `bpf()` syscall (`BPF_MAP_CREATE`), fixed at creation time in type, `key_size`, `value_size` and `max_entries`, and identified in user space by a file descriptor. Inside the kernel the BPF program touches it only through helpers — `bpf_map_lookup_elem()`, `bpf_map_update_elem()`, `bpf_map_delete_elem()`. User space issues the matching `bpf()` commands on that descriptor, usually through libbpf wrappers, which **copy** the key in and the value out; it is not a pointer into the program's memory. Because the handle is a file descriptor, the map is reference-counted: when the last descriptor and program reference go away the map is freed, unless it has been pinned into the bpffs filesystem to outlive the loader.
code
c · 26 lines#include "vmlinux.h"
#include <bpf/bpf_helpers.h>
struct {
__uint(type, BPF_MAP_TYPE_HASH);
__uint(max_entries, 10240);
__type(key, __u32);
__type(value, __u64);
} openat_count SEC(".maps");
SEC("tracepoint/syscalls/sys_enter_openat")
int count_openat(void *ctx)
{
__u32 pid = bpf_get_current_pid_tgid() >> 32;
__u64 init = 1;
__u64 *val = bpf_map_lookup_elem(&openat_count, &pid);
if (!val) {
bpf_map_update_elem(&openat_count, &pid, &init, BPF_ANY);
return 0;
}
__sync_fetch_and_add(val, 1);
return 0;
}
char LICENSE[] SEC("license") = "GPL";go deeper
Be able to say plainly that a map is a key/value store in the kernel, that BPF code reaches it with helper calls, and that user space reaches it through a file descriptor.
Explain that type, key size, value size and capacity are fixed at BPF_MAP_CREATE, and that user-space lookups copy through the bpf() syscall rather than sharing memory.
Show you know the cost model: aggregate in the kernel, poll summaries rather than per-event records, and know that iteration with BPF_MAP_GET_NEXT_KEY is not a consistent snapshot.
Own the interface decision — which state lives in maps, which maps are pinned and therefore become a contract between processes, and who is allowed to hold a descriptor to them.
## Why maps exist at all An eBPF program is a short piece of verified bytecode that the kernel runs at a hook point — a kprobe firing, a packet arriving, a tracepoint being hit. It gets a context pointer, a stack of a few hundred bytes, and no ability to write into any process's memory. When it returns, everything local to it is gone. So there has to be somewhere to put results, and somewhere to keep state between invocations. That somewhere is a **map**. A map is a kernel-owned data structure, most commonly a key/value store, that survives across program invocations and is reachable from both the BPF side and user space. Almost every real eBPF tool is structured the same way: a small program in the kernel that filters and aggregates into a map, plus a user-space agent that reads the map and prints, exports, or acts on it. ## The shape is fixed when the map is created Maps are created by the `bpf()` system call with the `BPF_MAP_CREATE` command, which takes: - `map_type` — `BPF_MAP_TYPE_HASH`, `BPF_MAP_TYPE_ARRAY`, `BPF_MAP_TYPE_PERCPU_HASH`, `BPF_MAP_TYPE_LRU_HASH`, `BPF_MAP_TYPE_RINGBUF`, `BPF_MAP_TYPE_PERF_EVENT_ARRAY`, and a few dozen more specialised ones. - `key_size` and `value_size` in bytes. - `max_entries` — the capacity. - flags. All of these are fixed for the life of the map. Keys are exactly `key_size` bytes; there is no variable-length key and no automatic growth when the map fills. In practice you rarely call the syscall by hand — you declare the map in the BPF C source and the loader creates it: ```c struct { __uint(type, BPF_MAP_TYPE_HASH); __uint(max_entries, 10240); __type(key, __u32); __type(value, __u64); } bytes_by_pid SEC(".maps"); ``` The loader creates the map, gets a file descriptor, and patches that descriptor into the program's instructions before loading it, so the verifier knows exactly which map each helper call touches and can check that the key and value sizes the program uses match the map's declaration. ## The kernel side: helpers only BPF code cannot reach into the map's internals. It calls helpers: - `bpf_map_lookup_elem(&map, &key)` returns a pointer to the value **inside the map**, or NULL. The verifier forces a NULL check before you dereference it; because the pointer is into the map itself, writing through it updates the stored value directly. - `bpf_map_update_elem(&map, &key, &value, flags)` inserts or overwrites. Flags are `BPF_ANY`, `BPF_NOEXIST` (fail if present) and `BPF_EXIST` (fail if absent). - `bpf_map_delete_elem(&map, &key)` removes an entry. All of these return an integer status that a careless program ignores — a real source of silent data loss when a map is full. ## The user-space side: it is a file descriptor User space never sees a kernel pointer. It holds a file descriptor and issues `bpf()` commands against it: `BPF_MAP_LOOKUP_ELEM`, `BPF_MAP_UPDATE_ELEM`, `BPF_MAP_DELETE_ELEM`, and `BPF_MAP_GET_NEXT_KEY` for iteration (you walk a hash map by asking for the key after the one you have; passing a NULL/absent key gets the first). Newer kernels add batch commands so you can drain many entries per syscall. libbpf wraps these as `bpf_map_lookup_elem(fd, &key, &value)` and friends, and `bpf_map__fd(skel->maps.bytes_by_pid)` gets the descriptor from a generated skeleton. The important property: **each of these calls copies**. The key goes into the kernel, the value comes back out, one syscall at a time. Polling a large map every second is real overhead, which is why the aggregation belongs in the BPF program and only the summary belongs in the map. The one genuine exception is `BPF_MAP_TYPE_RINGBUF`, whose buffer user space memory-maps and consumes without a syscall per record. ## Lifetime Because the handle is a file descriptor, a map is reference-counted like any other kernel object. References come from open descriptors, from loaded programs that use it, and from bpffs pins. When the last one goes, the map and its contents are freed. That is why a tool that loads a program and exits leaves nothing behind — and why long-lived or shared maps are pinned into the BPF filesystem so a separate process can reopen them later. ## What people get wrong They assume the map is shared memory both sides poke at directly; it is not, outside of the ring buffer. They assume a map grows when it fills; it does not. And they assume the data outlives the loader by default; it does not.
- If maps are freed when the last reference goes, how would you keep one alive after the loading process exits?Pin it into the BPF filesystem, normally mounted at /sys/fs/bpf. A pin is an extra reference held by a filesystem entry, so the map survives the loader's exit and any other privileged process can reopen it by path and keep reading or writing it. Removing the pin drops that reference again.
- How do you iterate over every entry of a BPF hash map from user space, and what should you be careful about?Repeatedly issue BPF_MAP_GET_NEXT_KEY, feeding back the key you just received, and look up each key as you go. It is not a snapshot: the BPF program keeps inserting and deleting while you walk, so you can miss entries or see ones that vanish before the lookup. Batch commands cut the syscall count but not the raciness.
- Why is it better to aggregate inside the BPF program than to export every event and aggregate in user space?Every record you export costs a copy and, for per-key polling, a syscall; a busy hook can fire millions of times a second, so the export path becomes the bottleneck and starts dropping data. Counting or histogramming in the kernel turns millions of events into a handful of map entries that user space reads cheaply.
saying these in an interview costs you the question
- Thinks a map is shared memory both sides write to directly
- Says the map grows automatically when max_entries is reached
- Assumes map data survives the loader process by default
- Believes BPF programs can write into user-space process memory
- Thinks keys can be any length rather than exactly key_size