skip to content

An eBPF tool needs to stream individual event records to a user-space agent. What does BPF_MAP_TYPE_RINGBUF give you that BPF_MAP_TYPE_PERF_EVENT_ARRAY does not, and when would you still reach for the perf buffer?

level: middleimportance: must knowfreq 66%

answer

  1. one shared buffer versus one ring per CPU
  2. memory that scales with core count
  3. reserve first, fill in place, then submit
  4. ordering across CPUs survives in one of them
  5. 5.8 is the dividing line

basics

~20 s

BPF_MAP_TYPE_RINGBUF is one buffer shared by all CPUs: events keep their global order, memory does not scale with CPU count, and bpf_ringbuf_reserve() lets a program fill a record in place. The perf buffer is per-CPU and is the only option below Linux 5.8.

solid answer

~50 s

The perf buffer, `BPF_MAP_TYPE_PERF_EVENT_ARRAY`, is one ring per CPU. You size every ring, so memory scales with core count, a program on a quiet CPU cannot borrow space from a busy one, and because the consumer polls the rings independently, records arrive interleaved rather than in true time order. `bpf_perf_event_output()` also copies the record from the program's stack, so you pay for the copy even when the buffer is full and the record is thrown away. The BPF ring buffer, added in 5.8, is a single MPSC buffer shared by all CPUs: one size to tune, global ordering, and a reservation API — `bpf_ringbuf_reserve()` returns a pointer straight into the buffer, you fill it, then `bpf_ringbuf_submit()` or `bpf_ringbuf_discard()`. If the reserve returns NULL you drop the event before doing any work. It is the default choice today. Stay on the perf buffer for kernels older than 5.8, or when the shared reservation lock becomes the bottleneck at extreme per-CPU event rates.

code

c · 27 lines
c
#include "vmlinux.h"
#include <bpf/bpf_helpers.h>

struct event {
    __u32 pid;
    char comm[16];
};

struct {
    __uint(type, BPF_MAP_TYPE_RINGBUF);
    __uint(max_entries, 256 * 1024);   /* bytes: power of two, page multiple */
} events SEC(".maps");

SEC("tracepoint/syscalls/sys_enter_execve")
int on_execve(void *ctx)
{
    struct event *e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);
    if (!e)
        return 0;                      /* buffer full: drop, no copy wasted */

    e->pid = bpf_get_current_pid_tgid() >> 32;
    bpf_get_current_comm(&e->comm, sizeof(e->comm));
    bpf_ringbuf_submit(e, 0);
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

go deeper

for a junior

Know that per-event records leave the kernel through a buffer map, and that the modern one is BPF_MAP_TYPE_RINGBUF while older tools use the per-CPU perf buffer.

for a middle

Explain the mechanics: shared versus per-CPU rings, memory scaling with core count, global ordering, and the reserve/submit/discard sequence that avoids a wasted copy.

for a senior

Show operational judgment — count and export dropped events, filter and aggregate in-kernel before enlarging the buffer, and know that reservation contention is the one case that still favours per-CPU rings.

for a principal

Own the export contract for a fleet: minimum kernel version, whether tools must run both paths, and what loss rate is acceptable before the event stream stops being trustworthy evidence.

## The problem Counters aggregate nicely into a hash map that user space polls once a second. Per-event records — a process exec, a failed connect, a slow request with its arguments — do not. They are variable in volume, they arrive at whatever rate the workload dictates, and losing them silently is a bug. What you need is a queue from the kernel to a user-space consumer, and eBPF offers two. ## The perf buffer `BPF_MAP_TYPE_PERF_EVENT_ARRAY` predates eBPF's ring buffer and is built on the perf subsystem's per-CPU mmap'd rings. The map is an array indexed by CPU; user space opens a perf event and a buffer for each CPU and polls them all. The BPF side calls: ```c struct event e = {}; e.pid = bpf_get_current_pid_tgid() >> 32; bpf_perf_event_output(ctx, &events, BPF_F_CURRENT_CPU, &e, sizeof(e)); ``` That works, and it worked for years — bcc's tools are full of it. But the per-CPU design has consequences: **Memory scales with CPUs.** You choose a per-CPU page count; a 96-core box multiplies it 96 times. Sizing generously is expensive, sizing tightly means the one hot CPU overflows while 95 rings sit empty. **No sharing of capacity.** Events on a busy CPU cannot spill into another CPU's ring, so loss is decided per CPU. **Ordering is lost.** Each ring is drained independently, so the consumer sees CPU-local order, not global order. If you are correlating an entry and exit event that happened on different CPUs, you must reorder by timestamp yourself. **You copy before you know there is room.** `bpf_perf_event_output()` takes a filled-in struct — usually built on the program's small stack — and copies it in. If the ring is full, the work and the copy are wasted, and the loss is reported to the consumer as a separate lost-samples callback. ## The BPF ring buffer `BPF_MAP_TYPE_RINGBUF`, available since Linux 5.8, is a single multi-producer/single-consumer buffer shared by every CPU. `max_entries` is the buffer size in bytes, and must be a power of two and a multiple of the page size. The API is what makes it better, not just the topology: ```c struct event *e = bpf_ringbuf_reserve(&events, sizeof(*e), 0); if (!e) return 0; /* no space — drop before doing work */ e->pid = bpf_get_current_pid_tgid() >> 32; bpf_get_current_comm(&e->comm, sizeof(e->comm)); bpf_ringbuf_submit(e, 0); ``` Reserve hands you a pointer *into the buffer*. You write the record once, in place. Then you either `bpf_ringbuf_submit()` it, making it visible to the consumer, or `bpf_ringbuf_discard()` it if you decide mid-way that it is not interesting — the space is released and nothing is ever copied. Reservation also gives you honest backpressure: a NULL return is your signal that the consumer is behind, checked *before* you spend cycles filling anything in. (`bpf_ringbuf_output()` also exists, taking a filled struct like the perf helper, for cases where the record's size is not known until it is built.) Because reservation is globally ordered, records reach user space in submission order across all CPUs. Memory is one number to tune rather than one number times the core count. And consumption is epoll-driven with adaptive wakeups — the producer can suppress or force a wakeup with `BPF_RB_NO_WAKEUP` and `BPF_RB_FORCE_WAKEUP` — so a high-rate stream does not wake the consumer per record. libbpf exposes `ring_buffer__new()` and `ring_buffer__poll()`, against `perf_buffer__new()` and `perf_buffer__poll()` for the older type. ## When the perf buffer is still the answer Two honest cases. First, kernel support: if you must run on kernels older than 5.8, the ring buffer does not exist, and portable tools ship both paths and pick at load time. Second, contention: reservation on a shared buffer takes a lock, so at extreme event rates spread across many cores, per-CPU rings can sustain higher aggregate throughput precisely because producers never touch the same cache line. That is a measure-it situation, not a default. ## The part that matters more than either choice Both are bounded queues, and a hook that fires a million times a second will outrun any consumer. The design fix is upstream of the buffer: filter in the BPF program, aggregate what can be aggregated into a hash or per-CPU map, and reserve the event stream for records that genuinely need to be seen individually. Whichever buffer you use, count your drops — the ring buffer's NULL reserve, the perf buffer's lost-sample callback — and surface that number, because a tool that quietly discards events is worse than one that admits it is behind.

  • What is the practical difference between bpf_ringbuf_submit() and bpf_ringbuf_discard() after a successful reserve?
    Submit publishes the reserved record so the consumer sees it; discard releases the same space and the consumer never does. Discard is what lets you reserve early, inspect the event as you fill it, and abandon it — a failed-syscall filter, say — without having built and copied a record you then throw away.
  • Your ring-buffer-based tool reports missing events under load. What do you do about it?
    First make the loss visible: a NULL from bpf_ringbuf_reserve() means the consumer is behind, so count those in a separate map and export the number. Then reduce the stream rather than only enlarging the buffer — filter harder in the BPF program, aggregate what can be aggregated, and only then size the buffer up. A bigger buffer delays overflow; it does not fix a producer faster than its consumer.
  • How does a portable tool support both kernels with and without BPF_MAP_TYPE_RINGBUF?
    Ship both output paths in the same object and decide at load time — detect ring buffer support, then disable the unused program and map before loading so the verifier never sees the unsupported one. libbpf's autoload controls make this routine, and the user-space side selects the matching consumer API accordingly.

saying these in an interview costs you the question

  • Thinks the BPF ring buffer is also one ring per CPU
  • Assumes perf buffer records arrive in global time order
  • Believes a bigger buffer removes event loss
  • Ignores a NULL return from bpf_ringbuf_reserve()
  • Calls bpf_ringbuf_submit() on a reserve that returned NULL

context