The verifier rejects an eBPF program because its stack usage is too large. How much stack does an eBPF program get on Linux, why is the limit so small, and where should a large scratch buffer live instead?
answer
- 512 bytes, and that is all
- it borrows the kernel's stack
- the whole call chain shares the budget
- big buffers belong in a map
- one entry, per-CPU, key zero
basics
~20 sAn eBPF program gets 512 bytes of stack, and with subprograms it is the combined depth of the call chain that must fit. The limit is small because the program runs on the ordinary kernel stack in any context. Large buffers belong in a per-CPU array map used as scratch space.
solid answer
~60 sThe verifier enforces a hard **512-byte** stack budget, and it checks the *combined* depth of a call chain rather than each frame in isolation — nested subprograms share the budget, and the call depth itself is capped. The reason is that eBPF does not get a stack of its own: it borrows the kernel stack of whatever context the hook fired in, which on x86-64 is typically only 16 KB total and may already be deep inside a syscall or an interrupt handler. A program that grabbed a few kilobytes could overflow it and take the machine down. So a `char buf[1024]` local is a load failure, not a runtime problem. The standard fix is a **per-CPU array map with a single entry** whose value is the large structure: you look up index 0, get a pointer to per-CPU scratch memory of any size, and no other CPU can be using the same copy. Global variables in a libbpf program have the same property for free, since libbpf places them in an internal map rather than on the stack.
code
c · 28 lines#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>
struct scratch {
char buf[4096]; /* far beyond the 512-byte stack budget */
};
struct {
__uint(type, BPF_MAP_TYPE_PERCPU_ARRAY);
__uint(max_entries, 1);
__type(key, __u32);
__type(value, struct scratch);
} scratch_map SEC(".maps");
SEC("tracepoint/syscalls/sys_enter_openat")
int big_buffer(void *ctx)
{
__u32 zero = 0;
struct scratch *s = bpf_map_lookup_elem(&scratch_map, &zero);
if (!s)
return 0;
s->buf[0] = 'x'; /* per-CPU scratch, no stack used */
return 0;
}
char LICENSE[] SEC("license") = "GPL";go deeper
Remember the number: an eBPF program has 512 bytes of stack, so large local buffers are a load failure. Know that the answer is to put the buffer in a map.
Explain that the program runs on the borrowed kernel stack, that the verifier statically checks the combined depth of the call chain, and that a single-entry per-CPU array map is the standard scratch-space idiom.
Show judgment about the alternatives — per-CPU scratch, libbpf globals, writing straight into a ring buffer reservation — and name the concurrency trap of using a shared array map instead of a per-CPU one.
Frame it as a deliberate design constraint: eBPF attaches to arbitrary hot kernel paths, so a small statically-provable frame is what keeps a tracing tool from panicking a production host.
## The number and the shape of the rule An eBPF program may use at most **512 bytes** of stack. This is not a soft target: the verifier computes the maximum stack depth statically and refuses the load if it exceeds the budget. Two details matter beyond the headline number: - **Call chains share the budget.** If a program calls a subprogram which calls another, the verifier checks the combined depth of the chain against the limit, so splitting a big buffer across two functions does not help. Nested call depth itself is also capped at a small number of frames. - **It is a static check.** The verifier does not observe usage; it computes the worst case from the bytecode. Code paths you never expect to run still count. ## Why 512 bytes eBPF programs do not run on a stack of their own. They execute on the kernel stack of whatever context the hook fired in — a syscall entry, a network softirq, an interrupt handler, a scheduler tracepoint. On x86-64 that whole stack is typically **16 KB** for the entire kernel call chain, and by the time you are deep inside a filesystem or networking path a substantial fraction is already consumed. Kernel stack overflow is not a graceful failure; it corrupts adjacent memory or panics the box. Given that eBPF's whole premise is that untrusted-ish code can be attached to arbitrary hot paths, allowing that code a large frame would give away exactly the safety property the verifier exists to guarantee. A tight, statically checkable ceiling is the cheap way to keep the guarantee absolute. ## What triggers the rejection in practice The usual culprits are all recognisable: - a local buffer for a path, a command line, or a filename — `char path[4096]` is eight times the whole budget on its own; - a large event struct built on the stack before being pushed to user space; - an on-stack array indexed in a loop, especially one that also unrolls; - several medium-sized structs in nested subprograms whose combined depth crosses the line. ## The per-CPU scratch map pattern The idiomatic answer is to move the buffer into a map: ```c struct scratch { char buf[4096]; }; struct { __uint(type, BPF_MAP_TYPE_PERCPU_ARRAY); __uint(max_entries, 1); __type(key, __u32); __type(value, struct scratch); } scratch SEC(".maps"); ``` Look up key 0, NULL-check the result, and you have a pointer to memory of arbitrary size that costs no stack. Choosing **per-CPU** rather than a plain array matters: each CPU gets its own copy, so two CPUs running the program simultaneously cannot scribble on each other's scratch space. That said, per-CPU is not a full mutual-exclusion guarantee — a program interrupted by another that uses the same scratch entry on the same CPU can still collide, which is why per-hook scratch entries or careful use are worth thinking about in tracing programs attached to nested paths. Two related patterns worth knowing: - **Global variables cost no stack.** In a libbpf program, a file-scope variable lives in an internal array map backing the object's data sections, not on the stack. Declaring your big struct as a global — or `static` at file scope — sidesteps the limit without an explicit map. - **Ring buffer reservation writes in place.** For events destined for user space, reserving space in a `BPF_MAP_TYPE_RINGBUF` and filling it directly avoids building the record on the stack and copying it, which saves both the stack and the copy. ## What people get wrong The most common wrong instinct is trying to shrink the struct until it just fits — squeezing a filename into 200 bytes because the budget allows it. That trades a load failure for silent truncation in production. The second is assuming the limit is per function, and refactoring a big local into a helper; the combined-depth rule defeats that immediately. The third is reaching for a plain (non-per-CPU) array map as scratch, which loads fine and then produces corrupted events under concurrency — a bug that only appears under load, which is the worst kind.
- Why per-CPU rather than an ordinary array map for scratch space?An ordinary array map has one shared copy, so two CPUs running the program at once overwrite each other's working data — corrupted events that only appear under load. A per-CPU array gives each CPU its own copy of the value, which removes that race without any locking cost on the fast path.
- Does moving the buffer into a subprogram help?No. The verifier checks the combined stack depth of the whole call chain, not each frame separately, and it also caps nesting depth. Refactoring a large local into a helper function moves where the bytes are declared without changing the total, so the same rejection comes back.
- Is there a way to build an event without a stack copy at all?Yes — reserve space in a ring buffer and write the record in place, then submit it. That avoids both the stack frame and the copy from stack into the buffer, which is why it is the preferred shape for high-rate event output as well as a way around the stack budget.
saying these in an interview costs you the question
- Thinks the limit applies per function, not per call chain
- Shrinks the buffer until it fits and accepts truncation
- Uses a shared array map as scratch and races under load
- Believes eBPF has a private stack of its own
- Assumes the check is at runtime rather than at load