You need packet filtering and load balancing at close to line rate on commodity Linux servers. How would you decide between an XDP/eBPF datapath and a kernel-bypass framework such as DPDK, and what does each approach cost you?
answer
- one takes the NIC away entirely
- idle cost is not the same
- what happens to tcpdump and ss
- the unhandled packets are the question
- AF_XDP sits between them
basics
~20 sDPDK takes the NIC away from the kernel for maximum throughput, at the cost of the whole kernel network stack, its tooling, and dedicated busy-polling cores. XDP stays inside the kernel, slightly slower but selective, interrupt-driven and operable with normal tools. Choose by how much of the traffic is exceptional.
solid answer
~60 sThe decision comes down to how much of the stack you are willing to reimplement. DPDK unbinds the NIC from the kernel driver and polls it from userspace, so you get the raw hardware and nothing else — no kernel TCP stack, no routing table, no `ss`, no tcpdump, no iptables on that interface — and you burn whole cores at 100% whether or not traffic is arriving. XDP keeps the NIC under the kernel and runs your program per packet in the driver, so anything you do not handle can be passed up to the normal stack with `XDP_PASS`; the machine remains an ordinary Linux box you can debug. DPDK is faster at the top end and gives you total control of the datapath; XDP gets you most of the way with a fraction of the operational surface. In between sits AF_XDP, which keeps the kernel driver but gives a userspace program a fast, optionally zero-copy path. For filtering and L4 load balancing — where the program is small and most traffic is uniform — XDP is usually the right default, and DPDK earns its cost only when you are building a dedicated appliance.
go deeper
Know that kernel bypass means a userspace program drives the NIC directly, while XDP runs your code inside the kernel, and that only one of the two leaves the machine looking like a normal Linux host.
Be able to explain the mechanics: unbinding to vfio-pci and polling versus a verified program in the driver receive path, and what XDP_PASS means for traffic your program does not handle.
Argue the operational side concretely — lost tooling, hugepages and pinned cores, idle CPU burn, what happens to management traffic — and place AF_XDP correctly as the selective userspace path.
Make it a decision with stated criteria: the fraction of traffic that is exceptional, who operates the result, what has to be reimplemented, and the total cost of a fleet of appliances against a fleet of ordinary servers running a small in-kernel program.
## What each one actually does **DPDK** is a bypass. You unbind the NIC from its kernel driver and bind it to a userspace I/O driver (`vfio-pci`), after which the kernel no longer knows the device exists as a network interface. A userspace poll-mode driver maps the device's queues and spins on them. Packets are delivered to your process as raw buffers, and everything above the wire — ARP, IP reassembly, TCP, routing — is your problem, either implemented by you or borrowed from a userspace stack. **XDP** is an in-kernel hook. The NIC keeps its kernel driver and remains a normal interface; your BPF program runs in the driver's receive path and returns a verdict per packet. Whatever you do not want to handle returns `XDP_PASS` and continues up the ordinary stack. That difference — total ownership versus selective interception — drives everything else. ## Performance DPDK's ceiling is higher. Nothing sits between your code and the ring, batching and prefetching are entirely under your control, and there is no per-packet kernel involvement at all. Published XDP results reach tens of millions of packets per second per core for simple programs, which is enough to saturate typical server NICs on a handful of cores; DPDK will beat that, and the gap widens as the per-packet work shrinks toward the trivial. But the shape of the cost differs as much as the size. DPDK's poll-mode driver spins: an idle DPDK core reads 100% CPU and consumes power like a busy one. XDP is driven by the ordinary NAPI mechanism, so an idle machine is idle. If your traffic is bursty or your fleet is large, that difference in the average case can outweigh the difference in the peak. ## The operational surface This is where most real decisions are made. Under DPDK the interface is gone from the kernel's point of view. `ip`, `ss`, `tcpdump`, `ethtool` statistics, netfilter, the routing table and every runbook that uses them stop applying to that NIC. Your process needs hugepages, pinned cores and, typically, exclusive hardware. Anything the kernel used to do for you — including management traffic on that link — must be re-provided, which is why DPDK deployments usually dedicate one NIC to the datapath and another to management. You have, in effect, built an appliance. Under XDP the machine stays a Linux machine. You keep SSH on the same NIC, you keep the stack for everything the program passes, and `bpftool` inspects the program and its maps live. Programs are attached and detached without restarting anything, and they are verified before they load, so a buggy program is rejected rather than crashing the kernel. The cost is real constraints. Your code must satisfy the kernel's safety checks: bounded execution, provable memory access, a limited helper set. There is no arbitrary library code, no floating point, no blocking. Complex per-packet logic is genuinely harder to express in XDP than in a userspace C program, and some of it cannot be expressed at all. ## AF_XDP as the middle road AF_XDP is worth knowing precisely because it dissolves the false dichotomy. The kernel driver stays in place; an XDP program redirects selected packets into a socket bound to a specific receive queue, and a userspace process reads them from a shared memory region, with zero copy where the driver supports it. You get userspace flexibility for the traffic you care about, the kernel stack for everything else, and no unbinding of the device. It does not match DPDK's absolute ceiling, and it introduces its own complexity, but for many "we need userspace processing of a subset of traffic" problems it is the correct answer. ## How to decide Ask what fraction of the traffic is exceptional. If nearly every packet needs custom handling and the box exists solely to do that — a router, a firewall appliance, a telco function — DPDK's ownership model matches the problem and its costs are already accepted. If most packets are ordinary and a minority need custom treatment, or the machine has another job as well, bypass forces you to reimplement the ordinary path, and XDP's selectivity is worth far more than the last increment of throughput. Then ask who operates it. XDP keeps the box debuggable with the tools your on-call already has; DPDK requires a team that owns a userspace stack, its dependencies and its own diagnostics. For filtering and L4 load balancing specifically, the program is small and uniform, most traffic is either dropped or forwarded by a fixed rule, and the industry's large deployments run in-kernel — which is a strong signal that the extra ceiling rarely justifies the extra surface.
- Why is an idle DPDK core still at 100% CPU, and when does that actually matter?Its poll-mode driver spins on the receive ring rather than waiting for interrupts, which is how it avoids interrupt latency and per-packet kernel entry. It matters whenever utilisation is not near-constant: bursty traffic, multi-tenant hosts, or a large fleet where the power and the reserved cores are charged continuously for peak-rate capability you use occasionally.
- What can a DPDK application do that a verified eBPF program cannot?Run arbitrary code: unbounded loops, large data structures, third-party libraries, floating point, its own memory management, and a full userspace TCP stack. eBPF must satisfy the kernel's safety checks before it loads, so complex stateful per-packet logic is constrained in ways userspace code is not.
- You have chosen XDP, but one class of traffic needs processing too complex to express there. What is the path forward?Redirect just that class into userspace instead of moving the whole datapath. The XDP program handles the bulk in-kernel and redirects the exceptional flows into an AF_XDP socket, where a userspace process does the complex work. Everything else keeps passing to the normal stack, and the NIC stays a kernel interface.
- How does the safety story differ between a crash in a DPDK datapath and a fault in an XDP program?A DPDK crash takes down the process that owns the NIC, and with it the traffic on that interface, since nothing else is driving the device. An XDP program cannot crash the kernel — it is verified before loading — but it can still blackhole traffic through incorrect logic, and it can be detached instantly, which a wedged bypass process cannot be.
saying these in an interview costs you the question
- Choosing DPDK purely on peak packets per second
- Assuming tcpdump still works on a bypassed NIC
- Thinking eBPF can run arbitrary userspace-style code
- Ignoring that DPDK cores burn CPU while idle
- Treating AF_XDP as identical to full kernel bypass