skip to content

In Linux XDP (eXpress Data Path), where in the packet receive path does an XDP program run, and what do the return codes XDP_PASS, XDP_DROP, XDP_TX, XDP_REDIRECT and XDP_ABORTED each tell the driver to do?

level: middleimportance: must knowfreq 62%

answer

  1. earliest programmable point on receive
  2. before the sk_buff exists
  3. one return value, five verdicts
  4. TX means the same NIC it came in on
  5. ABORTED is the error drop

basics

~20 s

An XDP program runs inside the NIC driver's receive routine, before the kernel allocates an sk_buff. Its return value is a verdict: PASS to the normal stack, DROP the packet, TX back out the same NIC, REDIRECT elsewhere, ABORTED for error drops.

solid answer

~50 s

XDP is the earliest programmable hook on the receive path. In native mode the driver calls the program from its NAPI poll routine on the raw received frame, before an `sk_buff` has been allocated, which is why a dropped packet costs almost nothing. The program gets a `struct xdp_md` carrying `data` and `data_end` pointers and returns one of five verdicts. `XDP_PASS` hands the frame on to the normal network stack. `XDP_DROP` frees it immediately — that is what makes line-rate DDoS filtering possible. `XDP_TX` bounces it back out the same interface it arrived on, so the program has to fix up the Ethernet header itself. `XDP_REDIRECT` sends it to another target — a different netdev, a specific CPU, or an AF_XDP socket — via `bpf_redirect_map()`. `XDP_ABORTED` also drops, but signals a program bug and fires the `xdp_exception` tracepoint.

code

c · 27 lines
c
#include <linux/bpf.h>
#include <linux/if_ether.h>
#include <linux/ip.h>
#include <linux/in.h>
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_endian.h>

SEC("xdp")
int drop_icmp(struct xdp_md *ctx)
{
	void *data = (void *)(long)ctx->data;
	void *data_end = (void *)(long)ctx->data_end;
	struct ethhdr *eth = data;

	if ((void *)(eth + 1) > data_end)
		return XDP_ABORTED;
	if (eth->h_proto != bpf_htons(ETH_P_IP))
		return XDP_PASS;

	struct iphdr *ip = (void *)(eth + 1);
	if ((void *)(ip + 1) > data_end)
		return XDP_ABORTED;

	return ip->protocol == IPPROTO_ICMP ? XDP_DROP : XDP_PASS;
}

char LICENSE[] SEC("license") = "GPL";

go deeper

for a junior

Know that XDP is BPF code that runs in the network driver as packets arrive, and that its return value decides the packet's fate rather than any rule table.

for a middle

Be ready to name all five verdicts, say that the hook runs before the sk_buff is allocated, and explain that this placement — not the language — is where the speed comes from.

for a senior

Show you have run one: bounds checks on every header access, counters in a map because dropped packets are invisible downstream, L2 rewriting before XDP_TX, and XDP_ABORTED reserved for genuine bugs.

for a principal

Frame where the hook belongs in a defence design: volumetric drops at the driver, policy that needs connection context further up, and the operational cost of putting code that can silently blackhole traffic into the driver path.

## Where the hook sits A NIC writes an incoming frame into a receive ring and raises an interrupt. The driver's NAPI poll routine walks that ring, and normally the next thing it does is allocate an `sk_buff` — the kernel's per-packet metadata structure — and hand it to `netif_receive_skb()`, from where the packet travels through the tc ingress hook, netfilter, routing, the transport layer and finally into a socket. XDP inserts a call *before* that allocation. In native mode the driver invokes the attached BPF program directly on the receive buffer, so if the verdict is "drop", the kernel never built any per-packet state to throw away. This is the whole performance story: the cost of a dropped packet collapses from "allocate an skb, walk several layers, free it" to "run a few hundred instructions and recycle the page". Published numbers for simple XDP drop programs on a single core sit in the tens of millions of packets per second; the equivalent netfilter rule is an order of magnitude slower. ## The context object The program receives a `struct xdp_md`. The important fields are `data` and `data_end`, both integers you cast to pointers, plus `data_meta` for a small scratch area you can hand to a later tc program, and `ingress_ifindex` and `rx_queue_index` for the arrival point. There is no protocol parsing done for you and no metadata: you are looking at raw bytes starting at the Ethernet header, and every read must be preceded by a bounds check against `data_end` or the program will not load. ## The five verdicts `XDP_PASS` means "not my business" — build the skb and continue up the stack normally. A real filter passes the overwhelming majority of traffic, so this is the common path. `XDP_DROP` frees the buffer on the spot. Nothing further in the kernel sees the packet, which has an important consequence: tools that tap the stack later, such as tcpdump, will not show it either. Keep your own counters in a map if you want to know what you dropped. `XDP_TX` retransmits the frame out of the *same* interface it arrived on. It cannot choose another device. Because the program is operating below the routing layer, nothing rewrites the Ethernet header for you — if you are bouncing a packet back, you swap the source and destination MAC addresses yourself. This is the primitive behind XDP load balancers and behind trivial things like an in-driver ICMP echo responder. `XDP_REDIRECT` is the general escape hatch. Combined with `bpf_redirect_map()` it can push the frame to another network device, to another CPU for further processing, or into an AF_XDP socket for a userspace program. Redirecting to a different netdev requires that device's driver to implement the transmit-side XDP entry point, so not every target works. `XDP_ABORTED` drops the packet like `XDP_DROP` but declares the drop to be an error. It fires the `xdp_exception` tracepoint, which is what you attach to when you want to know that your program hit an impossible branch. Use it in the arms that should never be taken — a failed bounds check, a malformed header — and never as your ordinary drop. ## Attaching and inspecting Attachment is a property of a network device, not of a qdisc: ```bash ip link set dev eth0 xdp obj xdp_drop.o sec xdp ip link show dev eth0 # shows the attached mode and prog id bpftool net show dev eth0 # same, from the bpftool side ip link set dev eth0 xdp off ``` Because it is per-device and ingress-only, XDP has no say over anything the host transmits. Outbound filtering belongs to the tc egress hook instead.

  • Your program returns XDP_TX after rewriting the destination IP. Why does the packet still arrive at the wrong machine?
    Because XDP sits below routing and neighbour resolution, nothing recalculates the Ethernet header for you. If you only touched L3, the frame goes back out with its original destination MAC and lands wherever that address points. You must rewrite the MAC addresses yourself, and fix the IP checksum, before returning `XDP_TX`.
  • How would you count how many packets your XDP program dropped, given that the packets never reach the rest of the stack?
    Keep the counters in the program: increment a per-CPU array map keyed by drop reason, and read it from userspace with `bpf_map_lookup_elem()` or `bpftool map dump`. Many drivers also export XDP counters through `ethtool -S`, and `xdp_exception` gives you the `XDP_ABORTED` count without any extra code.
  • Why must every packet read in an XDP program be preceded by a comparison against ctx->data_end?
    The program runs in the kernel on an attacker-influenced buffer, so the kernel will not accept any access it cannot prove is in bounds. Each dereference has to be dominated by an explicit check against `data_end`; without it the program is rejected at load time rather than failing at runtime.

XDP is the bouncer at the door of the building rather than the receptionist on the fifteenth floor: turning someone away costs a sentence, not a lift ride, a visitor badge and a walk back down.

saying these in an interview costs you the question

  • Thinking XDP can also filter outbound traffic
  • Believing XDP_TX can pick any interface
  • Using XDP_ABORTED as the ordinary drop code
  • Assuming XDP sees a fully built sk_buff
  • Claiming XDP runs after netfilter's hooks

context