skip to content

Networking: XDP and tc

eBPF in the datapath: XDP actions (PASS, DROP, TX, REDIRECT) at the earliest point in the stack, tc ingress and egress hooks, socket filtering, load-balancing patterns, and how it compares with kernel bypass such as DPDK. Interviewers ask because XDP is how modern DDoS filtering and L4 load balancers reach line rate without leaving the kernel.

on this pageshow

questions

6

In Linux eBPF networking, what does a program attached to the tc (traffic control) hook see that an XDP program does not, and when would you choose tc over XDP?

level: middleimportance: must knowfreq 55%

answer

  1. one hook has an sk_buff, one does not
  2. metadata versus raw bytes
  3. only one of the two sees egress
  4. clsact and direct-action
  5. driver support only matters for one

basics

~20 s

A tc BPF program runs after the kernel has built an sk_buff, so it sees packet metadata, works on every driver, and can be attached on egress as well as ingress. XDP is faster but ingress-only and metadata-free.

solid answer

~40 s

XDP trades context for speed; tc trades speed for context. By the time a tc program runs, the kernel has allocated an `sk_buff`, so the program's `struct __sk_buff` exposes things XDP cannot know: the packet's `mark`, `priority`, `ifindex`, protocol fields the stack already parsed, and on egress the owning socket and cgroup. It also works on any interface regardless of driver support, because there is nothing driver-specific about it, and it has an egress hook — XDP has none. You attach it to the `clsact` qdisc on ingress or egress and return `TC_ACT_OK` or `TC_ACT_SHOT` in direct-action mode. So: XDP for volumetric drop and forwarding at the earliest possible point, tc when you need policy that depends on who sent the packet, or you need to touch outbound traffic at all.

code

bash · 6 lines
bash
# clsact provides ingress and egress attach points without queueing
tc qdisc add dev eth0 clsact
tc filter add dev eth0 ingress bpf da obj tc_prog.o sec classifier
tc filter add dev eth0 egress bpf da obj tc_prog.o sec classifier
tc filter show dev eth0 egress
bpftool net show dev eth0

go deeper

for a junior

Know that eBPF can be attached at more than one point in the network path, and that the tc hook is the one that also handles outbound traffic.

for a middle

Be able to contrast the two contexts concretely — raw bytes versus an sk_buff with mark, priority and ifindex — and name the tc action codes and the clsact attach points.

for a senior

Show you would place logic deliberately: floods and forwarding at XDP, context-dependent policy and egress at tc, with the metadata and driver-support constraints driving the split rather than a preference for the faster hook.

for a principal

Own the coexistence problem: multiple agents attaching to one interface, ordering and ownership guarantees, and whether your platform can depend on a kernel new enough for tcx across the whole fleet.

## Two hooks, two data structures The two hooks differ mainly in what has already happened to the packet when your code runs. An XDP program runs in the driver on a raw receive buffer. Its context, `struct xdp_md`, is little more than a pair of pointers into those bytes. Nothing has parsed the packet, nothing has classified it, and nothing has associated it with a socket, a process or a cgroup. A tc program runs after `sk_buff` allocation. Its context, `struct __sk_buff`, is a sanitised view of that structure and carries real metadata: `len`, `protocol`, `ifindex`, `mark`, `priority`, `tc_classid`, `cb` for passing values between programs, plus VLAN and tunnel fields. On egress you can additionally reach socket-level facts, which is how per-workload policy gets implemented. ## What the sk_buff buys, and what it costs The metadata is the point, and so is the cost. Allocating and initialising an `sk_buff` is precisely the work XDP exists to skip, so a tc drop is materially more expensive than an XDP drop. For a volumetric flood — millions of packets a second you want gone — that difference decides whether the machine survives. For a few hundred thousand packets a second of policy enforcement it is irrelevant, and the metadata is worth far more than the cycles. The skb also means you are working with the stack's own representation of the packet, so helpers can grow, shrink, encapsulate and checksum it with the stack's cooperation: `bpf_skb_store_bytes()`, `bpf_l3_csum_replace()`, `bpf_l4_csum_replace()`, `bpf_skb_adjust_room()`, `bpf_redirect()`. XDP has a narrower and lower-level set. ## Ingress, egress, and portability XDP is a receive-path hook. There is no general XDP hook on transmit, so anything you want to do to packets the host sends — outbound filtering, marking, encapsulation — is tc's job. XDP in native mode also depends on the driver implementing it. Drivers such as mlx5, ixgbe, i40e, bnxt and virtio-net do; plenty of others do not, and stacked virtual devices are their own story. tc BPF has no such requirement: it lives in generic kernel code, so it works everywhere, which matters when you are shipping something that has to run on hardware you have not seen. ## Verdicts and attachment tc programs return traffic-control action codes rather than XDP verdicts: `TC_ACT_OK` continues processing, `TC_ACT_SHOT` drops, `TC_ACT_REDIRECT` (via `bpf_redirect()`) moves the packet to another device, `TC_ACT_UNSPEC` falls through to the next filter. In *direct-action* mode — the `da` keyword — the program's return value is taken as the action verdict directly, instead of returning a classid that a separate action then interprets. Effectively everyone uses direct-action mode. Attachment goes through the `clsact` qdisc, which exists specifically to provide ingress and egress attach points without doing any queueing: ```bash tc qdisc add dev eth0 clsact tc filter add dev eth0 ingress bpf da obj tc_prog.o sec classifier tc filter add dev eth0 egress bpf da obj tc_prog.o sec classifier tc filter show dev eth0 ingress ``` Since kernel 6.6 there is also the `tcx` attach point, which gives tc programs proper `bpf_link` ownership semantics and a defined multi-program ordering, instead of the older model where two pieces of software could quietly fight over the same filter list. `bpftool net show` lists what is attached either way. ## Choosing Use XDP when the answer is "this packet should never have been built": flood drops, an L4 load balancer's forwarding path, anything measured in millions of packets per second. Use tc when you need to know something about the packet's context, when you need to act on egress, when the traffic volume does not justify the constraints, or when you cannot guarantee driver support. Systems that do both — Cilium being the well-known example — put the frontend drop and forwarding logic at XDP and the per-workload policy at tc.

  • What does the direct-action mode flag change about how a tc BPF program's return value is interpreted?
    Without it, the classifier is expected to return a class identifier, and a separate tc action object then decides what happens to the packet. With `da`, the return value *is* the verdict — `TC_ACT_OK`, `TC_ACT_SHOT`, `TC_ACT_REDIRECT` — so one program does both classification and action, with no second lookup and no separate action configuration.
  • You need to encapsulate outbound packets in a tunnel header. Why is tc the natural hook for that and not XDP?
    Two reasons. XDP has no egress hook at all, so outbound traffic never reaches it. And growing a packet is far easier on an skb: `bpf_skb_adjust_room()` makes room and cooperates with the stack's checksum and offload handling, whereas XDP's headroom is whatever the driver reserved.
  • Why does the tcx attach point introduced in kernel 6.6 matter for software that ships tc programs?
    The old `tc filter` model has no ownership: two agents attaching to the same interface can overwrite or reorder each other's programs, and a crashed agent leaves programs behind. tcx gives each attachment a `bpf_link` with defined lifetime and explicit relative ordering, so independent pieces of software can coexist on one interface predictably.

saying these in an interview costs you the question

  • Saying tc BPF also runs before the sk_buff
  • Thinking XDP has an egress hook too
  • Assuming tc needs driver support like XDP does
  • Believing tc programs return XDP_DROP or XDP_PASS
  • Treating XDP as strictly better because it is faster

context

open as a page

In Linux XDP (eXpress Data Path), where in the packet receive path does an XDP program run, and what do the return codes XDP_PASS, XDP_DROP, XDP_TX, XDP_REDIRECT and XDP_ABORTED each tell the driver to do?

level: middleimportance: must knowfreq 62%

basics

~20 s

An XDP program runs inside the NIC driver's receive routine, before the kernel allocates an sk_buff. Its return value is a verdict: PASS to the normal stack, DROP the packet, TX back out the same NIC, REDIRECT elsewhere, ABORTED for error drops.

open as a page

In an XDP-based L4 load balancer, what is the difference between returning XDP_TX and returning XDP_REDIRECT for a forwarded packet, and why do such designs typically encapsulate the packet and use direct server return rather than plain NAT?

level: seniorimportance: should knowfreq 33%

basics

~20 s

XDP_TX resends the packet out the interface it arrived on; XDP_REDIRECT hands it to a different target through a redirect map. Encapsulation with direct server return lets backends reply straight to the client, so only the request half crosses the load balancer.

open as a page

You attach an XDP drop program to a 25 GbE Linux server's interface. The filter works, but throughput and CPU use are no better than the iptables rule it replaced, and `ip link show` reports the attachment as xdpgeneric. What are XDP's three attach modes, and what has happened here?

level: seniorimportance: should knowfreq 40%

basics

~20 s

XDP attaches in native mode inside the driver, generic mode in the shared kernel receive path after the sk_buff is allocated, or offloaded onto a capable NIC. Generic mode is a correctness fallback with none of the performance benefit — the driver did not support native XDP.

open as a page

You need packet filtering and load balancing at close to line rate on commodity Linux servers. How would you decide between an XDP/eBPF datapath and a kernel-bypass framework such as DPDK, and what does each approach cost you?

level: principalimportance: should knowfreq 42%

basics

~20 s

DPDK takes the NIC away from the kernel for maximum throughput, at the cost of the whole kernel network stack, its tooling, and dedicated busy-polling cores. XDP stays inside the kernel, slightly slower but selective, interrupt-driven and operable with normal tools. Choose by how much of the traffic is exceptional.

open as a page

Your XDP program on a Linux host is dropping flood packets with XDP_DROP, but `tcpdump -i eth0` on that host shows none of those packets at all. Why can tcpdump not see them, and where does its own filter actually run?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

tcpdump reads packets from an AF_PACKET socket that is fed later in the receive path, after the driver has built an sk_buff. Native XDP runs before that, so packets it drops never reach the tap. Count drops in a BPF map instead.

open as a page