skip to content

In Linux eBPF networking, what does a program attached to the tc (traffic control) hook see that an XDP program does not, and when would you choose tc over XDP?

level: middleimportance: must knowfreq 55%

answer

  1. one hook has an sk_buff, one does not
  2. metadata versus raw bytes
  3. only one of the two sees egress
  4. clsact and direct-action
  5. driver support only matters for one

basics

~20 s

A tc BPF program runs after the kernel has built an sk_buff, so it sees packet metadata, works on every driver, and can be attached on egress as well as ingress. XDP is faster but ingress-only and metadata-free.

solid answer

~40 s

XDP trades context for speed; tc trades speed for context. By the time a tc program runs, the kernel has allocated an `sk_buff`, so the program's `struct __sk_buff` exposes things XDP cannot know: the packet's `mark`, `priority`, `ifindex`, protocol fields the stack already parsed, and on egress the owning socket and cgroup. It also works on any interface regardless of driver support, because there is nothing driver-specific about it, and it has an egress hook — XDP has none. You attach it to the `clsact` qdisc on ingress or egress and return `TC_ACT_OK` or `TC_ACT_SHOT` in direct-action mode. So: XDP for volumetric drop and forwarding at the earliest possible point, tc when you need policy that depends on who sent the packet, or you need to touch outbound traffic at all.

code

bash · 6 lines
bash
# clsact provides ingress and egress attach points without queueing
tc qdisc add dev eth0 clsact
tc filter add dev eth0 ingress bpf da obj tc_prog.o sec classifier
tc filter add dev eth0 egress bpf da obj tc_prog.o sec classifier
tc filter show dev eth0 egress
bpftool net show dev eth0

go deeper

for a junior

Know that eBPF can be attached at more than one point in the network path, and that the tc hook is the one that also handles outbound traffic.

for a middle

Be able to contrast the two contexts concretely — raw bytes versus an sk_buff with mark, priority and ifindex — and name the tc action codes and the clsact attach points.

for a senior

Show you would place logic deliberately: floods and forwarding at XDP, context-dependent policy and egress at tc, with the metadata and driver-support constraints driving the split rather than a preference for the faster hook.

for a principal

Own the coexistence problem: multiple agents attaching to one interface, ordering and ownership guarantees, and whether your platform can depend on a kernel new enough for tcx across the whole fleet.

## Two hooks, two data structures The two hooks differ mainly in what has already happened to the packet when your code runs. An XDP program runs in the driver on a raw receive buffer. Its context, `struct xdp_md`, is little more than a pair of pointers into those bytes. Nothing has parsed the packet, nothing has classified it, and nothing has associated it with a socket, a process or a cgroup. A tc program runs after `sk_buff` allocation. Its context, `struct __sk_buff`, is a sanitised view of that structure and carries real metadata: `len`, `protocol`, `ifindex`, `mark`, `priority`, `tc_classid`, `cb` for passing values between programs, plus VLAN and tunnel fields. On egress you can additionally reach socket-level facts, which is how per-workload policy gets implemented. ## What the sk_buff buys, and what it costs The metadata is the point, and so is the cost. Allocating and initialising an `sk_buff` is precisely the work XDP exists to skip, so a tc drop is materially more expensive than an XDP drop. For a volumetric flood — millions of packets a second you want gone — that difference decides whether the machine survives. For a few hundred thousand packets a second of policy enforcement it is irrelevant, and the metadata is worth far more than the cycles. The skb also means you are working with the stack's own representation of the packet, so helpers can grow, shrink, encapsulate and checksum it with the stack's cooperation: `bpf_skb_store_bytes()`, `bpf_l3_csum_replace()`, `bpf_l4_csum_replace()`, `bpf_skb_adjust_room()`, `bpf_redirect()`. XDP has a narrower and lower-level set. ## Ingress, egress, and portability XDP is a receive-path hook. There is no general XDP hook on transmit, so anything you want to do to packets the host sends — outbound filtering, marking, encapsulation — is tc's job. XDP in native mode also depends on the driver implementing it. Drivers such as mlx5, ixgbe, i40e, bnxt and virtio-net do; plenty of others do not, and stacked virtual devices are their own story. tc BPF has no such requirement: it lives in generic kernel code, so it works everywhere, which matters when you are shipping something that has to run on hardware you have not seen. ## Verdicts and attachment tc programs return traffic-control action codes rather than XDP verdicts: `TC_ACT_OK` continues processing, `TC_ACT_SHOT` drops, `TC_ACT_REDIRECT` (via `bpf_redirect()`) moves the packet to another device, `TC_ACT_UNSPEC` falls through to the next filter. In *direct-action* mode — the `da` keyword — the program's return value is taken as the action verdict directly, instead of returning a classid that a separate action then interprets. Effectively everyone uses direct-action mode. Attachment goes through the `clsact` qdisc, which exists specifically to provide ingress and egress attach points without doing any queueing: ```bash tc qdisc add dev eth0 clsact tc filter add dev eth0 ingress bpf da obj tc_prog.o sec classifier tc filter add dev eth0 egress bpf da obj tc_prog.o sec classifier tc filter show dev eth0 ingress ``` Since kernel 6.6 there is also the `tcx` attach point, which gives tc programs proper `bpf_link` ownership semantics and a defined multi-program ordering, instead of the older model where two pieces of software could quietly fight over the same filter list. `bpftool net show` lists what is attached either way. ## Choosing Use XDP when the answer is "this packet should never have been built": flood drops, an L4 load balancer's forwarding path, anything measured in millions of packets per second. Use tc when you need to know something about the packet's context, when you need to act on egress, when the traffic volume does not justify the constraints, or when you cannot guarantee driver support. Systems that do both — Cilium being the well-known example — put the frontend drop and forwarding logic at XDP and the per-workload policy at tc.

  • What does the direct-action mode flag change about how a tc BPF program's return value is interpreted?
    Without it, the classifier is expected to return a class identifier, and a separate tc action object then decides what happens to the packet. With `da`, the return value *is* the verdict — `TC_ACT_OK`, `TC_ACT_SHOT`, `TC_ACT_REDIRECT` — so one program does both classification and action, with no second lookup and no separate action configuration.
  • You need to encapsulate outbound packets in a tunnel header. Why is tc the natural hook for that and not XDP?
    Two reasons. XDP has no egress hook at all, so outbound traffic never reaches it. And growing a packet is far easier on an skb: `bpf_skb_adjust_room()` makes room and cooperates with the stack's checksum and offload handling, whereas XDP's headroom is whatever the driver reserved.
  • Why does the tcx attach point introduced in kernel 6.6 matter for software that ships tc programs?
    The old `tc filter` model has no ownership: two agents attaching to the same interface can overwrite or reorder each other's programs, and a crashed agent leaves programs behind. tcx gives each attachment a `bpf_link` with defined lifetime and explicit relative ordering, so independent pieces of software can coexist on one interface predictably.

saying these in an interview costs you the question

  • Saying tc BPF also runs before the sk_buff
  • Thinking XDP has an egress hook too
  • Assuming tc needs driver support like XDP does
  • Believing tc programs return XDP_DROP or XDP_PASS
  • Treating XDP as strictly better because it is faster

context