skip to content

Your XDP program on a Linux host is dropping flood packets with XDP_DROP, but `tcpdump -i eth0` on that host shows none of those packets at all. Why can tcpdump not see them, and where does its own filter actually run?

level: middleimportance: nice to knowfreq 32%

answer

  1. the capture tap needs an sk_buff
  2. one hook runs before the tap, one after
  3. the filter runs per socket, not on the NIC
  4. its return value is a byte count
  5. counters in a map are the instrument

basics

~20 s

tcpdump reads packets from an AF_PACKET socket that is fed later in the receive path, after the driver has built an sk_buff. Native XDP runs before that, so packets it drops never reach the tap. Count drops in a BPF map instead.

solid answer

~50 s

tcpdump does not read the wire; it opens an `AF_PACKET` socket, and the kernel copies packets to that socket from a tap point in the shared receive path — after the driver has handed the packet to the stack. A native XDP program runs earlier than that, inside the driver, so a packet it drops is freed before anything can be copied to the tap. That is a feature, not a bug: it is the same reason the drop is cheap. The filter you type is compiled by libpcap into classic BPF, attached to that socket with `setsockopt(SO_ATTACH_FILTER)` and translated by the kernel into eBPF, and its return value is the number of bytes of each packet to keep — zero meaning ignore this packet. So the filter runs per packet delivered to the socket, not on the NIC. To see what XDP dropped, keep counters in a map and read them with `bpftool map dump`, or temporarily return `XDP_PASS`.

code

bash · 11 lines
bash
# The filter expression compiled to classic BPF, before attachment
tcpdump -d -i eth0 'tcp port 443'

# XDP-dropped packets never reach that tap - read the program's own counters
bpftool prog show
bpftool map dump name drop_count

# Is the program running at all?
sysctl -w kernel.bpf_stats_enabled=1
bpftool prog show
ethtool -S eth0 | grep -i xdp

go deeper

for a junior

Know that tcpdump captures from inside the kernel rather than from the wire, so code running earlier in the path can remove a packet before any capture tool can show it.

for a middle

Be able to say where the capture tap sits relative to XDP and to tc, explain that the filter is compiled BPF attached to an AF_PACKET socket, and name a way to count XDP drops instead.

for a senior

Use the tap position as evidence during an incident — captured but not delivered points at a different hook than never captured — and instrument programs with map counters and driver statistics so the datapath is observable without changing its behaviour.

for a principal

Set the expectation for the fleet: once packet handling moves into the driver path, capture tools stop being the source of truth, so the datapath must ship its own counters and the on-call runbooks must be rewritten around them.

## What tcpdump is actually doing Running `tcpdump -i eth0 'tcp port 443'` does three things. It opens an `AF_PACKET` socket bound to the interface, so the kernel will copy packets destined for the stack to it. It compiles the expression `tcp port 443` with libpcap into a classic BPF program — you can see the instructions with `tcpdump -d` — and attaches it to the socket with `setsockopt(SO_ATTACH_FILTER)`. Then it reads matched packets and prints them. The filter therefore runs in the kernel, per packet, on the way to that one socket. Its position matters twice over: it saves the copy to userspace for packets that do not match, which is why filtering in the expression is far cheaper than piping everything to userspace and grepping — and it is only ever reached by packets that make it to the tap in the first place. The return-value convention of a socket filter is a genuine oddity worth knowing: it is not a boolean. It returns the number of bytes of the packet to accept, so returning 0 discards the packet for that socket and returning a large value keeps the whole thing. That is exactly how the snap length is applied, and modern kernels translate the classic program into eBPF internally before running it. A program compiled as eBPF can be attached the same way with `SO_ATTACH_BPF`. Critically, none of this affects other consumers. A socket filter decides what *this* socket receives; the packet continues to the rest of the stack regardless. ## Why the XDP drop is invisible Order of operations settles it. In native mode the XDP program runs in the driver's receive routine, before an `sk_buff` exists. The AF_PACKET tap that feeds tcpdump lives further along, in the shared receive path, and needs that `sk_buff`. A packet that XDP drops is freed at the driver, so there is nothing for the tap to copy, and no socket filter ever runs against it. The same reasoning explains the cost saving. If tcpdump could see XDP-dropped packets, the kernel would have had to build the very per-packet state XDP exists to avoid. ## tc is not the same This is where the two hooks diverge in an operationally visible way. The AF_PACKET ingress tap is fed before the ingress hook where tc BPF programs run, so a packet that a tc ingress program drops has *already* been copied to tcpdump. If you are debugging and you can see the packet in tcpdump but the application never receives it, an ingress tc program dropping it is entirely consistent with what you are seeing; an XDP program dropping it is not. On the transmit side the mirror image holds. The tap for outbound packets sits at the hand-off to the driver, after the queueing layer, so a packet dropped by a tc *egress* program never appears in tcpdump output either. So "tcpdump sees it" is a real piece of evidence about where in the path something happened — as long as you know where the tap sits relative to the hook you are asking about. ## How to observe XDP instead Since the packet is gone before any tap, instrument the program itself: * **Counters in a map.** Increment a per-CPU array or hash map keyed by drop reason, then read it with `bpftool map dump id <id>` or from your own userspace loader. This is the standard approach and costs almost nothing. * **`XDP_ABORTED` for genuine errors**, which fires the `xdp_exception` tracepoint — useful precisely because it is out-of-band from your counters. * **Driver statistics.** Many drivers export XDP counters through `ethtool -S <iface>`, which is independent of your program and therefore a good cross-check. * **Program run statistics.** Enabling `kernel.bpf_stats_enabled` makes `bpftool prog show` report run count and accumulated run time per program, which tells you whether the program is being invoked at all — the first thing to establish when you suspect it is attached but not running. * **Temporarily pass.** In a controlled test, return `XDP_PASS` on the packets you would drop and capture them normally. You lose the performance property while you do it, which is the point of doing it only briefly. The habit to build is that once you put code in the driver path, the ordinary capture tools stop being the source of truth for what arrived. The program's own counters become the instrument.

  • A packet appears in tcpdump but the listening application never receives it. Does that rule out a BPF program dropping it?
    It rules out a native XDP program, which runs before the capture tap. It does not rule out a tc ingress program: that hook runs after the tap, so a packet can be captured and then dropped. Check `tc filter show dev <iface> ingress` and `bpftool net show` before concluding the application is at fault.
  • Why does putting the filter expression in the tcpdump command matter more than piping the output through grep?
    The expression is compiled to BPF and attached to the socket, so non-matching packets are discarded in the kernel and never copied to userspace. Capturing everything and filtering afterwards pays the copy, the userspace wake-up and the formatting for every packet on the interface — which on a busy link can drop packets and disturb the very system you are debugging.
  • A socket filter's return value is a byte count rather than a boolean. What is that used for?
    Truncation. Returning N accepts the first N bytes of the packet, which is how the snap length works — capture the headers you need without copying full payloads. Returning 0 means this socket ignores the packet entirely. Other consumers of the packet are unaffected either way; the filter only governs this socket.

saying these in an interview costs you the question

  • Assuming tcpdump sees everything on the wire
  • Thinking the capture filter runs on the NIC
  • Believing a socket filter drops the packet system-wide
  • Treating an empty capture as proof nothing arrived
  • Expecting tc ingress drops to be invisible too

context