skip to content

On a trunk-port capture, why does the tcpdump filter `vlan and host 192.0.2.10 or host 192.0.2.10` miss untagged traffic to that host, and how do you fix it?

level: seniorimportance: should knowfreq 10%

answer

  1. first vlan moves every later offset
  2. applies to the rest of the expression
  3. untagged test goes first
  4. 4 bytes per tag

basics

~20 s

The first vlan keyword shifts the decoding offsets by 4 bytes for the rest of the expression, so the second host test reads untagged frames at the wrong place. Put the untagged test first: host 192.0.2.10 or (vlan and host 192.0.2.10).

solid answer

~40 s

In pcap-filter, the **first `vlan` keyword changes the decoding offsets for the remainder of the expression**, on the assumption that the packet carries an 802.1Q tag, and each further `vlan` adds another 4 bytes. That shift is applied in parse order, not just to the primitives ANDed with `vlan`. So in `vlan and host 192.0.2.10 or host 192.0.2.10`, the second `host` test is compiled to look 4 bytes further in, and an untagged frame fails it. The fix is ordering: `'host 192.0.2.10 or (vlan and host 192.0.2.10)'` compiles the untagged test with normal offsets before `vlan` shifts anything. `vlan 100 and vlan 200` matches VLAN 200 inside VLAN 100. `mpls`, `pppoes`, `geneve` and `vxlan` shift offsets the same way.

go deeper

for a junior

Recall that vlan is a special keyword and that filters on trunk links need to handle tagged frames explicitly.

for a middle

Explain that the first vlan shifts offsets by 4 for the rest of the expression and that each extra vlan adds another 4.

for a senior

Write mixed tagged and untagged filters with the untagged test first, verify them with tcpdump -d, and know Linux may carry the tag as metadata.

for a principal

Where trunks and overlays like VXLAN are common, standardise filter templates per capture point so encapsulation never silently hides evidence.

## What the `vlan` keyword does In libpcap's **pcap-filter** language, `vlan [vlan_id]` is true if the packet is an IEEE 802.1Q VLAN packet, and with an id only if the tag carries that VLAN ID. Unlike `host` or `port`, it has a side effect on the compiler: - The **first** `vlan` keyword encountered **changes the decoding offsets for the remainder of the expression**, assuming the packet is a VLAN packet. - **Each** further `vlan` keyword increments the offsets by another 4 bytes, the size of one tag, so stacked tags can be matched. The effect is positional. libpcap's compiler applies it to everything parsed after the keyword, not only to the primitives logically ANDed with it. A comment in libpcap's `gencode.c` says so directly: `(vlan and ip) or ip` checks only for VLAN-encapsulated IP, not for IP over plain Ethernet as well. ## Why the example misses untagged traffic Take `vlan and host 192.0.2.10 or host 192.0.2.10` on a trunk carrying both tagged and untagged frames: 1. `vlan` tests for a VLAN TPID and shifts every later offset by 4. 2. The first `host 192.0.2.10` is compiled with the shifted offsets - correct for tagged frames. 3. The second `host 192.0.2.10` is **also** compiled with the shifted offsets, because it comes later in the expression. 4. An untagged frame fails `vlan`, falls to the `or` branch, and is tested 4 bytes off: the bytes read are not the IPv4 addresses, so it does not match. The filter compiles cleanly and catches tagged traffic, so the gap is easy to miss. ## The fix: untagged first, tagged second ``` host 192.0.2.10 or (vlan and host 192.0.2.10) ``` The first `host` is compiled before any `vlan` keyword, so it uses normal Ethernet offsets. Then `vlan` shifts, and the second `host` uses tagged offsets. The rule of thumb: **in any expression mixing tagged and untagged tests, write the untagged part first**. | Expression | What it actually matches | |---|---| | `host X or (vlan and host X)` | X untagged, or X inside one tag | | `(vlan and host X) or host X` | X inside one tag only | | `vlan 100 and vlan 200` | VLAN 200 encapsulated within VLAN 100 | | `vlan and vlan 300 and ip` | IPv4 in VLAN 300 inside any outer VLAN | | `vlan or host X` | every tagged frame, plus X tested 4 bytes off | ## What the compiled code checks `tcpdump -d 'vlan 4095'` on Ethernet shows the shape. It loads the EtherType at byte 12 and accepts any of three tag protocol identifiers - **0x8100** (802.1Q), **0x88a8** (802.1ad) and **0x9100** - then loads the tag at byte 14, masks it with `0xfff` and compares the VLAN ID. Only the VLAN ID is compared, not the priority bits. On Linux there is a twist. The kernel often removes the outer tag from the packet data and keeps it as metadata. When the kernel supports the BPF extensions for that, libpcap (since 1.9) emits code that checks the metadata first (`ldb [vlanp]`, `ldh [vlan_tci]` in `-d` output) and falls back to the in-packet tag, deciding at run time. You still write `vlan`; the compiled program covers both cases. ## Other keywords with the same side effect - `mpls [label]` shifts offsets by 4 per label, and a `vlan` test after `mpls` is rejected ("no VLAN match after MPLS"). - `pppoes [session_id]` shifts offsets to the PPPoE session payload. - `geneve [vni]` and `vxlan [vni]` shift offsets into the encapsulated packet; `vxlan` arrived in libpcap 1.11.0. The same ordering rule applies to all of them: put the tests on the outer, unencapsulated packet first, then the keyword, then the tests on what is inside. ## Practical checks - Confirm whether a capture point sees tags at all before writing VLAN filters: `tcpdump -e` prints link-level headers, including the 802.1Q tag. - Compile the expression with `tcpdump -d` and no `-i`, which gives plain Ethernet code without the Linux metadata checks, and check that the untagged branch loads the IPv4 addresses at bytes 26 and 30 while the tagged branch uses 30 and 34. - When a host seems silent on a trunk, suspect the filter's ordering before suspecting the network.

  • How would you match IPv4 traffic in VLAN 300 when it may be nested inside any outer VLAN tag?
    `vlan and vlan 300 and ip`. The first `vlan` matches any outer tag and shifts offsets by 4; `vlan 300` then tests the inner tag's VLAN ID and shifts by another 4; `ip` checks the EtherType after both tags. Order matters because each keyword moves the offsets used by everything after it.
  • On a Linux host, the NIC removes VLAN tags before the packet reaches the capture socket. Does a tcpdump `vlan 10` filter still work?
    Usually yes. When the kernel supports the BPF extensions for VLAN metadata, libpcap compiles `vlan 10` to check the tag in packet metadata first and fall back to an in-packet tag, so both cases are handled at run time. In `tcpdump -d` output this shows as loads from `[vlanp]` and `[vlan_tci]`.

The first vlan keyword is like sliding a ruler 4 centimetres along a page: every measurement you write down after that uses the moved ruler, even for pages that were never shifted. Measure the unshifted pages before you slide it.

saying these in an interview costs you the question

  • Believes vlan only affects primitives ANDed with it
  • Writes (vlan and host X) or host X to catch both cases
  • Thinks vlan 100 or vlan 200 matches either tag on one frame
  • Concludes a host is silent on a trunk without checking filter order
  • Assumes mpls or vxlan leave later offsets unchanged