skip to content

In a Wireshark display filter, when do you use `contains` rather than `matches`, and why is `http.response.code contains "50"` rejected?

level: middleimportance: should knowfreq 17%

answer

  1. substring versus regular expression
  2. case sensitivity differs
  3. PCRE2, double-quoted pattern
  4. not on numbers or addresses
  5. escape twice or use r""

basics

~20 s

contains looks for a literal string or byte sequence and is case-sensitive; matches applies a case-insensitive Perl-compatible regular expression. contains is refused on http.response.code because it cannot be used on atomic fields such as numbers; compare the number instead.

solid answer

~30 s

`contains` is a literal search: `http.user_agent contains "curl"` or `tcp.payload contains 47.45.54`, exact bytes, case included. `matches` (or `~`) takes a double-quoted PCRE2 pattern and is case-insensitive unless the pattern starts with `(?-i)`: `http.host matches "\\.example\\.com$"` catches every subdomain in any case. Backslashes are escaped twice in a normal string; a raw string, `r"\.example\.com$"`, avoids that. `contains` cannot be used on atomic fields such as numbers or IP addresses, so `http.response.code contains "50"` fails; numbers want `>= 500` or `in {...}`, and if you really need text, `string(http.response.code) matches "^50"`. Addresses want CIDR: `ip.addr == 192.0.2.0/24`.

go deeper

for a junior

Recall that contains is a literal, case-sensitive search and matches is a regular expression that ignores case by default.

for a middle

Explain the double escaping in quoted patterns, raw strings, and why contains is refused on numbers and addresses while comparisons and CIDR are not.

for a senior

Know what a frame-level search cannot see, such as decrypted data and tokens split across segments, and prefer field-level searches that carry less noise.

for a principal

Weigh broad frame searches against precise field filters when a team hunts for leaked secrets in captures, and what each misses.

## Two search operators, two jobs Wireshark's display-filter language has two operators that search inside a value rather than compare it whole. Both exist only in English form, apart from the `~` alias for `matches`. | | `contains` | `matches` / `~` | |---|---|---| | Right-hand side | a string or a byte array | a double-quoted regular expression | | Engine | literal substring or byte search | PCRE2, Perl-compatible | | Case | exact | case-insensitive by default | | Left-hand side | a protocol, a string or bytes field, or a slice | a string, or a field converted to one | | Typical use | a known token in a header or payload | a pattern with alternatives or anchors | ## `contains`: exact bytes or characters `contains` asks whether a value holds a given sequence. A double-quoted string is converted to its UTF-8 bytes when the left side is a byte field, so these two filters are equivalent: ``` tcp.payload contains "GET" tcp.payload contains 47.45.54 ``` It works on whole protocols too: `http contains "/api/v1/orders"` searches the bytes the HTTP dissector covers. It is **case-sensitive**. For a case-insensitive literal, wrap the field in `lower()` or `upper()`: `lower(http.server) contains "nginx"`. ## `matches`: regular expressions `matches` applies a **PCRE2** pattern (Wireshark 4.0 replaced GRegex with PCRE2). The pattern must be a double-quoted string, and matching is **case-insensitive by default**; prefix the pattern with `(?-i)` to make it case-sensitive. - `http.host matches "\\.example\\.com$"` keeps every subdomain of example.com, whatever the case. - `http.request.uri matches "^/api/v[12]/"` separates two API versions. - `http.user_agent matches "(?-i)^curl/"` keeps only the exact-case curl agent. The double backslash is the trap. Inside a normal double-quoted string, a backslash is an escape character, so a regex `\.` must be typed `\\.`. **Raw strings**, available since 3.6, keep backslashes literal: `http.host matches r"\.example\.com$"`. ## Why `http.response.code contains "50"` is rejected The manual says plainly that `contains` cannot be used on **atomic fields** such as numbers or IP addresses. `http.response.code` is an unsigned integer and `ip.src` an IPv4 address, so a substring search over them makes no sense to the engine. Use the type's own tools: 1. Numbers: compare them. `http.response.code >= 500` for every server error, `http.response.code in {502, 503, 504}` for a few. 2. Addresses: use CIDR. `ip.src == 192.0.2.0/24` rather than looking for "192.0.2" as text. 3. If you really need text, convert explicitly: `string(http.response.code) matches "^50"`. The `string()` function turns integers and addresses into their decimal or dotted text; it is not meant for string or byte fields, which are already searchable. ## What a search can and cannot see - `frame contains "password"` searches the captured bytes of each frame. The `frame` protocol does not include secondary data sources such as decrypted data, so a token inside a TLS session will not be found there even when decryption is working; filter on the decrypted protocol's fields instead. - A search sees one packet at a time. A token split across two TCP segments is only visible on a field the dissector built from the reassembled data. - Field-level searches are more precise than frame searches: `http.cookie contains "session="` will not fire on the same text in a request body. ## Functions that help searches A few filter functions make searches sharper: - `lower()` and `upper()` normalise case before `contains`: `lower(http.server) contains "apache"`. - `len()` returns the byte length of a string or bytes field: `len(http.request.uri) > 100` finds unusually long request URIs, a common sign of probing or of a client bug. - `string()` turns a number or address into text for `matches`, as in `string(ip.dst) matches r"\.255$"` for addresses ending in .255. Combine them with field-level tests rather than reaching for `frame` first; a narrow field keeps the result small enough to read. ## Choosing between them Use `contains` when you know the exact token and its case. Use `matches` when you need alternatives, anchors or case-insensitivity, and keep patterns anchored where you can so they say exactly what they mean. Use neither on numbers or addresses: compare those as numbers and subnets.

  • Your filter `http.host matches "example.com"` also keeps hosts such as exampleXcom.net. Why?
    In a regular expression `.` matches any character and the pattern is unanchored, so it matches anywhere in the string. Escape the dot and anchor it: `http.host matches "\\.example\\.com$"`, or with a raw string `r"\.example\.com$"`. If you only want a literal substring in exact case, `contains` is simpler.
  • How do you make a `contains` test case-insensitive?
    Lower-case the field first: `lower(http.user_agent) contains "python-requests"`. `contains` itself compares exactly, while `matches` is already case-insensitive, so the alternative is `http.user_agent matches "python-requests"`; escape any regex metacharacters if you go that way.

saying these in an interview costs you the question

  • contains ignores case, so CURL and curl match the same filter.
  • matches is case-sensitive unless you add a flag such as (?i).
  • contains works on any field, so ip.src contains "192.0.2" is fine.
  • A regex dot in a normal double-quoted string needs only one backslash.
  • frame contains finds text inside decrypted TLS payloads.