Which single-byte change from an attacker defeats a Suricata content match anchored with depth, and what does widening the anchor cost?
answer
- offset and depth are absolute
- distance and within are relative to the last match
- byte-exact unless told otherwise
- one byte of padding moves everything
- widening is paid in CPU and dropped sessions
basics
~20 sOne byte of padding pushes the pattern past the depth window, and one changed letter defeats a match written without nocase. Widening the anchor means scanning the whole buffer on every packet and accepting more false positives, which on an inline sensor drops real traffic.
solid answer
~50 sAnchoring keywords bound where a pattern may appear. `offset` says start searching N bytes in, `depth` says search only the first N bytes, and after a previous match `distance` skips N bytes while `within` requires the next match inside N bytes. Each bound is an assumption about the attacker's layout. Prepend one byte and a pattern that had to start at offset 0 now starts at 1, so a `depth` equal to the pattern length misses it. Change one letter's case and a content match without `nocase` misses it, because the normalised URI buffer decodes percent-encoding and traversal but does not fold case. Insert one separator and a `distance`/`within` pair between two contents breaks. The fix is not free: dropping `depth` scans the whole buffer per packet, `nocase` and `pcre` cost more per byte, and the extra matches become false positives - which on an inline sensor means dropped business traffic, not just a noisy alert.
code
text · 9 linesalert http $EXTERNAL_NET any -> $HOME_NET any ( \
msg:"admin login path"; \
flow:established,to_server; \
http.uri; content:"/admin/login"; depth:12; \
sid:1000042; rev:1; )
... matches: GET /admin/login?next=/
... misses: GET /ADMIN/login (no nocase; the URI buffer is not case-folded)
... misses: GET /app/admin/login (pattern now starts past the depth window)go deeper
Know what offset, depth, distance and within each bound, and that a content match is byte-exact unless nocase is present.
Explain how each anchor encodes an assumption about the attacker's formatting, and demonstrate two edits that preserve meaning while breaking the match.
Show the trade in both directions: what a narrow rule misses, what a widened rule costs per packet, and why a false positive on an inline sensor is a business outage rather than noise.
Be ready to argue that a rule ships with its evasion margin stated, and to explain who accepts the residual risk when widening it would cost more capacity than the estate is funded for.
## What the anchoring keywords actually assert A `content` match on its own says "these bytes appear somewhere in this buffer". That is cheap to write and expensive to run, and it matches far too much. The anchoring keywords narrow it: - **`offset:N`** - do not start searching until N bytes into the buffer. - **`depth:N`** - search only the first N bytes (measured from the buffer start, or from `offset` if one is given). - **`distance:N`** - relative to the end of the *previous* content match, skip at least N bytes before looking. - **`within:N`** - relative to the end of the previous match, this content must be found within the next N bytes. `offset` and `depth` are absolute; `distance` and `within` are relative and only make sense chained after another content. Together they turn a vague "contains" into a statement about structure: this token is the first thing in the request line, this parameter follows that one closely, this magic value sits at a fixed offset in the header. That precision is what makes a rule cheap and specific. It is also exactly the surface an adversary works on, because every bound is a prediction about the attacker's formatting, and the attacker controls the formatting. ## The one-byte moves **Padding shifts the offset.** A rule written to catch a token at the start of a buffer with `depth` equal to the token length has asserted position 0. One extra byte in front - a leading slash, a space, a duplicated separator, an extra path segment - and the token now begins at position 1. The pattern is still there in full, and the rule does not fire. **Case defeats a match written without `nocase`.** Content matching is byte-exact by default. Suricata's normalised HTTP URI buffer decodes percent-encoding and resolves directory traversal, but it does not lowercase the path. `/ADMIN/login` therefore carries the same meaning to most application stacks and a different byte string to the rule. **A separator breaks a relative chain.** Two contents joined by `within:4` assert that the second follows the first closely. An inserted byte, a doubled delimiter, or extra whitespace where the parser tolerates it pushes the second match outside the window. **Encoding changes the bytes without changing the meaning.** Whether the sensor sees the encoded or decoded form depends on which buffer the rule is written against. A rule on a raw buffer sees what was on the wire; a rule on a normalised buffer sees what the parser produced. Writing against the wrong one is a common way to produce a rule that passes a lab test and misses in production. None of this is clever. That is the point of the question: the cost of an anchored rule is that it is defeated by trivial, meaning-preserving edits, and the author must know which ones. ## What loosening costs The naive answer is "so remove `depth` and add `nocase`". Both have a price, and it lands on the sensor and on the business: | Change | What it buys | What it costs | | --- | --- | --- | | Remove `depth` | Matches the pattern anywhere in the buffer | Full-buffer search on every candidate packet | | Add `nocase` | Survives case changes | Case-insensitive search is more expensive per byte | | Add a `pcre` for variants | Catches many spellings at once | The most expensive test in the engine; must stay anchored behind a cheap content or it runs constantly | | Widen to a shorter content | Harder to shift out of range | Short, common strings make weak fast patterns, so far more rules get fully evaluated | And the second bill is false positives. On a passive sensor, a false positive costs someone's attention. On an inline sensor with a drop action, a false positive drops a legitimate session - a payment call, a build fetching a dependency, a customer's upload. That asymmetry is the reason a rule author cannot simply widen everything until nothing evades. ## The honest posture A well-written content rule is deliberately narrow and its narrowness is documented. The author should be able to say: this matches the technique when the token appears in this buffer within this window in this case; it does not match padded, re-cased or differently-encoded variants; here is why widening was rejected. That statement is worth more than an unqualified claim of coverage, because the person consuming it can decide what else to fund. The corollary is that a single content rule is a narrow instrument, not a control. It is at its best when it is cheap, specific and known-narrow, sitting alongside detections built on different observations entirely.
- Why is adding a pcre for every variant not the answer?A regular expression is the most expensive test the engine runs, and one that is not anchored behind a cheap content match runs on far more packets than it should. The usual pattern is a short, distinctive content that pre-filters, with the pcre only refining what survives. Replacing structure with regex is how a rule set stops keeping up with the wire.
- How do you decide between the raw buffer and the normalised one?Ask what the target will act on. If the application sees the decoded, traversal-resolved form, write against the normalised buffer so encoding tricks collapse before you match. If the point is to catch the encoding itself as an anomaly, write against the raw buffer. A rule tested against one and deployed to inspect the other is a classic silent miss.
- Should the rule's known evasions be written down, or is that just admitting weakness?Written down. A narrow rule with a stated margin is usable - someone can decide whether the gap matters and what else to fund. An unqualified rule gets recorded as coverage, and the gap is then discovered by an adversary rather than by a reviewer.
An anchored content match is a guard told to read only the first twelve characters of a badge. One space typed in front of the name and the badge reads clean.
saying these in an interview costs you the question
- Thinks content matching is case-insensitive by default
- Confuses depth with distance, or offset with within
- Removes all anchors to be safe, ignoring the CPU bill
- Assumes the normalised URI buffer lowercases the path
- Treats a false positive on an inline sensor as merely noisy