In a Suricata rule, what must an attacker's packet satisfy before the engine scans payload bytes, and what does that ordering save?
answer
- cheapest test first
- tuple before bytes
- direction and session state gate the scan
- payload scanning is the expensive step
- a header matching everything is a CPU bill
basics
~20 sThe rule header - protocol, source and destination address and port, and direction - plus flow keywords are checked first. Only packets that pass reach the payload content match, so most traffic is discarded without any byte scanning.
solid answer
~40 sA signature is two parts. The header is a tuple: action, protocol, source address and port, an arrow giving direction, destination address and port. The body then adds pre-conditions such as `flow:established,to_server` before any `content` match. The engine groups rules by that header, so a packet is only ever considered against the group its tuple belongs to, and inside the group a multi-pattern pre-filter picks which rules are worth evaluating fully. That ordering is the whole performance model: payload scanning is the expensive step, and the header exists to make sure an attacker's packet only reaches it if it is already plausible. It also means the reverse is true - traffic on a port or direction the header does not cover is invisible to that rule no matter what bytes it carries.
go deeper
Be ready to read a rule aloud and point at the header, the direction arrow, the flow keyword and the content match, and to say which of them the engine tests first.
Explain the grouping and pre-filter mechanics: why a packet is a candidate for only a fraction of the rule set, and why payload scanning is the step everything else exists to avoid.
Show that you size a rule set by its cost per packet, and that you can name the coverage a narrow header silently gives away - the ports, directions and address pairs where the same payload is never inspected.
Own the trade as a capacity decision: broad headers buy coverage with CPU that turns into dropped packets, and someone has to state how much sensor the estate is funded to run.
## The two halves of a signature A Snort or Suricata rule reads as `action protocol source_addr source_port -> dest_addr dest_port (options)`. Everything left of the parenthesis is the **header**; everything inside is the **body**, where `msg`, `flow`, `content`, anchoring keywords, `pcre`, `sid` and `rev` live. The header is not documentation. It is the first and cheapest filter in the engine, and it determines whether the rule is even a candidate for a given packet. ## Why the order exists Comparing an address against a range and a port against a number is a handful of CPU cycles. Searching a payload for a byte string is orders of magnitude more expensive, and a sensor may hold tens of thousands of rules while packets arrive at line rate. So the engine spends the cheap tests first: 1. **Protocol and header tuple.** Rules are grouped by protocol, port and address set. A packet is matched against the groups its tuple selects; rules in other groups are never considered at all. 2. **Flow and state keywords.** `flow:established,to_server` restricts the rule to packets in a tracked session travelling toward the server. This throws away stray packets, backscatter, and the server's own responses. 3. **The multi-pattern pre-filter.** Across all rules in a group, one content string per rule (the fast pattern) is searched for in a single pass. Only rules whose fast pattern was seen are then evaluated in full. 4. **Full evaluation.** Now the remaining content matches, their anchoring keywords, and any `pcre` are run in the order written. The address variables in the header are usually named: `$HOME_NET` for the estate you defend, `$EXTERNAL_NET` for everything else, typically defined as `!$HOME_NET`. A rule written `$EXTERNAL_NET any -> $HOME_NET any` is asserting a direction of attack. ## What this buys the defender The cost the defender pays for a sensor is CPU per packet, and CPU that runs out does not produce an error - it produces dropped packets, and a dropped packet is a detection gap nobody sees. Header pre-conditions are how you keep that bill down. A rule set of twenty thousand signatures is affordable only because any single packet is a candidate for a small fraction of them. The inverse is the cost of getting it wrong. A rule whose header matches everything - `any any -> any any` - makes every packet a candidate, so the pre-filter and then full payload evaluation run far more often. Ten such rules are survivable; a rule set where the address variables have gone degenerate is a capacity problem. ## What this costs the defender in coverage The same ordering that saves CPU is also a blind spot an adversary can occupy for free. If the rule says `-> $HOME_NET 80`, the identical payload delivered to port 8443 never reaches the content match. If it says `flow:to_server`, the same bytes coming back from a compromised server are not evaluated. If `$EXTERNAL_NET` is defined as `!$HOME_NET`, an attack from one internal host to another is not external-to-home traffic and the rule sits idle. This is why an honest answer to "is that technique detected?" always includes the header. The rule matches that payload *in that direction, on those ports, between those address sets*. Outside that box it matches nothing, and the box is chosen by the rule author, not by the attacker. ## What an interviewer is listening for The common wrong answer is that rules are evaluated top to bottom like a firewall access list, first match wins. They are not: an IDS evaluates every rule that survives grouping and pre-filtering, and multiple rules can alert on one packet. The second wrong answer is that every rule scans every packet - if that were true, no sensor would keep up. Being able to say which tests are cheap, which are expensive, and which order they run in is what separates someone who writes rules from someone who has only enabled a vendor rule set.
- Why put `flow:established,to_server` in a rule at all if the content match would catch the same bytes?Two reasons. It keeps the expensive payload scan off stray packets, scan noise and backscatter that were never part of a real session. And it stops the rule firing on the server's own echo of the pattern in a response or an error page, which is a large share of the false positives on naive content rules.
- If a rule's header never matches any traffic on your sensor, what does it cost you?Almost nothing in CPU - it is eliminated at grouping and never reaches payload evaluation. It costs you in a more dangerous currency: it looks like coverage in a rule count and detects nothing. The expensive rules are the ones whose headers match everything, and the useless ones are the ones whose headers match nothing.
- Does a rule matching mean the packet is blocked?Only if the sensor is inline and the rule's action is a drop or reject. A rule with an alert action on a sensor fed by a mirror port records the event and the packet still arrives. The action word and the deployment position, not the rule body, decide whether an adversary is stopped or merely noticed.
Checking the header before the payload is like a mailroom sorting by address label before anyone opens an envelope. Opening every envelope would be accurate and would also stop the post.
saying these in an interview costs you the question
- Thinks rules are evaluated top to bottom, first match wins
- Believes every rule scans every packet's payload
- Cannot say what the arrow in the header specifies
- Treats the header address variables as documentation
- Assumes any rule match also blocks the traffic