A purple-team replay of known-bad traffic raised no alert — how do you tell silent capture loss from a rule miss?
answer
- chain of delivery, not chain of rules
- ask each stage what it counted
- one stage counts nothing by design
- the connection record is the discriminator
- compare byte counts against what you sent
basics
~20 sWalk the delivery chain stage by stage and ask each one what it counted. If the sensor's own connection record for the replay is absent or truncated, packets were lost before matching; if the record is complete with the expected byte counts, the packets arrived and the detection is at fault.
solid answer
~60 sTreat it as a chain-of-delivery question, not a rule question. Four stages carry the replay: the mirror or tap, the aggregation broker, the sensor's NIC and capture ring, and the engine. Each has counters except the tap, which is passive and reports nothing at all, and the mirror engine, which on many platforms exposes no drop counter either — so the first two stages must be bounded by arithmetic and by the destination port sitting at line rate. The decisive evidence is the sensor's own connection record for the replay: absent, one-sided or short against the byte count you know you sent means capture loss; present and complete means the packets were matched against and the detection is what failed. Do that comparison against a source of truth outside the sensor — the switch interface counters for the same window, or the sending host's record of the transfer. And say the honest thing: a zero drop counter on the sensor does not prove zero loss, because it cannot see the stages above it.
code
text · 16 linesleaf switch, monitor session 1 (destination port at 10G)
destination tx: 9.98 Gb/s sustained for the replay window <- at ceiling
mirrored frames discarded: not counted on this platform
packet broker, sensor egress port 3
ingress 14.2 Gb/s egress 9.8 Gb/s discards 41,203,118
sensor NIC (interface receive statistics)
dropped 0 missed 2,884,110 per-queue: q3 2.88M, others 0
capture ring (AF_PACKET tpacket_stats)
tp_packets 4,102,338,915 tp_drops 512,006
engine
connection record for the replay: absent
alerts in the replay window: 0go deeper
Know that a missing alert has more than one cause, and that the first thing to establish is whether the sensor ever received the traffic.
Name the stages between the wire and the engine and say which of them keep counters — and that a passive tap keeps none at all, so that stage is bounded by arithmetic rather than measured.
Show the discrimination in practice: the sensor's own connection record against a known replay byte count, corroborated by a source outside the sensor, before any rule is opened.
Own the fact that this investigation is only possible if the counters, the retention and the scheduled replay were funded in advance, and be able to state an honest bound when a stage is unmeasurable.
## Why this is the question and not a rule-tuning question When a replay produces no alert, the competent-but-wrong reflex is to open the rule. It is the wrong first move, because a rule that never saw the bytes cannot be tuned into seeing them, and every hour spent on it is an hour spent confirming the wrong theory. The correct first move is to establish whether the packets were ever presented to the matching engine. ## The four stages, and which of them can be counted **Stage 1 — the mirror or the tap.** An optical tap is entirely passive: it splits light and reports nothing, ever. A mirror session on many platforms exposes no counter for frames its mirror engine discarded. This stage frequently cannot be measured directly. It is bounded indirectly: what the source link carried in the window (from the switch's ordinary interface counters, which are reliable) against the destination port's capacity. A destination port at line rate for any part of the window is a positive finding. **Stage 2 — the aggregation broker.** This one does count. Per-port ingress and egress rates and discard counters exist; the question is whether anyone collects them at a granularity finer than the five-minute poll that hides bursts. **Stage 3 — the sensor's NIC and capture ring.** Interface receive statistics distinguish frames dropped by the driver from frames the NIC missed because no buffer was free. Per-queue counters matter here: loss concentrated on one queue points at an uneven hash rather than a global capacity shortfall. The capture ring keeps its own statistics — packets accepted against packets dropped. **Stage 4 — the engine.** Packets and flows processed, its own internal drops, and the connection log. ## The evidence that actually decides it The sensor's own **connection record for the replay** is the discriminator, because it is written by the same process that would have matched: - **No record at all** — the traffic never reached the engine. Look upstream. - **A record present but one-sided**, or with byte counts far below what you know you sent — partial delivery. Loss upstream, or a direction that never converged on this worker. - **A record complete, with byte counts matching what you sent** — the packets were there and were matched against. Now the problem is genuinely the detection: the rule, the protocol parser, the port assumption, an exception, a disabled ruleset. Corroborate against something outside the sensor. The switch's interface counters for the replay window, the sending host's own record of what it transferred, or the receiving service's log all give you an independent number of bytes to compare against. ## The direction of every claim in this investigation - A zero capture-drop counter proves that **nothing was lost after arrival at the NIC**. It says nothing about the mirror or the broker. - A silent tap proves nothing whatsoever; silence is its normal state. - An absent alert proves that **no rule fired**, which is compatible with loss, with a rule that did not match, with a rule that was disabled, and with a genuinely quiet estate. - A present, complete connection record is the one artefact that lets you say the sensor **had** the data. That is why it, and not the drop counters, is the thing to reach for first. ## What answering this at all costs you None of the above is available retrospectively unless you were already paying for it: the stage counters must be collected continuously, at a sub-minute granularity that survives bursts, and retained long enough to look back at the window in question. That is a metrics-collection and storage cost that has to be argued for before the incident, and it is invariably argued for after. The replay capability itself is a cost too — a scheduled, repeatable, known-payload exercise with a known byte count, run at a known time, is what makes the comparison possible at all. Running it only at average-load hours is a wasted budget: schedule it into the peak window, because the peak window is where the loss lives. ## What you say when you cannot prove it Sometimes the answer is that the mirror stage is unmeasurable on this platform and you cannot distinguish the two causes for last quarter's window. Say that plainly, then state what you can bound — the link carried X, the destination port could carry at most Y, so at least X minus Y was never offered to the sensor — and put the instrumentation gap on the record as a finding of its own. An honest bound is a better answer than a confident story built on a counter that does not measure what the story needs it to.
- The connection record exists and the byte counts match what you sent. What do you conclude?That delivery is not the problem: the packets reached the engine and were matched against. The failure is in detection — the rule itself, a protocol parser that did not recognise the traffic, a port or direction assumption in the rule, a suppression or exception, or a ruleset that is not loaded. That is now a legitimate rule-tuning investigation, which it was not before this evidence existed.
- Capture loss shows up almost entirely on one receive queue. What does that suggest?Not a capacity shortfall but an uneven distribution: the hash is landing a disproportionate share of traffic on one queue, often a few very large flows, and that queue's worker cannot keep up while the others idle. Adding overall capacity will not help much. Look at how traffic is being spread across queues and at whether a small number of heavy flows should be filtered upstream instead.
- What should the replay itself look like so this comparison is possible next time?Known payload, known byte count, known five-tuple, run at a recorded time — and scheduled into the busy window, not a convenient quiet afternoon. Without a byte count you know, a truncated connection record is indistinguishable from a small transfer, and the whole discrimination collapses. Repeat it on a schedule so a regression is caught by trend rather than by the next incident.
saying these in an interview costs you the question
- Opens the rule before proving the packets arrived
- Reads a zero sensor drop counter as zero loss on the feed
- Expects a passive optical tap to report discards
- Treats no alert as evidence the traffic was benign
- Polls stage counters every five minutes and calls bursts absent