A protocol-parsing sensor is fed thousands of short sessions by an adversary - why does it exhaust memory on session count, not bandwidth?
answer
- cost is per flow, not per bit
- concurrency plus attached analyzers
- reassembly and file buffers on top
- at the cap it sheds records, not packets
- eviction order nobody deliberately chose
basics
~20 sState is held per connection and per file in flight, not per bit, so ten thousand tiny sessions cost far more table space than one large transfer. At the ceiling the sensor sheds records, not traffic.
solid answer
~50 sA parsing sensor allocates a state record for every tracked flow, plus reassembly buffers and analyzer state for every stream it is decoding and every file it is extracting. That cost scales with concurrent sessions and with how many analyzers attach, not with link utilisation - which is why a flood of short connections, or a scan that opens a connection per host, is far more expensive than a single saturating transfer. When the cap is reached the engine does not stop: it expires or stops tracking flows, abandons reassembly, or skips file extraction. In a tap position that means you silently lose *records*, not packets, and the loss is worst exactly where you are busiest. The dangerous part is that the shedding policy - idle-first, oldest-first, or simply refusing to track anything new - was almost never chosen deliberately, so nobody can say which evidence went missing during the hour it mattered.
go deeper
Know that a sensor keeps a record for every connection it is watching, so many small connections cost more than one big one. Bandwidth is not the unit that fills it up.
Explain the three consumers - flow state, stream reassembly, analyzer and file buffers - and what each does when its cap binds. Be specific that shedding produces missing records rather than dropped traffic on a tap.
Show that you monitor the sensor's own counters and treat a cap-exceeded window as a declared evidence gap. Interviewers listen for whether you can bound what you lost rather than assume you lost nothing.
Be ready to say which facts you will stop collecting, in writing, so the trade is one you chose rather than one an eviction policy made for you at the worst moment.
## Where the memory actually goes A sensor that parses protocols pays in three places, and none of them is proportional to bandwidth. 1. **Flow state.** One record per tracked connection: endpoints, timers, counters, TCP sequence bookkeeping. It exists from the first packet until an inactivity timeout or a clean close, so its total size is *concurrent flows multiplied by per-flow size*. 2. **Stream reassembly.** To hand an analyzer an ordered byte stream, the engine buffers out-of-order and unacknowledged segments. A stream with a large gap holds its buffer until the gap closes or the engine gives up. 3. **Analyzer and file state.** Each attached analyzer keeps its own parse state - a half-read header, a pending request awaiting a response - and file extraction buffers an object while it is hashed or typed. Bandwidth touches only the second and third of these, and only briefly. Concurrency touches all three, continuously. ## Why the arithmetic surprises people A single 10 Gbps transfer is one flow. Ten thousand connections of 4 KB each are ten thousand flow records, ten thousand analyzer attachments, ten thousand file-extraction decisions - at a fraction of the throughput. That is why a horizontal scan, a chatty microservice pattern, or an adversary who deliberately fans a transfer across many short sessions will push a sensor into its caps long before the link is full. It is also why sizing a sensor on "we have a 10 Gbps link" is the wrong unit; the unit is concurrent sessions at peak, plus how many analyzers you have turned on. ## What happens at the ceiling, and why it is quiet Engines expose caps on flow tables and reassembly memory precisely because the alternative is being killed by the operating system. When a cap binds, the engine sheds work, and the shed comes in recognisable shapes: - **Refuse to track new flows.** The newest sessions become invisible above the transport. This is the shape that hurts during a flood, because the flood's own sessions are the new ones. - **Expire idle or oldest flows early.** This is the shape that hurts on an extranet, because your long-running partner transfers are exactly the idle-looking, oldest sessions - the ones you most wanted a full record of. - **Abandon reassembly or file extraction.** The connection record survives, the file record never appears, and the log looks like a session in which nothing was transferred. All three are recorded, if at all, in the engine's own counters rather than in your detection output. Nobody gets an alert saying "your evidence for the last forty minutes is partial". ## Records, not packets - and the position that decides it This is the fence that separates an out-of-band sensor's state price from a firewall's. A firewall or an inline IPS that runs out of connection-table space has to make a **traffic** decision: drop, or admit without inspection. A sensor on a tap or mirror port makes an **evidence** decision: the traffic is already on its way regardless, and what you lose is the ability to say later what it was. That difference changes who you argue with. A dropped session is a phone call from the business within minutes; a missing file record is discovered weeks later by an investigator who cannot tell whether the absence means nothing happened or means the sensor was full. ## The adversarial version An adversary who has learned where your sensor sits and roughly how it is sized does not need to defeat a rule. Generating enough concurrent, cheap sessions to push the sensor to its caps is enough, because the shedding is untargeted: it damages visibility across the whole estate, including the one session that carries the actual transfer. Nothing about this requires the adversary to touch the sensor. Detecting the condition is therefore a monitoring problem about the sensor itself, not about traffic. ## How to answer well Say the unit out loud: concurrent sessions and attached analyzers, not gigabits. Then say what the engine does at the ceiling and which of your own visibility that costs, and finish with the operational consequence - that you must alarm on the sensor's own drop and memory-cap counters, and treat a period in which they were non-zero as a known gap when someone later asks what happened, rather than reading a quiet log as a quiet estate.
- How would you know, a month later, that the sensor was over its cap during an incident window?Only from the sensor's own health telemetry - flow-table and reassembly memory-cap counters, packet-loss counters, and the rate of connection records per minute against baseline. Those must be collected and alarmed like any other production metric. If they were never recorded, the honest answer to "was anything missed" is that you cannot tell, which is a much worse position than knowing you lost forty minutes.
- What is the cheapest lever to reduce state cost without giving up the extranet visibility you care about?Narrow what the sensor is asked to do rather than what it sees: turn off file extraction and hashing for high-volume flows you have no intention of examining, restrict which analyzers attach, and shorten inactivity timeouts so dead flows release state. Each of those is an explicit, written-down loss of a specific fact, which is exactly what an implicit eviction policy is not.
saying these in an interview costs you the question
- Sizes a sensor by link speed instead of concurrent sessions
- Thinks an out-of-band sensor drops traffic when memory fills
- Assumes eviction is safe because flows were idle
- Reads a quiet log as evidence of a quiet estate
- Believes turning on all analyzers is free once the sensor is bought