East-west sensing on 300 nodes to catch an intruder's lateral hops: what grows fastest, and what does it consume?
answer
- no interior chokepoint to aim at
- points scale with nodes, not links
- interior volume dwarfs the edge
- copies compete with production on the uplink
- reduce on the node, ship verdicts
basics
~20 sSensor count grows with nodes, not links, and interior traffic volume dwarfs the perimeter feed you sized against. What it consumes first is node CPU and the uplink you backhaul copies over — the same link you are trying to monitor.
solid answer
~50 sA perimeter deployment is one or two aggregation points fed by a link whose bandwidth you already know. Interior sensing on a container platform inverts every term: capture points scale with node count, and the traffic to be inspected is the interior, which on a service-heavy estate is several times the north-south volume and is dominated by service-to-service calls, replication and health checks. The first thing that breaks is usually the node itself — capture and analysis CPU competing with scheduled workloads — and immediately after it the uplink, because shipping copies off the node doubles load on the link you are monitoring. Practical deployments therefore tier: full inspection on a small set of nodes or segments, cheaper header-level records on more, nothing on the rest, with filtering done on the node before anything is shipped. The honest conclusion is that full interior coverage is not reachable, so the output is a coverage map rather than a percentage.
go deeper
Know that interior traffic between services is normally much larger than internet-facing traffic, so perimeter sizing numbers do not carry over.
Explain why capture points scale with node count rather than link count, and why copying traffic off a node loads the uplink you are monitoring.
Demonstrate the ordering of failures — node CPU, then backhaul, then aggregation and storage — and design the reduction on the node so the ask becomes a bounded CPU budget.
Frame the outcome as tiers and a coverage map rather than a percentage, and be ready to defend asking the platform for a share of fleet capacity.
## The perimeter number does not transfer Perimeter sensing is easy to size because it is bounded by a link you bought: one or two aggregation points, a known bandwidth, one place to put the sensor. Every one of those properties fails inside. **Capture points scale with nodes.** There is no interior chokepoint on a container platform. Traffic between workloads either stays inside a node or crosses the fabric as node-to-node tunnels; neither passes a place you can put one sensor. So the count of capture points is the count of nodes, and it grows every time the platform scales out. Three hundred nodes is three hundred deployments, three hundred upgrade targets, and three hundred things that can silently stop. **Volume is the interior, not the edge.** North-south traffic is the small end of a service estate. Interior traffic includes service-to-service calls that fan out several deep per user request, database replication, cache traffic, health checks, service discovery, log and metric shipping. On a service-heavy platform it commonly runs several times the internet-facing volume, and none of the usual budgeting instincts account for it because nobody ever paid for it as bandwidth. **Most of it is worthless to inspect.** A large fraction of interior bytes is replication and telemetry between two components that will never be an intrusion path in the way you care about. Paying full inspection cost for those bytes is the fastest way to run out of budget before you have covered anything interesting. ## What breaks, in order 1. **Node CPU.** Capture, reassembly and matching run on the same cores the scheduler is handing to workloads. Platform owners notice this immediately, because it shows up as reduced schedulable capacity — you are asking for a percentage of the whole fleet. 2. **The uplink you are monitoring.** If the sensor ships raw copies off the node, interior traffic now traverses the fabric twice. On a busy node this is self-defeating: the copies compete with the production traffic you exist to protect. 3. **Aggregation and analysis.** Even filtered, three hundred feeds converge somewhere, and that somewhere is now sized for interior volume rather than edge volume. 4. **Storage.** Anything you want to be able to look back at multiplies retention by the same factor. ## The lever that actually works: reduce before you ship The design that survives does its reduction **on the node**, in the capture path, before anything crosses the uplink: - drop the known-uninteresting talkers by address or port at the hook itself; - keep full packets only for a defined set of paths, and summaries for the rest; - run detection locally where the logic is cheap and ship the verdict rather than the bytes. That converts an unbounded bandwidth problem into a bounded per-node CPU budget, which is a number the platform team can be asked to approve. ## Node density cuts both ways A densely packed node hides more traffic inside itself, so instrumenting it buys more coverage per sensor — and costs more CPU, because more of the estate's traffic is passing through that kernel. Sparse nodes are cheap to instrument and buy little. When you can only fund a fraction of the fleet, density is a real selection input, and it is the one people forget. ## Where this lands The arithmetic does not converge on full coverage; it converges on tiers. A workable shape is: full inspection on a small set of nodes or segments chosen deliberately, header-level records on a wider set, and an explicit nothing on the rest. That third category is not a failure to be hidden — it is the thing you have to write down, because an intruder whose entire lateral path stays inside one unsensed segment crosses nothing you instrumented, and the absence of any record from that path is what someone will later ask you to explain. A candidate who answers this by quoting a sensor's rated throughput has missed the question. The rated throughput is the last constraint you hit; node CPU and backhaul are the first two.
- Why is it wrong to size interior sensing from your internet link bandwidth?Because north-south is the small end. A single user request can fan out into many interior calls, and replication, caching, health checks and telemetry never touch the edge at all. Sizing from the internet link routinely underestimates interior volume by a large multiple, and the error is discovered after the hardware is bought.
- Why does backhauling raw copies over the production uplink make things worse?Because the copies traverse the same fabric as the traffic they duplicate, so interior load grows with the sensing rather than being observed by it. On a busy node the sensing becomes a contributor to congestion. The fix is to filter and summarise inside the node and ship far less than you captured.
- How does node density change which nodes you instrument first?Dense nodes hide more traffic inside their own kernel, so a sensor there converts more otherwise-invisible traffic into observable traffic — at higher CPU cost. Sparse nodes are cheap and buy little. With a limited number of sensors, density is a genuine selection input alongside what the workloads on the node are worth.
saying these in an interview costs you the question
- Sizes interior sensing from the internet link bandwidth
- Quotes a sensor's rated throughput as the binding constraint
- Plans to ship raw copies off every node
- Assumes an interior aggregation chokepoint exists
- Promises full east-west coverage across the fleet