Why does a switch-port mirror miss an intruder pivoting between two containers on the same node, and what does capturing it cost?
answer
- no wire means nothing to copy
- the hop stays in the node kernel
- virtual interface pair, not a switch port
- vantage has to move onto the host
- one sensor per node, on production CPU
basics
~20 sTwo containers on one node exchange packets inside that node's kernel, so no frame ever reaches the switch a mirror copies. Seeing that hop means running capture inside every node, paid in CPU on the machines running production workloads.
solid answer
~50 sA mirror copies frames a switch forwards, and a tap copies what crosses a cable. Traffic between two workloads scheduled on the same node is switched or routed by that node's own kernel across virtual interfaces and never touches the physical NIC, so neither device has anything to copy. An intruder who lands in one workload and pivots to a co-located one therefore crosses no monitored point at all — the segment sensor and the boundary firewall both record nothing, which is not the same as recording an allow. The only vantage that sees it is inside the node: a capture hook in the kernel data path at the virtual interface. That costs a sensor on every node, CPU and memory taken from the same budget the workloads are scheduled against, and a component sitting in the data path of production traffic.
go deeper
Be ready to say plainly that a mirror or tap can only copy what crosses a switch or a cable, and that two containers on one node never put a frame there.
Explain the mechanics: virtual interface pairs, kernel bridging or routing inside the node, and a capture hook in the node data path as the only vantage that sees the packet.
Show the production judgment: a sensor per node costs CPU beside the workloads and sits in the data path, and placement decided by the scheduler makes any per-service coverage claim unstable.
Own the framing that silence from a sensor over a path it cannot observe is not evidence, and that the honest output is a per-node coverage statement someone accepts as residual risk.
## Where the packets actually go On a container platform each workload runs in its own network namespace with a virtual interface whose other end lives in the node's kernel. When two workloads on the **same node** talk, the node's kernel bridges or routes the frame from one virtual interface to the other. The packet is real, it is TCP, it carries the intruder's session — and it never leaves the machine. A switch-port mirror (SPAN) copies frames the **switch** forwards. A tap copies signal on a **cable**. Neither exists for a hop that stayed inside one host. This is a property of the vantage, not a misconfiguration: no number of extra mirror sessions, no better sensor, and no larger analysis cluster changes it. ## Why this is a security statement, not a plumbing detail Workload placement is decided by the platform's scheduler for reasons that have nothing to do with security — bin-packing, resource requests, node pools, spread constraints. A densely packed node can host dozens of workloads belonging to several teams. So on any real platform a meaningful share of east-west conversations are same-node conversations, chosen for you, and they are exactly the hops an intruder uses first: land in one workload, reach whatever is nearest. The consequence for a coverage claim is sharp. When your interior sensor shows no lateral traffic, that proves **the sensor saw none**. It does not prove none happened. A sensor's silence about a path it structurally cannot observe is not evidence of absence, and an interviewer will push on exactly that distinction. ## What it costs to fix The fix is to move the vantage from the wire onto the host: a capture hook in the node's kernel data path, attached at the virtual interface pair, using the kernel's traffic-control or eBPF hooks. It sees the workload's packet in plaintext before any overlay encapsulation on egress and after decapsulation on ingress. The bill is specific and someone has to fund it: | what you buy | what you pay | | --- | --- | | one capture point per node | 300 nodes means 300 sensors to deploy, upgrade and monitor | | plaintext inner packets | CPU and memory drawn from the same node budget the workloads are scheduled against | | coverage of same-node hops | a component in the data path of production traffic — a bad version is a production incident, not just a sensing outage | | copies you can analyse | backhaul of the copied bytes, often over the same uplink you are already monitoring | That last row is why per-node capture is usually filtered or summarised on the node before anything is shipped. ## The coverage claim is unstable by construction Because the scheduler moves workloads, visibility is a property of **nodes**, not of workload pairs. Two services whose conversation was visible on the wire yesterday may be co-located tomorrow and disappear, with nobody changing a sensor. Any statement of the form "we can see traffic between A and B" is only true for a given placement. The honest statement is per node: "these nodes carry a capture hook, these do not." ## Partial answers that are worth offering - **Constrain placement.** If a set of workloads must be sensed, schedule them only onto instrumented nodes. You are buying coverage with scheduling flexibility rather than with hardware. - **Accept a different record source.** Host and workload records are not network capture and prove different things, but for some segments they are the only evidence that will exist. Say which it is. - **Write the gap down.** The deliverable of an interior sensing programme is a dated list of what is not sensed, not a claim of full coverage. ## The wrong answers The common failure is to assume the overlay pushes everything onto the wire — that an encapsulated network means every packet becomes a node-to-node tunnel packet the top-of-rack switch can mirror. Encapsulation happens on the way **out** of the node; a same-node hop never gets there. The second failure is treating host capture as free because it is software.
- Would a tap on the node's uplink cable solve it?No. A tap on the uplink sees only what the node sends onto the wire, which is cross-node traffic. A same-node hop is bridged or routed inside the kernel and never reaches the NIC, so the tap has nothing to copy. It helps for node-to-node lateral movement and not at all for co-located workloads.
- Why does the scheduler make a coverage statement unstable?Because visibility follows placement. Two workloads split across nodes today are visible on the wire; co-located tomorrow, they are not, and nothing in the sensing changed. That is why coverage must be stated per node — which nodes carry a capture hook — rather than per pair of services.
- If you cannot put a sensor on every node, what is the cheapest useful move?Constrain placement: pin the workloads you must be able to evidence onto the subset of nodes that do carry a capture hook, and let everything else schedule freely. You pay in scheduling flexibility and node utilisation instead of in sensors, and the coverage claim becomes true by construction rather than by luck.
A camera in the corridor films everyone who changes rooms, and nothing that happens inside one room.
saying these in an interview costs you the question
- Claims a top-of-rack mirror covers all container traffic
- Assumes the overlay forces every packet onto the wire
- Thinks more mirror sessions fix same-node blindness
- Treats an empty sensor view as proof no lateral movement happened
- Treats host capture as free because it is software