skip to content

Segmenting a flat 4,000-host site into per-floor VLANs to contain an intruder hairpins east-west traffic through one firewall — what fills first?

level: seniorimportance: should knowfreq 50%

answer

  1. state before bandwidth
  2. the flow moves, it does not appear
  3. each flow crosses the trunk twice
  4. your flow records never saw it
  5. one unit carries it all during failover

basics

~20 s

Usually the firewall's session table and connection-setup rate, plus an uplink now carrying every flow twice. And you cannot size it from flow records: the traffic you are relocating never crossed a routed hop, so none was ever exported.

solid answer

~50 s

Splitting the domain does not create traffic, it relocates it: flows that were switched inside one VLAN now travel up a trunk to the routed boundary and back down, so they consume the uplink twice and land on a device sized for north-south internet traffic. What saturates first is normally the firewall's state table and its connection-setup rate rather than raw throughput, because east-west workloads are chatty — file shares, backup, imaging, printing, monitoring polls. Latency also rises on every one of those flows, which is what users actually report. The sizing trap is that your flow records understate the load badly: intra-VLAN traffic never crossed a routed interface, so no exporter ever saw it. You have to estimate east-west volume another way — endpoint or server-side counters, a pilot on one floor, or switch port statistics — and build the exception list before enforcement day, not after.

go deeper

for a junior

Understand the basic shape: once two hosts are in different subnets, their traffic must go up to the router and back down instead of being switched, so a device that was not in the path is now carrying it.

for a middle

Explain why state, not bandwidth, is the scarce resource — short-lived internal connections consume session-table entries and setup capacity — and why the uplink sees each hairpinned flow twice.

for a senior

Show the sizing method under uncertainty: name the flow-record blind spot, propose a single-floor pilot with the specific counters you would watch, and account for the HA failover case. Then present the exception list as a security artefact, not admin.

for a principal

Own the trade between the capacity bill and the containment gained, and be able to defend a phased rollout to whoever funds the hardware — including saying which floors go last and what risk stays open until they do.

## What segmentation moves Before the split, two hosts on different floors of one flat site exchanged frames through the switching fabric. After the split they are in different subnets, so every packet goes up to the routed boundary and back down. Three things change at once: 1. **The uplink carries each flow twice** — once up, once down — so a floor's east-west volume is doubled on the trunk before it reaches the boundary device at all. 2. **A device that was never in the path is now in it.** The perimeter firewall was sized against internet traffic: a few gigabits, a moderate number of long-lived sessions. Internal east-west traffic has a different shape. 3. **Latency appears where there was none.** A switched hop is microseconds; a hairpin through a stateful device adds enough that chatty protocols — file share metadata operations, imaging, database chatter — feel it as user-visible slowness. ## What fills first, and why it is usually not bandwidth The instinct is to compare aggregate throughput against the box's rated figure and conclude there is headroom. The resource that usually runs out first is state: the **session table** and the **connection-setup rate**. Internal workloads open enormous numbers of short-lived connections — monitoring polls, printing, directory lookups, backup agents, health checks — each of which must be created, tracked and expired. A firewall comfortably passing 3 Gbps of internet traffic in 60,000 sessions can fall over on 800,000 short internal sessions at a fraction of the bytes. Watch, in order: session table occupancy, new connections per second, CPU on the flow-setup path, then throughput. A second failure is capacity you did not know you had already spent: if the boundary is a high-availability pair, the surviving unit must carry the whole hairpinned load during a failover, so the real budget is one unit's, not two. ## The sizing trap this question is really about The confident wrong answer is "I'll size it from the flow records." **The traffic you are about to move is precisely the traffic your flow data never contained.** Flow export happens on routed interfaces; intra-VLAN traffic never reached one, so no record of it exists. Sizing from that data measures your north-south traffic and calls it the total, and the result is a firewall that is correct for the load you could see and undersized for the load you are creating. Honest alternatives, roughly in order of effort: - **Pilot one floor.** Segment a single floor first, measure the boundary device's session counts and setup rate under real use, and multiply by floors with a safety factor. This is also the only way to find the broken applications cheaply. - **Server-side counters.** Most east-west volume terminates on a countable number of servers — file, print, backup, imaging, directory. Their own connection and byte counters describe the load from the other end. - **Switch port statistics.** Coarse, but they tell you a floor's aggregate volume even when nothing knows the flow detail. ## The bill beyond the box Capacity is not the only price, and a senior answer says so: - **An exception list appears immediately.** "Floor 3 imaging must reach the floor 7 lab" is a real requirement, and every such rule is a permitted lateral path an intruder inherits. Write them with owners and end dates on day one, because they are impossible to prune later. - **Address assignment must be relayed** on each new segment, and each segment needs its own pool. - **A change window and a rollback plan per floor.** Enforcement day is when you learn which application had a hardcoded neighbour address. ## What you actually gained Say this too, because it is the justification for the bill: the containment number drops from 4,000 hosts to one floor's worth, and — the part often forgotten — the traffic between floors is now visible and loggable for the first time. You bought both containment and evidence, and you paid for them in capacity, latency and an exception list.

  • Why do you expect the session table to run out before throughput does?
    Because east-west workloads are chatty rather than heavy. Monitoring polls, printing, directory lookups and backup agents open large numbers of short-lived connections that each consume a state entry and a setup operation, while carrying few bytes. Throughput ratings are measured with large, long-lived flows and do not predict that behaviour.
  • How would you estimate the east-west load without any flow data for it?
    Pilot one floor and measure the boundary device directly — session count, new connections per second, CPU — then extrapolate with a margin. Cross-check against the server side, since most internal volume terminates on a small set of file, print, backup and directory servers whose own counters describe the same traffic from the other end.
  • What would you insist on capturing during the pilot beyond capacity numbers?
    The exception list. Every application that breaks becomes a permit rule, and those rules are the lateral paths that survive the project. Recording each one with an owner, a reason and an end date during the pilot is far cheaper than reconstructing four thousand rules later.

saying these in an interview costs you the question

  • Sizes the new boundary from existing flow records
  • Compares only aggregate bandwidth against the device rating
  • Forgets each hairpinned flow crosses the uplink twice
  • Ignores that one unit of an HA pair must carry the full load
  • Treats the exception list as paperwork rather than as permitted lateral paths

context