Your organisation must send all VPC traffic through a fleet of third-party firewall appliances that need to inspect the original, unmodified packets. Why is an AWS Gateway Load Balancer the right component here rather than a Network Load Balancer?
answer
- addressed to it, or routed through it
- transparent bump in the wire
- route table target, not a hostname
- the original packet travels intact
- one flow, one appliance
basics
~20 sGateway Load Balancer is transparent: it is inserted as a route-table target and encapsulates the original packet in GENEVE to the appliance fleet, preserving the true source and destination. A Network Load Balancer is an endpoint clients must address directly, so it cannot sit invisibly in the path.
solid answer
~50 sAn NLB is a destination: clients connect to its address and port, and it forwards the flow to a target. That model cannot express "inspect everything on the way to somewhere else", because the traffic was never addressed to the load balancer. **Gateway Load Balancer** is built for exactly that. It is simultaneously a layer-3 gateway and a layer-4 load balancer: you place a Gateway Load Balancer endpoint in a subnet and point route tables at it, so traffic flows through it without any endpoint being aware. GWLB encapsulates each original packet in **GENEVE**, on UDP port 6081, and sends it to an appliance in the fleet; the appliance decapsulates, inspects, and returns the packet, which GWLB forwards to the real destination. Because the original packet is carried intact inside the encapsulation, the appliance sees the true source and destination addresses. Flow stickiness keeps a given flow pinned to one appliance so stateful inspection works.
go deeper
Know that Gateway Load Balancer exists for putting third-party network appliances inline, and that unlike NLB it is reached through routing rather than by connecting to its address.
Explain the mechanism: a Gateway Load Balancer endpoint as a route-table target, GENEVE encapsulation on UDP 6081 carrying the original packet to the appliance, and the return path back to the true destination.
Design the topology — an inspection VPC, endpoints or Transit Gateway attachments in the spokes, route tables that decide which traffic classes are inspected — and reason about flow stickiness and appliance health as production concerns.
Own the policy question: what must be inspected, whether the layer fails open or closed, who runs the fleet, and whether centralised inspection is worth the latency, cost and blast radius compared with distributed controls.
## The shape of the requirement "Everything must pass through inspection" is not a load balancing problem in the usual sense. The traffic in question is not addressed to the security fleet at all — it is a client talking to a server, and the appliances have to be inserted between them without either endpoint changing anything. That is a *bump in the wire*, and it is a fundamentally different insertion model from a load balancer that clients dial. ## Why NLB cannot do it An NLB is a destination. It has addresses, it has listeners on ports, and a client reaches a service by connecting to it. The load balancer then picks a target and forwards. Two properties make this unsuitable for transparent inspection: - **It must be addressed.** Nothing forwards traffic through an NLB on the way to a different destination. You cannot point a route table at it. - **It rewrites the destination.** The packet that arrives at the target is destined for the target. An inspection appliance that needs to know the original destination — the whole point of a firewall decision — has lost it. ## How Gateway Load Balancer works GWLB combines two roles in one component: **A layer-3 gateway.** You create a *Gateway Load Balancer endpoint* — a VPC endpoint type — in a subnet, and you make it the target of routes. Traffic sent toward a destination whose route points at that endpoint is diverted into the GWLB. Neither the client nor the server knows: this is expressed entirely in routing, which is why it can be imposed centrally on workloads that are never modified. **A layer-4 load balancer.** It distributes those diverted flows across a fleet of appliance targets, health-checks them, and adds and removes capacity like any other load balancer. The connective tissue is **GENEVE encapsulation on UDP port 6081**. GWLB wraps the entire original packet — headers and all — in a GENEVE tunnel to the chosen appliance. The appliance decapsulates, sees the packet exactly as the client sent it, applies its policy, and if it permits the traffic, re-encapsulates and hands it back. GWLB then releases the packet toward its real destination. Nothing about the packet's addressing has changed along the way. **Flow stickiness** matters as much as encapsulation. Stateful inspection only works if both directions of a conversation, and every packet within it, reach the same appliance — an appliance that saw only half a handshake has no basis for a decision. GWLB pins a flow to one appliance for its lifetime. ## The architecture that follows The usual production shape is a dedicated **inspection VPC** holding the GWLB and the appliance fleet, connected to workload VPCs through Transit Gateway or through GWLB endpoints placed in the workload VPCs themselves. Route tables in the spokes direct the traffic classes that must be inspected — often all egress, sometimes east-west between environments — at the endpoint. Security teams gain a single enforcement point that application teams cannot bypass, and appliance capacity scales horizontally instead of being a pair of hand-managed boxes. The costs are real and worth naming: extra hops add latency, the inspection fleet is a hard dependency in the path of production traffic, and you are paying both for the load balancer and for the appliance licences. The failure mode of an inspection layer is that everything stops, so its capacity and health-checking deserve the same care as the workloads it protects. ## Answering the interview version The crisp discriminator to state is insertion model. NLB is *addressed* — traffic goes to it. GWLB is *routed through* — traffic passes it on the way elsewhere, encapsulated so the appliance sees the original packet. If the requirement contains the words "transparent", "inline" or "the appliance must see the real addresses", the answer is GWLB. If clients can simply be pointed at an endpoint, you never needed it.
- How is a Gateway Load Balancer actually put into the traffic path?Through a Gateway Load Balancer endpoint placed in a subnet and referenced as the target of routes in a route table. Because the insertion is expressed in routing, workloads need no configuration change and cannot opt out — which is precisely why security teams like it. The client and server never learn that inspection happened.
- Why does flow stickiness matter for an inspection fleet?Stateful appliances build per-connection state, so every packet of a flow — in both directions — must land on the appliance that saw the handshake. If packets were spread across the fleet, no appliance would hold a complete view and policy decisions would be wrong or connections would be dropped. GWLB pins a flow to a single appliance for its lifetime.
- What are the operational costs of routing production traffic through such a fleet?You add hops and latency, you introduce a hard dependency whose failure stops all inspected traffic, and you pay for both the load balancer and the appliance licences. Treat the inspection layer as a tier-one production system: multi-AZ, scaled ahead of demand, health-checked carefully, and with an explicit, rehearsed decision about fail-open versus fail-closed.
saying these in an interview costs you the question
- Suggests pointing a route table at a Network Load Balancer
- Thinks appliances must be given the load balancer's DNS name
- Assumes the appliance sees a rewritten destination address
- Ignores flow stickiness for stateful inspection
- Treats the inspection fleet as a non-critical side path