What is the difference between an L4 (transport-layer) and an L7 (application-layer) load balancer, and what routing or scaling decisions can an L7 balancer make that an L4 balancer cannot?
answer
- L4 = TCP/UDP tuple, blind to payload
- L7 = terminates + parses HTTP
- path/header routing needs L7
- L7 costs a TLS handshake + extra hop
- layered: L4 edge, L7 behind it
basics
~20 sAn L4 load balancer only looks at IP addresses and TCP/UDP ports and just forwards packets, without knowing what's inside. An L7 load balancer reads the actual HTTP request (URL, headers, cookies) and can route based on that content, but doing so costs more CPU and adds latency.
solid answer
~50 sL4 load balancing operates at the transport layer: it sees source/destination IP and port and forwards or NATs packets or TCP connections without parsing the payload, so it's fast, protocol-agnostic (works for any TCP/UDP traffic), and cheap in CPU terms. L7 load balancing terminates the connection, parses the actual application protocol (typically HTTP), and can then make decisions based on URL path, headers, cookies, or request body - enabling path-based routing (e.g. /api/* to one service, /static/* to another), content-based rewriting, TLS termination, and cookie-based sticky sessions. The trade-off is that L7 requires terminating and re-establishing connections (adding latency and CPU cost for TLS/HTTP parsing) and needs to understand the specific application protocol, whereas L4 is essentially protocol-blind. Real systems commonly layer both: an L4 balancer (e.g. a cloud Network Load Balancer) in front for raw throughput and DDoS absorption, with L7 balancers (e.g. an ALB, NGINX, or Envoy) behind it doing smart routing.
go deeper
Should know L4 works on IP/port and L7 works on HTTP content like URLs, at a basic level.
Should explain path/header-based routing as an L7-only capability and name TLS termination as an L7 cost.
Should discuss the layered architecture pattern (L4 edge + L7 behind it) and articulate the latency/throughput trade-off concretely.
Should reason about DDoS posture, TLS passthrough vs termination trade-offs, and how choosing L4 vs L7 at each tier shapes the overall system's failure modes and capacity planning.
## Where the labels come from The 'L4' and 'L7' labels come from the OSI networking model: layer 4 is the transport layer (TCP/UDP), and layer 7 is the application layer (HTTP, gRPC, etc.). ## What an L4 balancer can see An L4 load balancer makes its routing decision using only information available at the transport layer, without ever looking inside the packet payload: - source IP, - destination IP, - source port, - destination port, - and protocol. Mechanically, this is often implemented as connection-level forwarding or even NAT/DSR (direct server return): the balancer picks a backend for a new TCP connection (or, for UDP, a flow) based on a hash of the connection tuple, and then simply relays packets back and forth, or in DSR setups, only handles the inbound direction while the backend replies directly to the client. Because it never parses the payload, an L4 balancer is **protocol-agnostic** - it works identically for HTTP, a database wire protocol, a custom TCP service, or arbitrary UDP traffic - and it is very fast, since forwarding packets is cheap compared to parsing them. ## What an L7 balancer unlocks An L7 load balancer sits one layer higher: it actually terminates the client's connection (including TLS, if used), reads the parsed application-layer request, and only then decides which backend should handle it, typically opening a separate connection to that backend (or reusing one from a pool). Because it understands the protocol, it can inspect the HTTP method, URL path, query string, headers, and cookies, and route on any of them. This unlocks capabilities an L4 balancer simply cannot offer: - **path-based routing**, where requests to `/api/orders/*` go to an orders microservice and `/static/*` go to a CDN-backed static file service, all sharing one public hostname; - **header- or cookie-based routing**, used for canary releases (route 5% of requests carrying a beta cookie to the new version) or A/B testing; - **content-based rewriting** and redirect rules; - **cookie-based session affinity**, where the balancer injects a cookie identifying which backend handled the first request and pins subsequent requests from that client to the same backend. L7 balancers can also do things like request retries and circuit breaking on 5xx responses, since they can actually see the response status, something an L4 balancer cannot do because it never decodes the payload. ## What that capability costs The cost of that capability is real. Terminating a connection means the L7 balancer must do a full TLS handshake with the client (unless TLS passthrough is used, which sacrifices most L7 features), buffer and parse the HTTP request line and headers, and then originate a fresh connection (or reuse a pooled one) to the chosen backend - effectively doubling the number of TCP connections in play and adding CPU work for parsing and, if applicable, TLS. This adds latency, typically low single-digit milliseconds for a well-tuned proxy but nontrivial under very high request rates, and it means the L7 balancer itself becomes a more resource-intensive piece of infrastructure that needs its own capacity planning. It also means the L7 balancer sits directly in the data path for every request rather than just the initial connection setup, so bugs or misconfiguration in it (a bad routing rule, an overloaded proxy process) can break traffic in a much more visible, protocol-aware way, such as returning malformed 502s instead of silently dropped packets. ## Why production layers both Because of this trade-off, production architectures very often layer both. - A cloud **Network Load Balancer** (a pure L4 device, such as AWS's NLB) sits at the outermost edge because it can handle extremely high packet-per-second throughput with minimal latency and survives volumetric attacks well, and it forwards traffic to a tier of L7 balancers or reverse proxies (an AWS ALB, NGINX, HAProxy in HTTP mode, or an Envoy/Istio service mesh sidecar) that do the smart, content-aware routing to individual microservices. - **Kubernetes Ingress controllers** are a familiar concrete example of L7 balancing: the Ingress resource defines path- and host-based routing rules ('/orders goes to the orders-service, /users goes to the user-service'), and the underlying controller (NGINX Ingress, Envoy-based Contour, etc.) implements them by terminating HTTP and dispatching per request. ## The rule of thumb A good rule of thumb when choosing: if you just need to spread raw TCP/UDP load across identical backends serving the same thing, L4 is simpler, cheaper, and faster; the moment routing needs to depend on what's inside the request - the URL, a header, a cookie, the response code - you need L7.
- Can an L7 load balancer also do TLS passthrough instead of termination, and what do you give up if it does?Yes - in TLS passthrough mode the L7 balancer forwards the encrypted bytes untouched based on SNI (the hostname in the TLS ClientHello, which is visible even though the rest is encrypted) and lets the backend terminate TLS itself. You keep host-based routing but lose everything that requires reading the decrypted HTTP request, such as path-based routing, cookie-based affinity, or header inspection.
- Why might an L4 balancer be preferred specifically for DDoS resilience?Because it does far less per-packet work - no TLS handshake, no HTTP parsing - it can absorb a much higher packet rate before it becomes the bottleneck, and volumetric attacks are fundamentally about overwhelming the cheapest layer that will fail first. Cloud providers' network-layer DDoS protection is typically built on this kind of high-throughput L4 infrastructure before traffic ever reaches an L7 proxy.
- How does an L7 load balancer typically implement path-based routing internally?It maintains a routing table mapping URL path prefixes (or hostnames, headers, etc.) to backend target groups, and on each incoming request it parses the request line, matches the path against that table, and forwards the request to a chosen backend from the matched group - often applying a secondary algorithm like least-connections within that group.
An L4 balancer is like a mail sorting machine that routes packages purely by the zip code on the envelope, never opening them. An L7 balancer is like a person who actually opens each envelope, reads the letter inside, and routes it to the right department based on what it says.
saying these in an interview costs you the question
- Says L4 can route based on URL path
- Thinks L7 balancers don't add any latency or resource cost
- Can't explain that L7 means terminating the connection
- Believes L4 and L7 are interchangeable terms for 'load balancer'
- Doesn't know TLS termination is an L7 concern