You are deciding whether to front a service with connection-level (layer 4) or request-level (layer 7) proxying. How do you make that call, and what situations make layer 4 the right answer even though layer 7 is more capable?
answer
- list the decisions first
- per connection or per request
- parsing implies decrypting
- not everything is HTTP
- layer them, don't pick globally
basics
~20 sChoose by the decisions you need to make per request. If routing, rewriting, per-request retries or quotas depend on reading the request, you need layer 7. If the payload is not HTTP, the proxy tier may not hold the key, or throughput and latency dominate, layer 4 is correct.
solid answer
~60 sI start from what the traffic layer has to decide, not from which product is nicer. If the decision is per request — route `/api` here and `/static` there, shift a weighted slice to a canary, retry an individual call, enforce a per-tenant quota, emit per-request latency and status — then the layer must be able to read requests, and I accept the termination point, the CPU and the extra failure surface that come with it. If the decision is per connection, or the payload is not HTTP at all — a database or broker protocol, an opaque binary RPC — layer 7 buys nothing and costs parsing that cannot even happen. Three constraints push me to layer 4 regardless of capability: the proxy tier is not allowed to hold the service's private key, the latency or throughput budget will not tolerate an extra parsing hop, or the protocol is one the proxy does not speak. In practice large systems run both: a connection-level edge for raw scale and non-HTTP traffic, request-level proxying close to the services that need it.
go deeper
Be able to say that layer 7 is needed when routing depends on the request itself, and that layer 4 is the choice for traffic that is not HTTP or that the proxy must not decrypt.
Explain the trade concretely: request visibility buys routing, rewriting, retries and per-request metrics, and costs a termination point, parsing CPU and an extra place requests can fail.
Demonstrate that you would decide per traffic class and per constraint — protocol, key custody, latency budget — and describe a layered arrangement rather than one global answer.
Own the organisational and reversibility angles: who edits the routing surface, how big its blast radius is, and why moving up an altitude later is far cheaper than unwinding capabilities built on request visibility.
## Start from the decisions, not the products The useful question is never "L4 or L7?" in the abstract. It is: **what does this layer have to decide, and how often?** Write the list down. If it contains anything that requires knowing what a request said, the layer must read requests. If it does not, reading requests is pure cost. Decisions that force layer 7: host, path, method or header routing; weighted traffic shifting between versions; per-request load spreading; per-request retries and timeouts; header injection and URL rewriting; response caching or compression; per-client quotas; and per-request telemetry (status, path, upstream latency) — which is often the real reason teams want it, and a perfectly good reason. Decisions that live happily at layer 4: which backend gets this connection, is that backend alive, and should this connection be admitted at all. ## Constraints that override capability Some situations settle the question before preferences enter: **The protocol is not HTTP.** PostgreSQL, MySQL, Redis, Kafka, SMTP, MQTT, custom binary RPC. A request-level proxy for HTTP simply cannot parse these, and forcing them through it means running it in a pass-through mode — which is layer 4 with extra steps. Some proxies do speak specific non-HTTP protocols at the request level; that is a per-protocol capability question, not a general one. **The proxy tier may not terminate.** If a compliance requirement, a customer contract or a key-custody boundary says the traffic must stay encrypted until it reaches the service, no layer 7 routing is possible on that path, because parsing implies decrypting. Re-encrypting upstream does not rescue it — the proxy still decrypted in the middle. Routing must then be made from connection-level facts such as the SNI name. **The budget will not take the hop.** Very high throughput per node, or a latency budget where a parsing and buffering hop is a meaningful fraction, argues for connection-level forwarding, which can approach kernel-speed copying. **Nothing per-request is being decided.** A failover front end for a database pool needs to know which primary is live. That is a connection decision. Adding an HTTP-aware tier gives it nothing to be aware of. ## Layer them rather than pick globally Mature systems rarely make one choice. A common shape: a connection-level edge that absorbs volume and handles non-HTTP listeners, then request-level proxying nearer the workload where routing and retries are decided — sometimes as a shared tier, sometimes as a per-workload sidecar. The virtue of that split is that the expensive, opinionated, frequently-reconfigured layer sits close to the team that owns it, while the cheap, stable layer takes the internet. The corresponding failure is a request-level tier deployed "for consistency" in front of traffic that never needed it: it pays parsing cost for opaque payloads, imposes header and body limits on protocols that do not have headers, and inserts a shared failure domain for no capability gained. ## Organisational cost is a real input At platform scale the choice is partly about who edits what. A shared request-level tier is a config surface that many teams change, a shared blast radius when a bad route is pushed, and a place where a bug affects everyone. That argues for keeping the shared tier as simple as the requirements allow, and pushing per-service request logic outward — to per-service proxies or sidecars — so that the change frequency of the routing rules matches the blast radius of the component holding them. ## Reversibility Moving from layer 4 to layer 7 later is usually straightforward: the addresses stay, you insert a parsing hop. Moving the other way is not, because everything built on request visibility — routing rules, weights, retries, quotas, dashboards — has to be re-homed. So it is fair to start at layer 4 when the requirement list is empty today, and fair to start at layer 7 when you can name the request-level decision you already need. What is not fair is asserting that layer 7 is "more advanced" and therefore better; the altitude that matches the decisions is the better one. ## How to answer this in an interview Name the decisions, name the constraint that could veto the choice (protocol, key custody, budget), state where you would put each layer, and say explicitly what trade you accepted — the parsing cost and the shared failure domain if you took layer 7, the loss of per-request routing, retries and telemetry if you took layer 4.
- A team wants path-based routing but is contractually forbidden from decrypting the traffic. What do you tell them?That the two requirements are mutually exclusive at that hop. Reading a path means terminating TLS, and re-encrypting upstream does not change that the proxy saw the plaintext. Either the routing key moves to something visible at the connection level, such as a hostname per route with SNI-based selection, or the request-level decision moves inside the trust boundary, to a proxy the service owner runs.
- When would you deliberately avoid a shared request-level tier in front of internal services?When it becomes a config surface many teams edit with one blast radius. If per-service routing rules change weekly, I would rather each service own its own request-level proxy or sidecar and keep the shared tier dumb and stable, so the change frequency of the rules matches the blast radius of the thing holding them.
- How does the choice change for internal service-to-service traffic rather than internet-facing traffic?Internal callers are known and few, so hostname-based edge concerns matter less, while per-request retries, balancing and telemetry matter more — which pushes toward request-level, often as a sidecar rather than a shared hop. Internal traffic is also more likely to be multiplexed RPC, where connection-level balancing distributes load poorly.
saying these in an interview costs you the question
- Choosing layer 7 because it is 'more advanced'
- Assuming re-encrypting upstream avoids decrypting at the proxy
- Forcing non-HTTP protocols through an HTTP-aware tier
- Ignoring that a shared layer 7 tier is one blast radius
- Treating the choice as global rather than per traffic class