skip to content

A team replaces a connection-level (layer 4) proxy in front of an HTTP service with a request-level (layer 7) proxy. What extra work does the proxy now do, and what does that cost in CPU, memory and failure surface?

level: middleimportance: should knowfreq 58%

answer

  1. byte pipe becomes protocol participant
  2. handshake per connection, parsing per request
  3. state now scales with in-flight requests
  4. errors the application never logged
  5. limits the app team does not own

basics

~20 s

A layer 7 proxy terminates TLS, parses and validates every request, holds per-request state and buffers, and maintains its own upstream connection pool. That costs handshake and parsing CPU, memory per in-flight request, added latency, and a new component that can reject or time out requests itself.

solid answer

~50 s

The proxy stops being a byte pipe and becomes a protocol participant. Per connection it now runs the TLS handshake; per request it parses and validates the request line and headers, applies routing rules, possibly rewrites and buffers, re-emits the request on a pooled upstream connection, then parses the response on the way back. The costs are CPU for handshakes and parsing, memory proportional to in-flight requests and buffer sizes, and some added latency — it can no longer just splice two sockets together. The bigger change is the failure surface: the proxy now authors responses of its own, so you get 502s and 504s the application never logged, plus 400s and 413s when a request violates the proxy's own header or body limits. Streaming responses, WebSocket upgrades and unusual protocols all need explicit handling, and you now have timeouts at two layers instead of one.

go deeper

for a junior

Know that a request-level proxy reads and understands every request, unlike one that just forwards bytes, and that it can therefore return errors of its own such as 502 or 504.

for a middle

Be ready to separate per-connection cost (the TLS handshake) from per-request cost (parsing, rules, buffers), and to name the limits — header size, body size — that the proxy now enforces on the application's behalf.

for a senior

Demonstrate that you would size and monitor the tier as a service: concurrent in-flight requests and buffer memory, new-connection rate for handshake CPU, and proxy-authored status codes as a distinct signal from application errors.

for a principal

Frame it as a boundary decision: a shared request-level tier concentrates CPU, memory and failure into one hop that many teams' traffic depends on, and every capability it offers is paid for in that shared blast radius.

## What changes on the hot path A connection-level proxy does a small, fixed amount of work: accept a connection, choose a backend, open a connection to it, then copy bytes. Many implementations can hand that copying to the kernel, so the per-byte cost is close to nothing and the per-request cost is literally zero, because the proxy has no idea requests exist. A request-level proxy replaces that with a loop it runs for every single request: 1. complete the TLS handshake for the connection (asymmetric crypto once, symmetric per byte after); 2. read and parse the request line and header block, validating framing, size limits and protocol conformance; 3. evaluate routing rules against host, path, method or headers; 4. optionally rewrite the path, add or remove headers, buffer part or all of the body; 5. lease an upstream connection from a pool, re-emit the request on it; 6. parse the response headers, possibly buffer, transform or compress the body, and stream it back. ## CPU Two distinct costs get confused here. TLS handshakes are per **connection**, not per request, and they dominate when clients connect frequently and dominate not at all when connections are reused — which is exactly why keep-alive matters so much at an HTTPS edge. Parsing, header manipulation and rule evaluation are per **request**, and their cost scales with header count and rule-set size rather than payload size. Bulk symmetric encryption and any compression scale with bytes. ## Memory and per-request state A connection-level proxy needs roughly a socket pair and a small fixed buffer per connection. A request-level proxy holds a parsed header structure, routing context, logging context and read/write buffers per in-flight request — and with HTTP/2 there can be many concurrent requests on one connection. If it buffers request bodies (to enable retries, or to shield the backend from slow clients) that is memory or disk proportional to body size times concurrency. Capacity planning shifts from "how many connections" to "how many concurrent requests, of what size". ## Latency Parsing adds microseconds; buffering can add much more. If the proxy waits for a complete response before forwarding it, time-to-first-byte becomes time-to-last-byte upstream, which is fatal for server-sent events, long-polling, chunked progress output and large downloads. Any team that moves to a layer 7 tier and then finds their streaming endpoint "stopped streaming" has met response buffering. ## The new failure surface This is the part worth leading with in an interview, because it is what actually pages people: - **The proxy authors responses.** A layer 4 proxy in trouble can only reset a connection; it has no vocabulary for anything else. A layer 7 proxy returns a real HTTP status — a 502 when the upstream connection failed or the response was unparseable, a 504 when its own timeout fired, a 503 when no backend is available. These appear in the proxy's logs and never in the application's, and that asymmetry is the standard clue that the edge, not the app, produced them. - **The proxy enforces its own limits.** Maximum header size, maximum body size, maximum URL length, allowed methods, protocol strictness. Requests the backend would have accepted can now be rejected with 400, 413 or 431, and the limit is configured somewhere the application team may not own. - **Protocol handling must be explicit.** WebSocket and other upgrades need to be allowed through; gRPC needs HTTP/2 support end to end; anything the proxy does not parse simply cannot pass at this altitude. - **Timeouts and retries now exist twice.** The client has its own, and the proxy has connect, read and idle timeouts of its own; getting the relationship between them wrong is a whole failure class in itself. - **Two implementations must agree on framing.** A front-end proxy and a backend that disagree about how a message is delimited is a security problem, not just a bug. ## What you buy for the price All of that is worth paying when you need what only request visibility provides: routing on path or header, per-request retries and load spreading, rewriting, quotas, caching, compression, and per-request logs and metrics with status and latency. The mistake is paying it for traffic where none of those decisions are ever made — an opaque internal RPC stream gets the parsing bill and none of the benefit.

  • Your proxy's access log shows 502s that the application logs know nothing about. What does that tell you?
    That the proxy generated them. A 502 means the upstream attempt failed or its response could not be parsed — connection refused, reset mid-response, an upstream that died, or a protocol mismatch. Because the application never handled a request, it has nothing to log. Compare the proxy's upstream-connect and upstream-response metrics for the same second to see which half failed.
  • Why does the TLS cost of a layer 7 edge depend so heavily on client behaviour?
    Because the expensive asymmetric work happens once per connection, not per request. Clients that reuse a keep-alive connection for hundreds of requests amortise it to nothing; clients that open a fresh connection per request pay a full handshake every time, and the edge's CPU tracks new-connection rate rather than request rate.
  • A streaming endpoint works when called directly but delivers nothing until it finishes once it goes through the new proxy. What is happening?
    Response buffering. The proxy is accumulating the response before forwarding it, so time-to-first-byte for the client becomes the upstream's time-to-last-byte. Streaming responses, server-sent events and progress output need buffering disabled on that route, which also gives up the proxy's ability to retry once bytes have been forwarded.

saying these in an interview costs you the question

  • Saying the TLS handshake runs once per request
  • Believing a layer 7 proxy is 'just slower', with no state cost
  • Assuming every request the backend accepts also passes the proxy
  • Blaming the backend for 502s that only the proxy logged
  • Expecting streaming to work unchanged through a buffering proxy

context