skip to content

Where does response compression belong in a middleware chain, and what breaks when it is installed too deep?

level: seniorimportance: nice to knowfreq 42%

answer

  1. it wraps bodies, not requests
  2. more producers than just the handler
  3. misplacement raises no error
  4. byte counters and signatures sit outside
  5. buffering stalls a streamed response

basics

~20 s

Compression wraps the response on its way out, so it belongs outside every element that can produce a body — handler, error stage and static assets alike. Placed too deep it silently misses responses produced outside it.

solid answer

~40 s

A compression hook works on the outbound half of the chain: it delegates, then wraps whatever body comes back. Its depth therefore decides which responses it covers. Put it deep — inside the error stage, or inside the element that serves static assets — and everything those produce goes out uncompressed, which shows up as an unexplained bandwidth pattern rather than as an error. Put it near the outside and it covers handler responses, error responses and short-circuited responses alike. Two constraints pull the other way. Anything that must measure or sign the bytes actually sent has to sit outside it, because compression changes both length and content. And a hook that buffers in order to compress will delay a long-lived streamed response, so streaming routes usually opt out.

go deeper

for a junior

Understand that compression happens to the response as it leaves, not inside the handler, and that a hook can only affect responses that travel out through it.

for a middle

List the elements that can produce a body — handler, error stage, short-circuiting hooks, asset serving — and explain why a compressor inside any of them silently misses that source.

for a senior

Bring the failure modes: no error is raised, a streamed endpoint appears hung because of block buffering, and length-sensitive or signing hooks must be hoisted outside the compression point.

for a principal

Decide the layer: compress at a shared edge or in the application, never accidentally in both, and set the thresholds, excluded content types and streaming exemptions as one policy rather than per service.

## A response-side hook's depth is about coverage Most discussion of the middleware chain is about the inbound journey, but compression is almost entirely an outbound concern. The hook delegates immediately, and does its work as the response travels back out. That makes its placement rule simple to state: **a response-side hook only affects bodies that pass through it, so it must wrap every element that can produce one.** In a typical chain, bodies come from more places than people remember: - the matched handler, on the happy path; - the error stage, when something throws; - an element that short-circuits — a limiter's rejection, a redirect, a cached response; - an element that serves static assets or packaged resources. A compression hook placed inside any of those never sees what that element produced. ## The symptom is a non-symptom This is what makes the misplacement interesting in an interview: **nothing fails.** An uncompressed response is a completely valid response. The client decodes it without complaint, tests pass, and the only evidence is quantitative — egress higher than expected, large error payloads or large asset responses that are inexplicably not shrinking, a gap between what a synthetic check reports and what real traffic costs. Teams find it months later while investigating a bandwidth bill. | Compression hook placed | Compresses | Misses | |---|---|---| | Near the outside, wrapping everything | handler, error stage, short-circuits, assets | nothing inside the chain | | Inside the error stage | handler responses | every error response | | Inside the asset-serving element | handler responses | asset responses | | Around one route group only | that group | every other route | ## What must stay outside it Compression changes two properties of the response: its byte length, and the bytes themselves. So anything downstream that depends on those has to be placed with care: 1. **A byte-counting metric or billing hook** must sit outside compression if it should report wire bytes, and inside if it should report payload bytes. Both are legitimate; picking one by accident is not. 2. **A hook that computes a digest or signature over the body** has to be unambiguous about whether it covers the compressed or uncompressed form, and the verifying side must agree. 3. **A hook that sets an explicit content length** must not run after the length has changed. ## Streaming is the real exception Compressors work on blocks: they accumulate input until they have enough to emit efficiently. For a normal response that is invisible. For a long-lived response that is supposed to deliver small events as they occur, it is fatal — each write is held in the compressor's buffer, and the client sees nothing for seconds or until the buffer fills. The endpoint looks hung while behaving exactly as configured. The usual resolutions are placement decisions in themselves: attach compression to the route groups that return ordinary payloads and leave streaming routes out; or use a compressor configured to flush at each write, accepting a worse ratio. What does not work is leaving a buffering compressor over a streaming route and treating the delay as a client problem. ## Other placement details worth knowing - **Do not compress twice.** If an element serves assets that were compressed ahead of time at build time, a compressor wrapping it must recognise that the body is already encoded and pass it through. Double encoding produces a body no client will decode. - **Small bodies are not worth it.** A size threshold below which the hook does nothing is standard; compressing a 40-byte error payload costs more than it saves. - **Already-compressed media gains nothing.** Images, video and archives are typically excluded by content type. - **The edge may already do it.** If a shared layer in front of the service compresses, doing it again in the application is wasted work — and worse, one of them will be doing it on already-encoded bytes. ## How to reason about it in an interview Start from the question *which elements produce a body?*, place the hook so it wraps all of them, then ask *what downstream of this point cares about length or content?* and hoist those out. That reasoning generalises: every response-side concern — decoration, sanitisation, buffering, size limits — is placed by the same two questions, and compression is simply the example where getting it wrong produces no error at all.

  • Where should an element that serves static assets sit relative to the authentication hook?
    Only outside it when the assets are genuinely public. An asset element placed outside authentication short-circuits before any credential check, so anything it can reach is served to anonymous callers. Private files behind such an element are a disclosure bug, not a caching optimisation, and the safe default is to keep it inside.
  • A byte-count metric reports much smaller responses than the bandwidth bill suggests. What placement would explain it?
    The counter sits inside the compression hook, so it measures payload bytes before encoding, while the bill reflects wire bytes. Neither number is wrong; they answer different questions. Move the counter outside compression to measure what left the machine, or label the two measurements distinctly and keep both.

saying these in an interview costs you the question

  • Assumes a misplaced compressor causes a visible failure
  • Forgets the error stage also produces a body
  • Leaves a buffering compressor over a long-lived streamed response
  • Compresses assets that were already encoded at build time
  • Measures wire bytes with a counter placed inside the compressor