skip to content

You are setting a request header size budget for a platform with a CDN, a load balancer, an ingress tier and dozens of services. How would you choose the number, decide what is allowed to consume it, and stop it from being exceeded again a year later?

level: principalimportance: nice to knowfreq 22%

answer

  1. One number, whole chain, written down
  2. p99 measured + proxy-injected bytes + headroom
  3. limit × concurrency = pre-auth memory
  4. Outer tiers never stricter than inner
  5. Opaque session ref beats claims-heavy token

basics

~20 s

Pick one number the whole chain honours (commonly 16–32 KB), sized from measured p99 plus proxy-added headers and checked against limit times peak concurrency for memory. Make outer tiers no more permissive than inner ones, allocate the budget explicitly, and monitor p99 against it.

solid answer

~60 s

Treat the header budget as a platform contract with a number, an owner, and an alarm. **Choose the number.** Measure current p99 header bytes at the edge, add what your own hops inject (forwarded, tracing, request-id fields — often 500 B to 2 KB), add headroom, then sanity-check the memory: the buffer is per in-flight connection, so limit × peak concurrency is memory anonymous clients can reserve. 16 KB or 32 KB is usually defensible; "as large as possible" is not. **Make it uniform.** The smallest tier decides, so align CDN, LB, ingress and app servers, and never leave an inner tier stricter than an outer one — that converts a clear 431 into an opaque 502. Where a managed edge caps you below your target, that cap *is* the budget. **Allocate it.** Say explicitly what may live in headers: session identifier, CSRF token, a bounded set of tracing fields. Push per-user authorization data behind an opaque reference rather than a claims-heavy token, and keep bulk data in the body. **Defend it.** Chart p99 header size against the limit, alert at ~60%, and fail the auth service's build when issued tokens exceed a threshold.

code

bash · 8 lines
bash
BUDGET=16384
under=$((BUDGET - 1024)); over=$((BUDGET + 1024))
for n in $under $over; do
  printf 'bytes=%s status=' "$n"
  curl -s -o /dev/null -w '%{http_code}\n' https://edge.example.com/health \
    -H "X-Budget-Probe: $(head -c $n < /dev/zero | tr '\0' 'a')"
done
# expect: under -> 200, over -> 431 (not 502, which would mean an inner tier rejected it)

go deeper

for a junior

Know that the limit exists per tier and that headers must be kept small; the design call itself is above this level.

for a middle

Be able to align the configured limits across tiers and explain why the smallest one wins.

for a senior

Measure p99 header size, account for proxy-injected headers, set a bounded limit uniformly, and drive the fix toward smaller tokens and scoped cookies.

for a principal

Own the budget as a platform contract: a justified number from measurement and memory modelling, explicit allocation of who may consume it, external ceilings acknowledged, conformance tests, monitoring at a fraction of the limit, and a review gate on new cookies and propagated headers.

## Why this is a design question and not a config question Header size is a shared, silent resource. Every team can add a cookie, a claim or a tracing field, each individually trivial, and none of them sees the total. The failure arrives all at once, affects the most privileged users first, produces no application logs, and is usually resolved by raising a number — which resets the clock without changing the trajectory. Handled as a platform budget instead, it becomes a bounded, monitored constraint. ## Step 1 — measure what you actually carry Instrument the outermost tier you control. In nginx, `$request_length` covers request line, headers and body; log it and build a histogram. What matters is the **p99 and the maximum by user cohort**, because the distribution is bimodal: anonymous requests are tiny, and heavily-entitled internal users sit far out on the tail. Break the number down: session cookie, CSRF cookie, third-party analytics cookies on the apex domain, authorization token, tracing headers, user agent. Then measure **growth from your own infrastructure**. Each hop typically adds `X-Forwarded-For`, `X-Forwarded-Proto`, `X-Forwarded-Host`, `Forwarded`, `traceparent`/`tracestate`, a request id, and vendor headers. Half a kilobyte to two kilobytes is normal, and it lands on the *innermost* tier — the one with the least slack. ## Step 2 — pick a number you can justify Three constraints bound the choice: - **Demand:** p99 today, plus injected headers, plus headroom for a year of feature growth. - **Memory:** the header buffer is allocated per in-flight connection *before* authentication. Limit × peak concurrent connections is memory an unauthenticated client can force you to reserve. 32 KB at 20,000 connections is ~640 MB per node — fine on some fleets, fatal on others. This is the calculation that turns "why not 1 MB?" into an answer. - **Ceilings you do not control:** managed CDNs and cloud load balancers impose their own maxima on total headers and on individual fields. If the edge caps a single field at 16 KB, no origin setting makes a 20 KB cookie work. The lowest external ceiling becomes the budget. Pick one number, write it down, and state the rationale next to it. ## Step 3 — order the tiers correctly The rule is **monotonic non-increasing outward**: no tier should be stricter than the tier in front of it. If the edge is more permissive than the origin, oversized requests are accepted at the front and fail inside, and the client sees a 502 with no explanation instead of a 431 that names the problem. Configure from the outside in, and add a conformance test — a synthetic request sized just under and just over the budget, run against every environment, asserting the expected status and that the rejection happens at the intended tier. ## Step 4 — allocate the budget explicitly A budget without line items gets consumed by whoever moves first. Decide what headers are entitled to space: - **Session/authn:** prefer an opaque identifier (tens of bytes) over a self-contained claims-heavy token (kilobytes). The trade is a server-side lookup — cacheable, and it also buys immediate revocation. If self-contained tokens are required for cross-service verification without a shared store, cap the claim set: no group lists, no profile data, no anything derivable at the resource server. - **Cookies:** scope with `Domain` and `Path` so cookies are not attached to every request. Serve static assets from a cookieless host. Forbid application cookies on the apex domain when subdomains never read them, because third-party scripts are the usual source of surprise bytes. - **Tracing:** bounded and standardised (`traceparent`, `tracestate` with a length cap). Baggage propagation is the classic silent inflator — an unbounded key/value bag that grows across services. - **Everything else:** goes in the body. Body limits are separate and orders of magnitude larger. ## Step 5 — make it hard to breach - **Monitoring:** chart p99 header size as a fraction of the budget; page when it crosses ~60%, not when it hits 100%. - **Shift left:** a test in the identity service that fails the build when an issued token exceeds N bytes; a lint that flags new cookies without a `Domain`/`Path` scope. - **Self-healing errors:** the 431 error page should clear known-disposable cookies, so a user whose browser state is already over the limit can recover without support instructions. - **Review gate:** new cookies and new propagated headers are a design-review item with a stated byte cost, the same way a new database index or a new synchronous dependency is. ## Step 6 — know what HTTP/2 does and does not buy you HPACK and QPACK compress repeated fields to near nothing on the wire, so browser-side traces look reassuring. But budgets (`SETTINGS_MAX_HEADER_LIST_SIZE`) are computed on the **uncompressed** list, and the instant a proxy downgrades to HTTP/1.1 for an origin the full bytes are back. Compression is a bandwidth optimisation, not a budget increase — and internal service-to-service hops are frequently HTTP/1.1 with no compression at all. ## The judgment being tested The interviewer wants to see you refuse the two lazy answers: "just raise it" (ignores memory, ignores external ceilings, ignores recurrence) and "never raise it" (ignores real requirements like enterprise assertions). The good answer sets a bounded number from measurement, enforces it uniformly, allocates it deliberately, and defends it with monitoring — while acknowledging that the durable engineering fix is to stop putting per-user state in headers at all.

  • A team argues that setting the limit to 1 MB removes the problem permanently. What is your response?
    That the header buffer is allocated per in-flight connection before any authentication, so 1 MB times peak concurrency is memory an anonymous client can force you to reserve — at 20,000 connections that is 20 GB per node, a trivially cheap denial-of-service. It also does not work: managed CDNs and load balancers enforce their own ceilings you cannot raise, so the request still fails, just at a tier with worse error reporting. The bounded answer is a justified number plus removing per-user bulk from headers.
  • Where would you draw the line between a self-contained token and an opaque session reference for a platform of this size?
    Size and revocation drive it. Self-contained tokens are attractive when many services must verify without a shared lookup, but their cost is paid on every request by every hop, and they grow with entitlement complexity. Once a token approaches a couple of kilobytes, or the business needs immediate revocation, an opaque reference plus a cached server-side lookup is the better trade: constant header cost, instant invalidation, and the lookup can be co-located or cached at the edge.

Airline cabin baggage: one published size, enforced identically at every gate, with an explicit allowance per passenger — otherwise every traveller optimises for themselves and boarding stops working.

saying these in an interview costs you the question

  • Treating it as a single config change rather than a cross-tier contract with an owner.
  • Proposing a very large limit without modelling per-connection memory against peak concurrency.
  • Ignoring managed CDN and load-balancer ceilings you cannot configure.
  • Leaving an inner tier stricter than an outer one, which turns a clear 431 into an opaque 502.
  • Assuming HTTP/2 header compression enlarges the budget; limits are measured on the uncompressed field list.

context