skip to content

Users report that your web app works in a fresh private browsing window but fails with an error page after they have been logged in for a while, and your application logs show nothing. How would you investigate, and what makes accumulated cookies or a large bearer token the likely cause?

level: seniorimportance: must knowfreq 45%

answer

  1. Private window works, logged-in fails
  2. No application log entry → rejected upstream
  3. Admins fail first → claims-heavy token
  4. Bisect the Cookie header with curl
  5. Fix order: shrink, scope, expire, then raise limits

basics

~20 s

That pattern means the request is rejected during header parsing, before the app runs. Cookies and token claims accumulate until the header section passes the server's ~8 KB cap, giving 400/431 or a reset. Measure header size at the edge, reproduce with curl, then shrink cookies and tokens.

solid answer

~60 s

The signature — works anonymously, fails once a session has accumulated state, no application log line — points to rejection at a front-end parser, not in the app. **Why it grows.** Cookies are sent on every matching request. A session cookie, a CSRF cookie, analytics cookies, feature-flag cookies, plus anything set on the parent domain and inherited by every subdomain, add up. A JWT in a cookie or `Authorization` header grows with claims — roles, group memberships, permissions — so the users who fail first are the ones in the most groups, which looks maddeningly like a permissions bug. **How to confirm.** Look at edge/proxy logs for header-size errors (nginx logs "too long header line"; check `$request_length`). Reproduce with `curl -H 'Cookie: …'` at increasing sizes to find the cliff. Compare tiers — an outer tier may allow more than the inner one, turning a 431 into a 502. **How to fix.** Shrink the token (drop claims, or swap a self-contained token for an opaque session id resolved server-side), scope cookies with `Domain`/`Path` so they are not sent everywhere, expire dead cookies, and only then consider raising limits consistently on every tier.

code

bash · 10 lines
bash
# Replay with the real cookie, then with half of it
curl -s -o /dev/null -w '%{http_code}\n' https://app.example.com/dashboard -H "Cookie: $(cat cookies.txt)"
curl -s -o /dev/null -w '%{http_code}\n' https://app.example.com/dashboard -H "Cookie: $(cut -c1-4000 cookies.txt)"

# Synthetic sweep to locate the threshold
for n in 4000 8000 12000 16000; do
  printf '%s -> ' "$n"
  curl -s -o /dev/null -w '%{http_code}\n' https://app.example.com/ \
    -H "X-Pad: $(head -c $n < /dev/zero | tr '\0' 'a')"
done

go deeper

for a junior

Recognise the pattern — clearing cookies fixes it — and know that oversized headers are rejected before the app runs.

for a middle

Locate the rejecting tier from its logs, reproduce with curl, and identify cookie accumulation or a large JWT as the driver.

for a senior

Run the full investigation across tiers, explain the 502-instead-of-431 variant, and prioritise fixes: shrink the token, scope cookies, expire dead ones, and only then raise limits everywhere.

for a principal

Treat header size as a platform budget with monitoring and an owner, and make the architectural call between self-contained tokens and opaque session references based on where the size cost lands.

## Reading the symptom Three facts, taken together, are almost diagnostic: 1. **A private window works.** Private browsing starts with no cookies, so the request is small. 2. **Long-lived sessions fail.** The failure correlates with accumulated state, not with a code path. 3. **Application logs are empty.** The request was rejected by a parser upstream of the application — nginx, a load balancer, a CDN, or the servlet container's HTTP connector — before any handler ran. Users describe it as "the site broke", "clearing cookies fixes it", or "it only breaks for the admins". The last variant is the tell for token bloat: privileged users carry more claims. ## Why headers grow over time **Cookies.** Every cookie whose `Domain` and `Path` match is attached to every request, including requests for images and API calls. Sources accumulate: - session and refresh cookies, - CSRF tokens, - analytics and A/B testing cookies (often set by third-party scripts on the apex domain, and therefore inherited by every subdomain including the API host), - consent-management state, - feature flags and UI preferences, - stale cookies from removed features that nobody expires. Browsers cap an individual cookie at about 4 KB and allow on the order of 180 cookies per domain, so a browser will happily hold far more total cookie bytes than a server will accept in one request. The browser drops silently; the server rejects loudly. **Tokens.** A JWT is base64url-encoded JSON, so it is bulky by construction: a signature plus a payload. Add `roles`, `groups`, `permissions`, `entitlements` arrays, and enterprise directory group lists, and 6–10 KB tokens are entirely realistic. Store one in a cookie and it is sent on every request; put it in `Authorization` and it is a single huge header line, which is worse under nginx's per-line rule. **Proxy-added headers.** Each hop may add `X-Forwarded-For`, `X-Forwarded-Host`, `X-Forwarded-Proto`, `Forwarded`, `traceparent`, `X-Request-ID`, and vendor-specific fields. A request that fits at the edge can exceed the origin's limit after enrichment — which is why the innermost tier must not be the most restrictive. ## Investigation, step by step 1. **Locate the rejecting tier.** Walk inward from the browser. Whoever logs the error owns the limit. nginx: `client sent too long header line` / `request header is too big` in the error log; Tomcat: a parse-error entry; a CDN or cloud LB: an access-log entry with no origin request. 2. **Measure size.** In nginx, log `$request_length` (request line plus headers plus body) and inspect the p99, not the mean. Browser devtools shows the request headers; copy them as curl to get the exact bytes. 3. **Reproduce deterministically.** Replay the captured request with curl, then bisect: halve the Cookie header until it passes. That both confirms the cause and tells you the effective threshold. 4. **Confirm the cliff matches a known default.** ~8 KB strongly suggests nginx `large_client_header_buffers` or Tomcat `maxHttpHeaderSize`; ~16 KB suggests Node or a raised setting. 5. **Check every tier's configured limit**, because a mismatch produces the confusing variant: the edge accepts, the origin rejects, and the edge reports 502 rather than 431. ## Fixes, in the order you should prefer them **1. Carry less.** Replace a claims-heavy self-contained token with an opaque session identifier and look up authorization server-side (cache it). This trades a lookup for a permanently bounded header and is usually the right architectural answer once tokens exceed a couple of kilobytes. If a self-contained token must stay, remove derivable claims — full group lists, display names, redundant profile data — and keep only what the resource server needs to authorize. **2. Scope cookies.** Set `Domain` to the narrowest host that needs the cookie and `Path` to the narrowest prefix. Do not set application cookies on the apex domain if API subdomains never read them. Serve static assets from a cookieless domain so 8 KB of cookies do not ride along with every icon. **3. Expire the dead.** Ship a one-time response that clears retired cookies (`Set-Cookie: old=; Max-Age=0; Path=/`). Crucially, if your error page itself is served by the app, make the *error* response clear the offending cookies — otherwise the user is stuck in a loop where every request, including the one that would fix them, is oversized. **4. Raise limits — last, and everywhere.** Pick a bounded value (16 KB or 32 KB), apply it from the outermost tier inward, and never leave an inner tier smaller than an outer one. Remember the buffer is per connection: the new limit times peak concurrency is memory an unauthenticated client can force you to allocate. **5. Add a guardrail.** Alert when p99 header size crosses, say, 60% of the limit. Header bloat grows monotonically with product features; without a monitor it re-emerges every year. ## The HTTP/2 wrinkle On an HTTP/2 or HTTP/3 hop, HPACK/QPACK compress repeated headers to almost nothing on the wire — so a browser trace can look healthy while the origin still rejects. The advertised limit (`SETTINGS_MAX_HEADER_LIST_SIZE`) is measured on the *uncompressed* list, and the moment a proxy downgrades to HTTP/1.1 the full bytes reappear. Never conclude from "HTTP/2 compresses it" that the budget is fine.

  • Only users belonging to many directory groups are affected. What does that tell you, and what is the durable fix?
    It says the size driver is per-user token content, almost certainly a group or roles claim embedded in a JWT, so header size scales with group membership. The durable fix is to stop shipping the group list in the token: issue an opaque session or token reference and resolve authorization server-side with a cache, or reduce the claim to a small set of coarse roles the resource server actually checks. Raising limits only postpones the next boundary.
  • Why can this failure appear as a 502 Bad Gateway rather than a 431 or 400?
    Because the tier that rejects is not the tier the client talks to. If an edge proxy accepts a 20 KB header but the origin's limit is 8 KB, the origin rejects or resets the upstream connection and the edge reports that as a 502. The give-away is that the 502 is instant rather than after a timeout and correlates with request size, so the fix is to align limits from the outside in.
  • How do you keep this from recurring after you fix it?
    Instrument request header size at the edge — log it, chart the p99, and alert at a fraction of the configured limit — and treat the header budget as an explicit platform constraint that new cookies and claims must fit within. Also add a test or lint on the auth service that fails when the issued token exceeds a size threshold, so growth is caught at the source rather than in production.

A doorway with a fixed frame: each new badge, lanyard and folder the employee accumulates is fine on its own, but one morning the whole armful no longer fits through, and the person on the other side of the door never even sees them arrive.

saying these in an interview costs you the question

  • Blaming the application or a permissions bug because only privileged users fail, without noticing that their tokens are larger.
  • Searching only application logs and concluding nothing is wrong.
  • Fixing it solely by raising the limit, so the same incident returns as the token keeps growing.
  • Assuming HTTP/2 header compression makes the size irrelevant — limits are measured uncompressed.
  • Serving an error page that sets more cookies, trapping the user in an unrecoverable loop.

context