Some users of a web app suddenly get HTTP 431 Request Header Fields Too Large (or a 400 from the CDN) on every request, and clearing cookies fixes it. How do you diagnose and permanently fix this?
answer
- 431 = header block too big, emitted by proxy not app
- clearing cookies fixes it → Cookie header is the cause
- OIDC state/nonce leftovers + fat JWT = classic pair
- measure p99 cookie bytes at the edge
- cookieless asset host + opaque session id
basics
~20 sAccumulated cookies pushed the Cookie request header past the proxy or server header limit. Measure total cookie bytes per user, find the writer that keeps adding cookies, cap and expire them, shrink the session cookie to an opaque id, and serve assets from a cookieless host.
solid answer
~60 s**Diagnose.** 431 (or a 400/494 from nginx, 413/494-class errors at CDNs) means the request header block exceeded the server's limit — commonly 8 KB total or 4–8 KB per header. Since clearing cookies fixes it, the `Cookie` header is the culprit. Reproduce by dumping `document.cookie` on an affected profile, then log total cookie bytes and cookie count at the edge for a sample of traffic so you can see the distribution, not one anecdote. **Find the writer.** Typical causes: per-tenant or per-feature cookies with unbounded names (`state_<uuid>`), OAuth/OIDC state and nonce cookies never cleaned up after the callback, a fat JWT with claims that grew, analytics and A/B cookies from third-party tags, and a `Domain=example.com` cookie set by many subdomains. **Fix.** Replace self-contained tokens with an opaque session id backed server-side; delete transient auth cookies at the end of the flow and give them short Max-Age; namespace and cap per-feature cookies; move static assets to a cookieless domain; raise the proxy limit only as a stopgap. Add an alert on p99 cookie-header bytes.
code
bash · 3 linescurl -s -o /dev/null -w '%{size_request}\n' \
-H 'Cookie: session=...; state_1=...; state_2=...' \
https://app.example.com/api/mego deeper
Recognise 431 as 'request headers too big' and that cookies are the usual cause; know that clearing site data is the user-side workaround.
Walk through measuring the Cookie header, identifying which cookie family dominates, and shrinking or expiring it.
Own the full loop: edge instrumentation, attribution by cookie family, opaque session id, transient-cookie lifecycle, cookieless asset host, and a recovery path for already-stuck browsers.
Set an org-wide cookie budget with ownership per name prefix, enforce it in CI or at the edge, and decide the token strategy (opaque vs self-contained) that keeps the budget viable as the estate grows.
## What 431 actually means `431 Request Header Fields Too Large` is the HTTP status for a request whose header section, or a single header field, exceeds what the server will accept. It is generated before the application sees the request, by the web server, reverse proxy or CDN. Common limits: nginx `large_client_header_buffers` defaults to 4 buffers of 8 KB, Node's HTTP parser defaults to about 16 KB total headers, many CDNs cap at 8–16 KB with a per-header cap around 4–8 KB. Some stacks answer with `400 Bad Request` or a vendor-specific code instead of 431, which is why the symptom is often reported as "random 400s". Cookies dominate because the browser concatenates *every* cookie matching the host and path into a single `Cookie` header and sends it on **every** request — XHR, navigation, images, fonts. Nothing else in a normal request grows without bound. The signature is exactly the one in the question: it affects a subset of users, it is sticky for a given browser profile, it survives reloads, and clearing site data fixes it instantly. That pattern is cookie accumulation, not a server bug. ## Diagnosis 1. **Confirm the header size.** On an affected profile, inspect the request in devtools and read the `Cookie` header length, or run `document.cookie.length` (note this omits `HttpOnly` cookies, so it is a lower bound — the real figure must come from the server or proxy). 2. **Measure the fleet, not the anecdote.** Log `Cookie` header bytes and cookie count at the edge for sampled requests. Plot the distribution. You are looking for a long tail that creeps upward with account age or session count — that shape means something writes a new cookie per event and never deletes. 3. **Attribute the bytes.** Group cookies by name prefix and see which family dominates. The usual suspects: - **OIDC/OAuth transient cookies**: `state`, `nonce`, `code_verifier`, often named with a random suffix so each login attempt leaves a new one. Abandoned logins accumulate forever if no `Max-Age` is set. - **Fat self-contained tokens**: a JWT that grew claims release by release, sometimes stored twice (access + refresh), sometimes chunked into `token.0`, `token.1`. - **Per-tenant / per-feature cookies**: one cookie per visited workspace or per experiment. - **Third-party tags** setting first-party cookies on the apex domain. - **Subdomain sprawl**: several apps all writing `Domain=example.com`, so every app pays for every other app's cookies. ## The fix, in order of value **Shrink the credential.** Move from a self-contained token in a cookie to an opaque session identifier resolved server-side. This usually reclaims kilobytes at a stroke and buys revocation as a bonus. **Give every transient cookie a lifetime and a delete path.** Auth-flow cookies should carry a short `Max-Age` (minutes) and be explicitly expired at the callback with `Set-Cookie: name=; Max-Age=0; Path=/`. Bound the number of concurrent flow cookies — keep the newest N and expire the rest. **Namespace and cap.** Give each subsystem a name prefix, and enforce in code review that no code path writes an unbounded number of cookie names. Where you need per-entity state, keep one cookie whose value is a bounded list rather than N cookies. **Scope tightly.** Prefer host-only cookies (no `Domain`) so `app.example.com` does not carry `marketing.example.com`'s cookies. Serve static assets from a separate host that your cookies never match — this removes cookie bytes from the majority of requests even before you shrink anything. **Raise limits only as a stopgap.** Bumping `large_client_header_buffers` or the CDN limit unblocks users today, but the header still grows, and you may have several hops each with its own limit — raising one just moves the failure. **Add a self-heal for stuck users.** Because a stuck browser cannot even reach the app, a 431 error page that instructs the browser to clear cookies (or a dedicated recovery endpoint on a sibling host that issues `Max-Age=0` for known cookie names) rescues the affected tail without a support ticket. **Alert.** Track p99 cookie-header bytes as a first-class metric with a threshold well under the smallest hop's limit. Because oversize failures are invisible until they are total, the metric is the only early warning you get.
- Why is raising nginx's large_client_header_buffers not an acceptable permanent fix?It treats a growth problem as a static one: whatever writes cookies without bound will exceed the new limit too. It also only fixes the hop you changed — a CDN, load balancer or the application server may each have their own cap — and larger buffers cost memory per connection. Use it to unblock users while you fix the writer.
- Why can document.cookie under-report the problem?document.cookie omits HttpOnly cookies, which usually include the session and auth cookies — exactly the large ones. It also shows only cookies matching the current path and host. The authoritative measurement is the Cookie header length observed server-side or at the proxy.
saying these in an interview costs you the question
- Blaming the browser or 'corrupt cookies' rather than measuring header size
- Fixing it only by raising the proxy limit
- Using document.cookie length as the authoritative size, ignoring HttpOnly cookies
- Deleting cookies without a Path/Domain match and assuming they are gone
- Not adding monitoring, so the same tail failure recurs next quarter