skip to content

Some GraphQL queries sent over GET fail with an HTTP error and no response body — how do you diagnose it?

level: seniorimportance: should knowfreq 44%

answer

  1. Present at the edge, absent inside
  2. Failures track size, not operation
  3. Every hop has its own ceiling
  4. Escaping triples what you must fit
  5. Find the byte threshold, then the hop

basics

~20 s

Suspect URL length. A percent-encoded document plus its variables can exceed the request-line or header limit of a browser, proxy or server, which rejects the request — often 414 URI Too Long — before any GraphQL execution.

solid answer

~50 s

The diagnostic tell is asymmetry: the requests appear in edge HTTP logs with a 4xx status and are completely absent from GraphQL operation metrics, meaning they never reached execution. On a GET that almost always means URL length. Nothing specifies a maximum — the browser, every intermediary and the server each impose one, commonly a few kilobytes across the request line and headers — and percent-encoding inflates escaped characters threefold, so the encoded length is far larger than the document you wrote. Confirm by correlating failures with request *size* rather than operation name, then replay one operation while growing a list variable until you find the binary cutover, then bypass hops until you find which one rejects it. Fix by shortening the URL with a persisted identifier, bounding list variables, or sending that operation over POST and accepting the loss of shared-cache eligibility.

code

graphql · 6 lines
graphql
query ScopedSegmentMatches($projectIds: [ID!]!, $sourceHash: String!, $targetLocale: String!) {
  segmentMatches(projectIds: $projectIds, sourceHash: $sourceHash, targetLocale: $targetLocale) {
    score
    target
  }
}

go deeper

for a junior

Know that a GET puts the whole operation and its variables into the URL, that URLs have practical length limits, and that a request rejected for length never reaches the GraphQL server at all.

for a middle

Explain where the limits live — browser, intermediaries, origin request line and header buffers — that none of them is specified by GraphQL, and that percent-encoding roughly triples every escaped character in what must fit.

for a senior

Demonstrate the diagnosis: spot the edge-versus-application metric asymmetry, correlate failures with byte length rather than operation, bisect to the exact threshold, and walk the hops to find the owner before changing anything.

for a principal

Own the prevention: a client-side length budget with a defined fallback, a build-time check on encoded operation size, alerting on statuses that never appear in application error metrics, and a policy on which operations may use GET at all.

## The signature: a failure that tracks size, not operation The tell in this failure is a mismatch between two vantage points. At the edge — a load balancer's access log, a CDN report, the browser's network panel — the requests exist and carry a 4xx status. In the GraphQL server's own operation metrics and logs, they do not exist at all. No parse, no validation error, no resolver, no timing. Traffic that is visible to HTTP and invisible to GraphQL means the request was rejected before execution, and on a GET the overwhelmingly common reason is that the URL was too long for some hop to accept. The second tell is *which* requests fail. A bug in an operation fails that operation every time. A length ceiling fails whichever requests happen to cross it, so failures cluster by variable payload rather than by operation name — the same query succeeding for one caller and failing for another is the shape to look for. ## Where the ceiling actually lives Nothing specifies a maximum URL length. What exists is a stack of independent limits, any one of which can be the binding one: - **The browser or client runtime**, which will refuse or truncate beyond its own bound. You do not control it and cannot configure it for your users. - **Every intermediary**: CDN, reverse proxy, load balancer, API gateway. These typically cap the request line and the total header block, with configurable defaults commonly in the 4-8 KB range. - **The server's own HTTP layer**, with its own request-line and header-buffer limits, applied before your application code runs. The status you get back depends on which hop refuses. **414 URI Too Long** is the purpose-built answer; many intermediaries answer 400 instead, and a limit hit on the header block rather than the request line can surface as 431. Because the refusal is generated by an intermediary, the body is that hop's error page — never a GraphQL envelope. A client that assumes otherwise reports a JSON parse failure, which is why this often reaches you as "the API returns garbage" rather than "the URL is too long". Remember also that percent-encoding inflates the payload roughly threefold for every escaped character, so the encoded length you must fit is far larger than the document you wrote. ## Diagnosing it, in order 1. **Correlate with size, not with operation.** Pull the failing request lines and measure their length. If the failures start above a threshold and there are no failures below it, you are done hypothesising. 2. **Find the exact threshold.** Replay one operation from a shell client, growing a list variable an element at a time. A clean binary cutover — every request under *n* bytes succeeds, every request over it fails — is conclusive, and the number usually lands on a recognisable configured value. 3. **Walk the hops.** Send the oversized request straight at the origin, bypassing the edge, then at each layer in turn. The first hop that rejects it owns the limit. Its access log will show the request; the next hop's will not. ## The red herring worth naming On a translation-memory graph running to a 1,200-request-per-minute peak, roughly 3.1% of GET reads began failing with no server-side trace. The failures clustered almost entirely on senior editors, which pointed the first day of investigation at authorization — specifically the theory that a permission check ran too late in the pipeline and was blowing up after the response had been assembled. It was not. Senior editors were simply the accounts with many project scopes, and the client passed those scopes as a list variable: 47 identifiers of about 40 characters each, which pushed the encoded URL past an 8,192-byte limit on one proxy tier. The correlation with privilege was real and entirely incidental. The lesson generalises — when a failure correlates with a user attribute, check whether that attribute also correlates with request *size* before believing it is about permissions. ## Fixes, in the order worth trying **Shorten the URL.** Replace the document text with a persisted operation identifier; that removes the largest and most variable component and leaves the variables as the only thing that can grow. **Bound the variables.** A list variable with no cap is an unbounded URL. Page it, or take a filter identifier instead of an inline list. **Fall back to POST for that operation.** Legitimate, and safe because the operation is a read, but you pay a round trip if you discover the limit by getting rejected, and the response loses its shared-cache eligibility — which was the entire reason for using GET. **Raise the limit.** Possible at hops you own, and useless unless you raise it at *every* hop, including ones you may not own. The browser's limit is not yours at all, so this can never be the whole answer. ## Keeping it from returning Put the guard on the client, where the length is known before the request leaves: compute the encoded URL, and if it exceeds a configured budget comfortably under the tightest hop, send that one operation over POST instead of discovering the ceiling from a 414. Back it with a check in the build that fails when a persisted operation's encoded form exceeds the budget, and an alert on the 414 and 431 rate at the edge, since by construction those never appear in application-level error metrics.

  • Why is raising the proxy's limit an incomplete fix?
    Because the ceiling is a stack of independent limits and the binding one moves as soon as you raise the first. Every intermediary and the origin server each cap the request line and header block, so a fix has to be applied at every hop — and the browser's own bound is not configurable by you at all, so client-originated requests can still fail after you have raised everything you own.
  • A client retries as POST after receiving 414. What are the tradeoffs?
    It is a legitimate fallback and safe here because the operation is a read, so repeating it changes nothing. The costs are a wasted round trip to discover the limit, and the loss of the shared-cache eligibility that motivated GET in the first place. Better to compute the encoded length client-side against a budget under the tightest hop and choose the method before sending.
  • Failures cluster on one group of users. Why is that not evidence of an authorization problem?
    Because a user attribute often correlates with request size. A group with more scopes, more projects or longer filter lists sends longer variable payloads, so a length ceiling looks exactly like a permission-dependent failure. Check the byte length of the failing requests against the succeeding ones before pursuing an authorization theory; if the split is a clean size threshold, permissions are incidental.

saying these in an interview costs you the question

  • Assumes a GET URL has no practical length limit
  • Looks for the failures in GraphQL operation metrics
  • Expects a data/errors body from an intermediary's rejection
  • Blames the operation instead of the payload size
  • Raises the limit on one hop and calls it fixed
  • Reads user-group clustering as an authorization fault

context