Why does one slow operation in a batched GraphQL request delay every other result in it?
answer
- One response cannot leave in pieces
- The total is a maximum, not an average
- Concurrency helps the sum, not the coupling
- It happens on an idle multiplexed connection too
- Nine members, nine chances to hit a tail
basics
~20 sBecause all the results travel in one JSON array in one HTTP response, which cannot be delivered until the last member has finished. Running members concurrently shortens the total wait, but the client still sees nothing until the slowest one completes.
solid answer
~50 sA batch is one request and one response. The response body is a single JSON array, so the server cannot serialize and flush it until every member has produced its envelope; the batch's latency is therefore the **slowest** member's latency plus overhead, however fast the others were. Concurrency inside the server helps the total but not the coupling. It is a property of the response body rather than of the connection, so it persists on an idle, fully multiplexed transport where the unbatched operations would have returned independently. The consequences are worse tail latency - a batch is slow whenever *any* member is slow, so the tails compound - and a fast operation on the rendering path arriving at the speed of a slow one it has nothing to do with. The fixes: stop co-batching different latency classes, cap batch size and the collection window, and give members a deadline that yields a member-level error rather than stalling the array.
code
pseudocode · 8 lineshandleBatch(members):
results = awaitAll(members.map(m -> execute(m))) # concurrent, but joined
# nothing has been written to the socket yet
writeStatus(200)
writeHeader("content-type", "application/json")
writeBody(encodeJson(results)) # one array, one flush
# latency seen by every caller = max(execute(m)) + encode + writego deeper
Hold on to the shape of the answer: one HTTP response carrying one array cannot be sent until every operation in it has finished, so the slowest one sets the wait for all of them.
Explain why concurrency does not save you - it converts the sum into a maximum but the maximum is still shared - and be able to say that this is caused by the shared response body, not by the connection.
Demonstrate the diagnosis and the arithmetic: the gap between per-member server timings and client-observed latency, and the fact that nine members give nine independent chances of a tail, so the batch's p99 rate is far worse than any member's.
Own the tradeoff at the level of what the client is allowed to fuse. Decide which operations may ever share a request, whether a per-member deadline is a platform guarantee, and whether batching earns its keep on a transport that already multiplexes.
## Where the coupling comes from The response body of a batched request is one JSON array. A JSON array is delivered as a unit: the server can only write element 2 after element 1, and in practice buffers the whole structure before writing anything, because the status and headers are committed at the top of the response and the array must be well formed. So the earliest moment any result can reach the client is the moment the **last** member finishes. That produces a rule worth stating plainly: `batchLatency = max(memberLatency) + overhead`. Executing members concurrently - which servers commonly do - changes the max from a sum into a maximum, which is a large win, but it does not decouple anything. Nine fast members and one slow one still leave the client waiting on the slow one. Crucially this has nothing to do with the connection. It is not requests queueing behind each other on a transport that cannot carry them in parallel; that is a separate, connection-level phenomenon with its own remedies. Batch coupling happens on a completely idle, fully multiplexed connection, because the constraint is that the ten results are one HTTP response by construction. Batching, in effect, reintroduces at the application layer a delay that the transport layer had already solved. ## What it looks like when it bites Take a charity donations graph whose supporter dashboard mounts nine operations within a few milliseconds and whose client library batches them. Eight are small - a campaign banner at 23 ms, a live total at 31 ms, a navigation menu at 19 ms. The ninth asks for a donor profile, a 37-field type that fans out into a giving history and a Gift Aid claim status held in a slower downstream system, and it takes 840 ms. Unbatched, the banner paints at 23 ms and the profile fills in later. Batched, nothing renders until 861 ms. Nobody wrote that regression; it arrived when someone enabled batching in the client, and every operation on the page inherited the worst latency on the page. The tail behaviour is worse than the average behaviour, and this is the part senior candidates are expected to reach for. Assume each of the nine members independently exceeds its own p99 one time in a hundred. The batch exceeds it whenever *any* member does, so the batch's rate is `1 - 0.99^9`, about 8.65% - roughly nine times worse than any single operation. Batching does not average latencies out, it takes their maximum, and maxima are dominated by tails. ## Diagnosing it The symptom that identifies it is a gap between server-side per-operation timings and client-observed latency for the fast operations. The server, if it is instrumented per member, will report the banner at 23 ms; the browser will report 861 ms for the request that carried it. When those two numbers disagree by roughly the duration of the slowest member in the same request, the batch is the cause. Two checks confirm it. Disable batching in one client build and see whether the fast operations' client-observed latency collapses to their server timings. Then look at the composition of slow batches: if the slow ones are exactly the batches that happen to contain one particular member, the coupling is doing the damage rather than the server being overloaded. ## What to do about it **Do not co-batch different latency classes.** This is the highest-value fix and it is a client-side one. Operations on the rendering path belong in their own request; anything known to fan out into a slow downstream belongs on its own. Many batching clients allow a per-operation opt-out for exactly this reason. **Cap the collection window and the batch size.** A window of a few milliseconds collects a burst without holding anything back; a window long enough to be generous is a fixed latency tax on every operation in it, paid before execution even begins. A member cap bounds how many operations one slow member can hold hostage. **Give members a deadline.** If the server enforces a per-member time limit and returns that member's envelope with a field error when it expires, the array completes on time and eight callers get their data. Without it, one pathological member sets the latency for everyone and, on a client timeout, fails everyone. **Question the batch at all.** On a multiplexed transport, batching buys header and framing overhead. It costs coupled latency, one status for many outcomes, and per-operation caching and attribution. The saving is often smaller than the costs, and turning batching off is a legitimate and frequently correct answer. Incremental delivery is sometimes suggested here - a spec-track direction that lets one operation's response arrive in several payloads - but it addresses slow fields inside a single operation, not the coupling between separate operations sharing a body.
- If the server executes the members concurrently, why is there still a problem?Concurrency turns the batch's cost from the sum of member latencies into the maximum of them, which is a real win, but the maximum is still shared. The response is one array in one HTTP response, so it cannot be flushed until the last member completes. Concurrency improves the number; it does not decouple the callers.
- How would you tell this apart from the server simply being slow?Compare per-member server timings against client-observed latency for the same request. If the server reports a fast operation at tens of milliseconds while the client sees hundreds, and the gap matches the slowest member in that same batch, the coupling is the cause. Confirming it is a matter of disabling batching in one build and watching the fast operations' client latency drop to their server timings.
- What is the single most effective mitigation?Stop co-batching different latency classes on the client. Operations on the rendering path go in their own request, and anything that fans out into a slow downstream goes on its own. Size caps, a short collection window and per-member deadlines all help, but none of them fixes a batch that deliberately mixes a 23 ms banner with an 840 ms profile.
- Does a per-member deadline change what the client receives?Yes, and deliberately. The timed-out member's envelope carries an error instead of data, so that caller degrades while the other eight get their results on time. Without it the whole array waits, and a client-side timeout then fails every caller in the batch rather than the one that was actually slow.
It is a minibus rather than nine taxis: everyone leaves together, so the last passenger to arrive sets the departure time for all of them.
saying these in an interview costs you the question
- Says server-side concurrency removes the coupling
- Blames the connection rather than the shared response body
- Assumes a multiplexed transport makes the problem disappear
- Thinks batch latency is the average of its members
- Ignores that tails compound across members
- Proposes a longer collection window to fix slow rendering