In GraphQL, why do per-view queries produce a request waterfall, and what does fragment colocation change?
answer
- Count round trips, not requests
- The child needs the parent's id first
- Needs known before the first byte
- One document, one latency budget
- Split when a panel is slow
basics
~20 sA nested view cannot form its query until its parent's response supplies the id it needs, so round trips run in series. Colocation hoists every view's fields into one operation the screen sends before rendering.
solid answer
~50 sA waterfall is a **data dependency**, not a concurrency limit: the child's variables come out of the parent's response, so request two cannot start until request one lands, and a six-level screen costs six serial round trips plus six rounds of parse, validate and authenticate. Colocation removes the dependency by making the whole screen's needs statically known before the first byte - the fragments are composed at author time, so the root sends one document. The judgement an interviewer wants next is the cost: one operation shares a single latency budget, so the slowest field gates the paint. Mitigations are incremental delivery with `@defer` on a slow spread, or deliberately keeping a second operation for a slow or optional panel so the shell is not held hostage. Colocation does not touch server-side fan-out; that stays a batching problem.
code
pseudocode · 6 linesrender OrderScreen(orderId):
order = fetch("{ order(id: $orderId) { id status courierId } }") # round trip 1
render CourierCard(order.courierId):
courier = fetch("{ courier(id: $cid) { displayName etaMinutes } }") # round trip 2
render CourierRating(courier.id):
fetch("{ courierRating(courierId: $id) { stars reviews } }") # round trip 3go deeper
Recall the basic shape: fetching per view means each nested view waits for its parent's response, while colocation lets the screen ask for everything in one go.
Explain why the requests cannot overlap - the child's variables come from the parent's data - and what one composed document saves beyond network time: one auth pass, one validation, one batch window.
Bring the tradeoff unprompted: a single response is gated by its slowest field, and name the two exits - incremental delivery where it is supported, or a deliberate second operation for a slow panel.
Frame it as a budget: what a screen may cost in round trips, where the cost cap sits, how nullability limits a panel's blast radius, and which parts of a page are allowed their own latency ceiling.
## Why per-view fetching serialises If each view fetches for itself, a nested screen does not issue N requests in parallel. It issues them in a **chain**, because each request's variables are values from the previous response: ```pseudocode render OrderScreen(orderId): order = fetch("{ order(id: $orderId) { id status courierId } }") # round trip 1 render CourierCard(order.courierId): courier = fetch("{ courier(id: $cid) { displayName etaMinutes } }") # round trip 2 render CourierRating(courier.id): rating = fetch("{ courierRating(courierId: $id) { stars } }") # round trip 3 ``` `CourierCard` cannot even be mounted, let alone form its query, until `order` has returned - so the requests are separated by full round trips, and the screen's time to last paint is roughly `depth x RTT` plus server time, not `max(request)`. On one order-tracking screen with 14 data-bound views nested 6 levels deep, that shape produced 9 distinct operations and 1,743 ms to last paint at the 75th percentile on mobile. Composed into a single colocated document of 118 fields, the same screen paints in 386 ms. The composed request is far larger than any of the nine it replaced; it is still four times faster, because the cost that dominated was serialised network latency. Three secondary costs travel with the waterfall. Each hop repeats authentication, parsing and validation. The same entity is often fetched by several views independently, and any per-request batching the server does is scoped to a request - so nine requests means nine tiny batch windows instead of one wide one. And the UI shows a cascade of spinners appearing in sequence, which reads as slowness even where the total is acceptable. ## What colocation changes structurally The important word is **static**. Under colocation, the fields the whole tree will render are known from the composed document before rendering starts, so the screen root can ask once. Composition happens when the code is written, not when components mount. Nothing about the render tree is negotiated with the network any more. That also widens every request-scoped optimisation the server has: one authentication pass, one validation pass, one request-scoped batch cache covering the whole screen rather than a slice of it. ## The tradeoffs a senior is expected to raise **One shared latency budget.** A single response is delivered when its slowest field is done. If the delivery-history panel joins a cold analytics store and takes 1.4 s, the status chip waits 1.4 s too - which the waterfall, for all its faults, did not do. Two answers exist. The first is **incremental delivery**: `@defer` on a slow fragment spread (and `@stream` for a long list) lets the server send the rest immediately and the deferred payload later. These are incremental-delivery additions on the specification's working track rather than settled parts of the released editions, and they need matching server and client support, so treat them as a capability to verify, not to assume. The second is simply **not composing everything**: keep a second operation for the slow or optional panel. Colocation is compatible with that - it says fragments compose upward, not that a screen may only have one root. **Blast radius.** A field error deep inside one view's fragment can null more of the response than that view, depending on nullability, so one panel's failure becomes a wider hole than it would have been as its own request. Budgeting nullability for that is its own discipline. **Static limits.** A composed screen document is bigger and deeper, and a server enforcing depth or cost caps may reject at 118 fields what it accepted at 13. That is a conversation with whoever sets the cap, not a reason to fragment the screen back into a waterfall. **Invisible UI.** Fragments for tabs, modals and drawers that are not on screen still ship their fields unless the spread is conditioned on a variable. **It does not fix server-side fan-out.** Asking for a list's children in one document still causes per-item backend work unless the server batches; colocation changes the number of HTTP round trips, not the number of database calls. ## How to prove it in an interview Say what you would measure. Count **operations per screen render**, not bytes; look for the stair-step in a network timeline where each request starts after the previous one finishes; compare time to last paint before and after. And be clear that "one request per screen" is a default, not a law - the goal is that no request waits on another for data it did not need.
- Doesn't HTTP/2 multiplexing remove the waterfall?No. Multiplexing lets many requests share a connection without head-of-line blocking, which helps requests that could already run at the same time. A waterfall is sequential for a different reason: request two's variables come from request one's response, so it cannot be issued early on any transport. Multiplexing shortens a wide fan-out, not a deep chain.
- When would you deliberately keep two operations on one screen?When part of the screen is materially slower, optional, or on a different cadence - a cold analytics panel, a permission-gated section, or a block that polls. The shell should paint on its own budget. If the server and client support incremental delivery, deferring that spread inside one operation is the tidier version of the same decision.
- Does composing to one operation reduce the number of backend calls?Not directly. The same fields resolve, so the same downstream work happens. What improves is that authentication and validation run once and any request-scoped batching now spans the whole screen. Eliminating per-item fan-out is a server-side batching concern and is unaffected by how the client grouped its documents.
saying these in an interview costs you the question
- Says HTTP/2 multiplexing removes the waterfall
- Insists one request per screen is always right
- Claims colocation fixes server-side fan-out
- Ignores that the slowest field gates the response
- Assumes a composed document passes any cost cap
- Optimises payload bytes instead of round trips