When does splitting a GraphQL response with @defer actually improve what the user sees?
answer
- it moves perception, not work
- measure the field, not the request
- can the screen paint without it
- what else assumes one payload per request
- two parallel operations are the baseline
basics
~20 sOnly when the deferred part is measurably slower than the rest and the interface can render usefully without it. If the client blocks on the whole tree, or the slow field is what the screen exists to show, deferring adds payloads and merge logic and moves nothing.
solid answer
~50 sAsk three things. First, is the gap real and measured per field rather than per request — a section that costs 610 ms next to 27 ms for everything else is worth splitting; one that costs 8 ms more than its siblings is not. Second, can the interface paint a deliberate-looking placeholder and swap in the section without reflowing. Third, what else in the system assumes one payload per request: whole-body response caching stops applying, a request timer has to be redefined as time-to-first-payload versus time-to-complete, and anything judging success by status code will misread a 200 whose later payload carried a field error. Weigh that against two plain operations issued in parallel, which need no unratified directive and give each half its own status, cache entry and retry. `@defer` wins when the slow part depends on the parent you are already resolving and the screen has a genuine two-stage design.
code
graphql · 16 linesquery AccountOverview($id: ID!) {
account(id: $id) {
displayName
availableBalance
statements(last: 12) {
period
closingBalance
}
... on Account @defer(label: "spendBreakdown") {
spendBreakdown(months: 12) {
category
total
}
}
}
}go deeper
Understand the basic trade: the user sees something sooner, but the server does the same amount of work and the client gains a loading state it has to design and eventually give up on.
Be able to justify a specific deferral with per-field numbers rather than intuition, and name what the client now owes — a placeholder that does not reflow and a deadline for the payload that may never come.
Show the second-order effects you would check before shipping: whole-body response caching, split latency metrics, status-based success checks, and a truncated stream that presents as a permanent spinner with no error signal.
Own the call against the alternatives. Argue when two parallel operations or a precomputed field beat betting on directives that no released specification edition defines and whose envelope has changed across drafts.
## The claim being tested `@defer` does not make anything faster. Every deferred field still resolves, on the same server, against the same backends, and the server pays a little extra to frame and serialize more than one payload. What `@defer` moves is **when the client can render something useful**. So the whole judgement reduces to one question: is there a part of this screen that is meaningfully slower than the rest, and can the screen be useful without it? If the answer to either half is no, deferring buys nothing and costs real complexity. ## The three tests **Is the gap real, and measured per field?** Request-level latency will not tell you. You need field-level timing. On a bank statements graph, an account overview resolved the header and twelve statement summaries in about 27 ms and the twelve-month spending breakdown in about 610 ms, against a 340 ms p99 budget for first paint. That is a real gap: the breakdown is not a contributor to the latency, it *is* the latency, and everything else fits the budget several times over. Contrast a field that is 8 ms slower than its siblings — deferring it adds a payload, a merge, a loading state and a failure mode to save nothing a user can perceive. **Can the interface paint without it?** A deferred section needs a design that looks deliberate while it is missing: a fixed-height skeleton that the real content replaces without reflowing the page. If the deferred field is what the screen exists to show — deferring the transaction list on a transactions page — the user watches a spinner either way, and you have added machinery to change what they are watching it inside of. **What else assumes one payload per request?** Response caching keyed on a whole body no longer applies. A request timer that stops at the first byte and one that stops at the last now measure two different things and must be named differently. Anything that infers success from the status code will misread a response whose later payload carried a field error, because the 200 was committed before that field ran. ## The failure mode to plan for A truncated stream does not look like an error. The client holds a well-formed tree, a 200 status, no error entry, and a last-read payload saying more is coming. On that statements screen, the breakdown panel spun indefinitely for a small slice of sessions whose connection dropped mid-body, and it produced no error-rate signal at all — the request had already succeeded as far as every dashboard was concerned. The fix is client-side and must be designed in from the start: a deadline for the outstanding payloads that turns silence into either a visible error or a plain follow-up operation fetching the deferred fields alone. "There is no terminal state when a payload never arrives" is the defect, and it is invisible in testing because test connections do not die. ## The alternatives you must have considered **Two operations.** The client sends the fast document and the slow document in parallel. It costs a second set of request headers and re-resolves whatever parent the slow field hangs from, but it needs no unratified directive, each response has its own status, its own cache entry, its own timing and its own retry. For a great many screens this is simply the better answer, and a senior candidate who reaches for it first is not being unambitious. **Make the field fast.** A 610 ms aggregate computed on every request is often a precomputed rollup waiting to happen. Deferring an expensive field can quietly become the reason nobody ever fixes it. **Do not select it.** The cheapest deferral is a client that asks for the breakdown only when the user opens the panel. `@defer` wins specifically when the slow part *depends on the parent you are already resolving*, so a second operation would duplicate that work, and the page has a genuine two-stage design. It is also, honestly, a choice about maturity: the directives are in no released specification edition, the payload envelope has been revised across drafts, and support has to exist end to end — in the server, in the client, and in whatever the response travels through. Adopting it is a bet that the ecosystem catches up, made for a latency win you should be able to state in milliseconds before you place it.
- What is the simplest alternative when one field is slow, and when would you prefer it?Two operations sent in parallel: a fast document for the shell and a second one for the expensive section. It costs another set of request headers and re-resolves whatever parent the slow field hangs from, but each response has its own status, cache entry, timing and retry, and it needs no unratified directive. Prefer it unless the slow part genuinely depends on work the first operation is already doing.
- What breaks in observability once responses become multi-payload?The word duration stops having one meaning — time to the initial payload and time to the final payload are different numbers and both matter, so the metric must be split rather than renamed. Error rates also drift, because a field error delivered in a later payload arrives after a 200 was committed and never reaches a status-based error counter.
- Does deferring reduce load on the server?No. Every deferred field resolves exactly as it would have, against the same backends, and the server holds the request open longer while paying extra framing and serialization per payload. If the goal is less work rather than earlier paint, the answers are precomputing the expensive field, or not selecting it until the user asks for it.
saying these in an interview costs you the question
- Defers a section the page cannot render without
- Claims deferring reduces total server work
- Judges the change by request duration alone
- Leaves no terminal state when a payload never arrives
- Defers a field only milliseconds slower than its siblings
- Treats a committed 200 as proof the response completed