skip to content

A slow dependency turns whole GraphQL responses null at peak — how do you reshape the schema to degrade partially instead?

level: seniorimportance: should knowfreq 47%

answer

  1. data present and null means execution ran
  2. Walk the chain, not just the field
  3. Ask which unit still renders without it
  4. Radius is the schema, frequency is the runtime
  5. Then sweep for the other cross-boundary fields

basics

~20 s

Find the unbroken Non-Null chain from the failing field to the root, decide the smallest unit consumers can render without that field, and make the link at that boundary nullable. Then bound the dependency with a timeout so it fails fast into that hole.

solid answer

~50 s

Treat it as two fixes: the schema decides how much is lost, the runtime decides how often. First confirm the shape of the loss — a response carrying `data: null` alongside an error entry whose path points deep into the graph tells you the chain from that field to the root has no nullable link. Map that chain and pick the boundary: in a payroll and benefits graph the benefits panel is expendable per payslip, so `payslip.benefitElections` loses its `!` and each affected payslip arrives with a hole while the rest of the pay run renders. Coordinate the change with consumers, since a field that was never null becoming null is a change they must handle. Then attack the frequency: a timeout and a circuit breaker around the provider so the field resolves to null in milliseconds at a 1,200-request-per-minute peak instead of holding the whole request open.

code

json · 9 lines
json
{
  "data": null,
  "errors": [
    {
      "message": "Benefits provider timed out after 2500ms",
      "path": ["payRun", "payslips", 417, "benefitElections"]
    }
  ]
}

go deeper

for a junior

Be able to recognise the symptom: a response with data present and null plus an error entry means a field failed and nothing on the way up could hold the hole. You are not expected to design the fix yet.

for a middle

Explain how to map the chain from the failing field to the root and why removing one ! at the right link changes the outcome. Be ready to say why the change affects existing consumers.

for a senior

Demonstrate that you separate blast radius from failure rate: the schema decides how much is lost, the timeout and breaker decide how often. Talk about choosing the boundary with the people who own the screens, and about sweeping the rest of the schema afterwards.

for a principal

Own the systemic view — a lint rule that no cross-boundary field sits under an unbroken Non-Null chain, field-level error-rate metrics as evidence of mis-marked fields, and the coordination cost of widening nullability across many consuming teams.

## What the symptom is telling you A response that arrives as `data: null` with a single error entry is not a transport problem and not a validation problem. Validation failures never execute, so they leave `data` absent rather than null. A `data` that is present and null means execution started, a field could not be produced, and the resulting hole had no nullable position to settle in between that field and the root. The error entry's path names where it started; the schema tells you why it did not stop. At low traffic this is invisible: the dependency answers, the promise holds, nobody notices that the schema has no firebreak. At a 1,200-request-per-minute peak the same dependency's tail latency crosses the resolver's timeout and every request that touches the field returns a blank page. The schema did not change; the failure rate did. That is why this shows up as an incident rather than as a review comment. ## Step one: map the chain, not just the field Take the path from the error entry and walk it in the schema, link by link, from the failing field up to the root operation field. Write down each link's nullability. What you are looking for is the first nullable link — the point at which the loss would have stopped. If there is none, the blast radius is the whole response. In a payroll and benefits graph the chain often looks like `payRun → payslips → benefitElections`, all Non-Null, with the elections served by an external benefits provider and everything else served from the payroll database. One unavailable provider costs thousands of correctly computed payslips. ## Step two: choose the boundary deliberately The schema change is not "make things nullable" but "decide the smallest unit a consumer can still render." Ask the people who build the screens. A payslip without its benefits panel is a usable payslip; a pay run without its payslips is not. So the `!` comes off `benefitElections`, and nothing else moves. Each affected payslip now carries a hole, an error entry names the exact path, and every other field in the response survives. Going coarser — making `payslips` nullable — would also stop the bubble, but it pays the whole pay run to protect one panel. Going finer is not possible: the field is the smallest thing there is. The change has a consumer cost. A field that was Non-Null and becomes nullable widens what every existing consumer must handle, so it is coordinated work rather than a quiet schema push — announce it, give consumers a window, and confirm that the screens actually render the hole rather than crashing on it. ## Step three: reduce the frequency, not just the radius A firebreak converts an outage into a degradation; it does not make the dependency reliable. The second half of the fix is at the resolver: * A **timeout** shorter than the client's patience, so the field yields a null quickly instead of holding the request open. Under load, an unbounded call is worse than a failed one, because slow requests pile up and cost the whole endpoint. * A **circuit breaker** so that once the provider is clearly down, the field fails immediately instead of every request paying the timeout. * A **fallback** where the domain allows one — a last-known-good value with a staleness marker is often better than a hole, and it is a schema decision too, because it changes what null means. ## Step four: prevent the next one The incident is evidence about the whole schema, not just this field. Sweep for other fields backed by a different service or a job that can lag, and check whether each one sits under an unbroken Non-Null chain. A schema linter can enforce a house rule mechanically — for example, that no field resolved across a service boundary may be Non-Null, or that no Non-Null chain from such a field may reach the root. Field-level error-rate metrics make the same point after the fact: a Non-Null field with a nonzero error rate is a mis-marked field by definition, because every one of those errors cost more than itself. ## What to say about the specification None of this is specified behaviour beyond the propagation rule itself. The specification says what happens when a Non-Null position cannot be filled; it does not say where the author should have put the `!`, and it has nothing to say about timeouts, breakers or fallbacks. Those are operational conventions layered on top, and an interviewer will notice whether you attribute them correctly.

  • Why not make the root operation field nullable and be done with it?
    Because that is where the failure was already going. A nullable root only changes `data: null` into `data: { payRun: null }`; the pay run and every payslip in it are still gone. A firebreak is only worth what it saves, so it belongs at the smallest unit a consumer can still use — the panel, not the page.
  • The screens already show a blank page. Do you ship the schema change or the timeout first?
    The timeout and breaker first, because they are server-side only and reduce the failure rate immediately without asking any consumer to change. The nullability change is the durable fix but it widens what every consumer must handle, so it needs coordination and a window. Shipping the runtime bound first also buys time to confirm which boundary the screens actually want.
  • How would you find the other fields with the same exposure before they fail?
    Enumerate the fields whose resolvers cross a boundary — another service, a cache that can miss, a job that can lag — and for each, walk upward to the root looking for a nullable link. A schema lint rule can encode the same check so it runs on every schema change, and per-field error-rate metrics catch the ones the rule missed: a Non-Null field with a nonzero error rate is by definition costing more than itself.

saying these in an interview costs you the question

  • Blames the transport for a data null response
  • Makes the root field nullable and calls it fixed
  • Adds retries without a timeout or a breaker
  • Returns an empty object to satisfy the Non-Null promise
  • Pushes the nullability change without telling consumers
  • Fixes the one field and never sweeps the schema

context