skip to content

How do you set a nullability policy for a large GraphQL schema owned by many teams?

level: principalimportance: should knowfreq 38%

answer

  1. Local marker, global consequence
  2. Key the rule off where the value comes from
  3. State it as a chain property
  4. Reviewers stop seeing one character
  5. Non-Null field with errors is mis-marked

basics

~20 s

Write a rule that derives ! from where the data comes from — Non-Null only for values produced with their parent, nullable across any boundary — then enforce it with schema lint in the build and check it against per-field error rates.

solid answer

~50 s

The specification is silent on placement, so the organization has to supply the rule, and the rule has to be derivable rather than stylistic. The one that survives contact with many teams ties nullability to the data source: a field produced from the same load as its parent may be Non-Null; a field that crosses a boundary — another service, a cache that can miss, a job that can lag — is nullable, and so is every ancestor between it and the root. Enforce it in the build with a schema linter rather than in review, because a `!` is one character and reviewers stop seeing it. Verify it in production with per-field error rates: a Non-Null field with a nonzero error rate is mis-marked by definition. And resist the opposite extreme — blanket nullability makes null ambiguous and pushes a branch into every consumer, so the policy must be a budget, not a ban.

code

pseudocode · 5 lines
pseudocode
for field in schema.allFields:
    if field.isFallible:                      # crosses a service, cache or job boundary
        for path in pathsFromRootTo(field):
            if every link in path is NonNull:
                fail(field, "fallible field under an unbroken Non-Null chain")

go deeper

for a junior

You are not expected to set policy, but know that where ! goes is a decision your organization makes, not something the specification dictates, and that the rule usually depends on where the field's data comes from.

for a middle

Be able to apply the rule to a field you are adding: say where its value comes from, whether it can fail independently of its parent, and what the chain above it looks like. That reasoning is what a policy is asking each author to do.

for a senior

Show how you would enforce and verify — lint in the build rather than review, per-field error rates as evidence, and a ranked migration for fields that are already wrong. Be ready to explain why widening a field to nullable is coordinated work.

for a principal

Own the cross-team framing: a marker written by one team decides how much of another team's response survives, so the chain property needs an owner, a default that is safe under inattention, and an exception record. Weigh it against the real cost of blanket nullability to consumers.

## Why a policy is needed at all On a schema owned by one team, nullability placement is taste and it works out. On a schema many teams contribute types to, it stops working out for a structural reason: **the blast radius of a `!` is not felt by the team that wrote it.** An author marks a field Non-Null because their screen wants it; the response it blanks belongs to a consumer they have never met, reached through ancestors owned by a third team. Local decisions, global consequence. That is the definition of a thing that needs a written rule. ## What the rule should key off Stylistic rules ("prefer nullable", "prefer non-null") do not survive, because they give a reviewer nothing to check. A rule that keys off the **provenance of the value** does: * A field produced from the same load as its parent — an identifier, a column of the row in hand, a value computed from data already present — **may** be Non-Null. If the parent exists, it exists. * A field that crosses a boundary — another service, a queue, a cache that can miss, a nightly job that may lag — **is** nullable, and no unbroken Non-Null chain may run from it to the root. The second clause is the one that carries the policy, and it is worth stating as a chain property rather than a per-field property. A nullable risky field with Non-Null ancestors is still a blank response waiting to happen for anything nested below it; the property you actually want is that between every fallible field and the root there is at least one nullable link. ## Where to enforce it Not in code review. A `!` is one character, it appears hundreds of times in a large schema, and reviewers habituate within a week. Enforcement belongs in the build: * A **schema linter** in the pipeline that fails the build on a violation. Rules of the shape *"this field's resolver crosses a boundary and every ancestor to the root is Non-Null"* need a way to know which fields are risky — usually an annotation or a naming convention that the team applies once and the linter reads. * A **required justification** for exceptions, recorded next to the field rather than in a ticket, so the next reviewer sees why. * **Ownership review at the boundary**, because in a schema assembled from many teams' contributions the chain crosses ownership lines. Someone has to own the chain property that no single team can see. ## How to know the policy is working Two signals, and the second is the honest one. * **Static**: the number of fallible fields sitting under an unbroken Non-Null chain, tracked over time. A nullability budget in this sense is a count you can drive toward zero. * **Runtime**: per-field error rates. A Non-Null field with a nonzero error rate is mis-marked by definition, because each of those errors cost more than the field — it took a region of somebody's response with it. This is the metric that turns the policy from an opinion into a measurement, and it is worth wiring before the lint rule, because it tells you where to start. ## The cost of the other extreme A policy of "make everything nullable" is easy to write, easy to enforce, and quietly expensive. Every nullable position is a branch in every consumer; a typed client generator turns the whole schema into optionals, and the type safety that made teams adopt GraphQL evaporates into null checks. Worse, null becomes ambiguous everywhere: a null may mean the value legitimately does not exist or that producing it failed, and only an error entry naming that path distinguishes them — so consumers that do not read the errors list under-report silently. So the policy is a budget with two sides. Spend nullability where it buys a firebreak in front of something that can genuinely fail; save it where the value cannot fail independently of its parent, and let consumers have the guarantee. ## The organizational part The last piece is the part interviewers are actually probing. A nullability policy is a cross-team availability contract, so it needs an owner, a place to record exceptions, a migration path for fields that are already wrong (widening a Non-Null field to nullable is a change every consumer must absorb, so it is coordinated work with a window, not a quiet push), and a default for new fields that is safe when nobody thinks about it. Set the default to nullable for anything crossing a boundary and you get the safe outcome from inattention, which is the only kind of policy that holds at scale. And say plainly that none of this comes from the specification. GraphQL defines what Non-Null means and what happens when the promise breaks; every word of the policy above is convention your organization chose.

  • How does the linter know which fields are fallible?
    It cannot infer it from the schema, so the information has to be supplied. In practice teams annotate the field — a schema directive, a metadata file or a naming convention — to declare that its resolver crosses a boundary, and the linter reads that. The declaration is cheap because the team writing the resolver already knows, and it has the useful side effect of making the risky surface of the schema enumerable for the first time.
  • A team insists their cross-boundary field must be Non-Null because their screen cannot render without it. What do you do?
    Separate the two claims. "Our screen needs it" is a client concern and is satisfied by the client treating a null as fatal for that view. "The schema must guarantee it" is an availability claim about a dependency they do not control, and it imposes the blast radius on every other consumer of the same ancestors. The usual resolution is a nullable field plus a documented client contract, with the exception recorded if they still want it.
  • How do you migrate a schema that is already over-marked?
    Rank by evidence rather than by sweep: order fallible Non-Null fields by observed error rate times the size of the region each one blanks, and fix the top of that list first. Each change widens what consumers must handle, so it needs an announcement and a window; batching several into one coordinated release is usually kinder than a trickle. Freeze the default for new fields first, so the backlog stops growing while you work it.

saying these in an interview costs you the question

  • Answers with a style preference and no derivable rule
  • Relies on code review to catch a single character
  • Proposes making everything nullable as the policy
  • Ignores that one team's marker blanks another's response
  • Never measures per-field error rates
  • Attributes the policy to the GraphQL specification

context