skip to content

In a GraphQL server, why can a field that only reads its parent object still hit the database?

level: seniorimportance: should knowfreq 40%

answer

  1. No code does not mean no work
  2. The cost belongs to the object
  3. Watch a property read fetch on access
  4. Calls scale with parent rows, not requests
  5. Time every field, including the resolver-less ones

basics

~20 s

Trivial describes the code, not the work performed. Reading a property off the parent runs whatever the object does on access, so a lazily loaded association fetches on the read - once per parent object, invisible in the resolver map.

solid answer

~50 s

Most fields in a real schema are trivial: one field loads an object and everything beneath is a plain read of it. That stays cheap only while the loaded object is fully materialised. If it is a lazily loading record - an association proxy, a getter that calls a cache or a service, a virtual property - then the read itself performs I/O, once per parent object, and nothing in the resolver map hints at it. A list of 2,614 animals with `herd { name }` selected therefore issues thousands of small fetches while the code shows only a single root resolver. Confirm it with per-field timing: a field you never wrote code for shows real latency, and call counts scale with parent rows, not requests. Fix it by materialising what the subtree needs when the parent is loaded, or by giving the field an explicit resolver that fetches deliberately and can be batched.

code

pseudocode · 14 lines
pseudocode
// One resolver, one query - looks like a single fetch
resolve Query.animalsByBirthYear(parent, args, context, info):
    return context.store.findAnimals(year: args.year)   // 1,138 records

// No resolver written for Animal.herd, so the value is read off the parent.
// The store's record type resolves the association on access:
class AnimalRecord:
    property herd:
        if not loaded(_herd):
            _herd = store.findHerd(this.herdId)   // one query, per record
        return _herd

// Selecting herd { name } therefore issues 1,138 single-row fetches
// from code that contains no fetch at all.

go deeper

for a junior

Remember that a field without its own resolver still reads a property off the parent object, and that reading a property can run code. Cheap in the schema does not mean cheap at runtime.

for a middle

Be ready to explain the mechanism: the value comes off the parent, the parent may load associations on access, and the resulting fetch happens once per parent object rather than once per request.

for a senior

Demonstrate the diagnosis. Per-field timing on fields you never wrote, call counts that scale with row counts, a controlled comparison with the field deselected - then a fix that materialises the data at load time rather than a cache bolted on top.

for a principal

Own the invariant rather than the incident: nothing handed to the execution layer may perform I/O on property access. Decide whether that is enforced by a mapping boundary or by convention, and make sure instrumentation covers fields with no code.

## "Trivial" is a claim about source code, not about cost A field is called trivial when nobody wrote a resolver for it: the value is taken straight off the parent object the field above produced. In a pedigree service, `Query.animal` fetches an animal and then `tag`, `name`, `birthWeightKg`, `sire` and `herd` are all reads of that one object. This is the ordinary shape of a GraphQL server - one loading field feeding a whole subtree of trivial ones - and it is why resolver maps are so much smaller than schemas. The assumption underneath is that reading a property is free. That assumption belongs to the *object*, not to GraphQL, and plenty of objects break it. ## How I/O hides behind a property read The field read is a call into whatever the parent object is. Three common shapes make that call do work: - **A lazily loaded association.** A record mapped from a store may expose related records as proxies that fetch on first access. Reading `animal.herd` issues a query. - **A computed getter.** A property that calls a service, consults a cache or derives a value from another system on each read. - **A wrapper around a remote object.** In a platform assembled from 11 services, an object handed to the schema layer may be a facade whose accessors call the owning service. In all three the GraphQL layer is doing exactly what it says: read the property. The cost is real, it scales with the number of parent objects, and it appears nowhere in the resolver code, which is why it survives review. ## The pedigree incident A breeding registry exposed `Query.animalsByBirthYear` returning animal records straight from its store. Adding `herd { name }` to a report document took the endpoint from 47 ms to over nine seconds. The document had grown by two fields; the resolver map had not changed at all. What changed was that a selection now touched a lazily mapped association: for the 1,138 animals in the largest year, the schema layer performed 1,138 single-row fetches, one per property read, all of them inside code nobody had written. The telling detail is the shape of the numbers. Backend call count tracked the *number of parent objects* rather than the number of fields in the document or the number of requests. That signature - work proportional to rows, from a field with no code - is the fingerprint of a trivial field that is not trivial. ## How to confirm it 1. **Per-field timing.** Instrument execution so that every field is timed, including the ones with no resolver of their own. A field you never wrote code for showing tens of milliseconds is the whole diagnosis. Teams that exclude "resolver-less" fields from their instrumentation to reduce noise delete exactly the evidence they need. 2. **Count backend calls per request.** Run the document twice, once with the suspect field selected and once without, and compare the call count. A jump proportional to the number of parent rows localises the field immediately. 3. **Inspect what the loading field actually returned.** The question is whether the object is fully materialised or a proxy. If it is the store's own record type, assume lazy behaviour until proven otherwise. 4. **Look at the store's own logs.** Thousands of identical single-row statements differing only in a key is unambiguous. ## Fixes, in order of preference - **Materialise what the subtree needs at load time.** If the parent-loading field can fetch the associated data in the same round trip, the trivial reads become genuinely trivial again. This is the fix that keeps the resolver map small. - **Make the field explicit.** Give it a real resolver that fetches deliberately. This does not by itself reduce the number of fetches, but it moves the work into a place where it is visible, reviewable and batchable - and per-request batching of exactly this pattern is a whole discipline of its own. - **Change what the schema layer is handed.** Map store records into plain, fully materialised values at the boundary, so no object reaching a resolver can perform I/O on a property read. This is the strongest guarantee and costs a mapping layer. ## The general principle "No resolver" is not a performance property. The real invariant a server should hold is *no object visible to the execution layer performs I/O on property access* - and once that invariant holds, the trivial-resolver majority is exactly the cheap projection everyone assumes it already is. Stating that invariant, rather than listing tools, is what marks a senior answer.

  • How do you tell this apart from a root fetch that is simply slow?
    By where the time and the calls sit. A slow root fetch shows one expensive field and a backend call count that does not move when you add or remove deeper fields. Hidden lazy loading shows latency on a field with no code, and a call count proportional to the number of parent objects, which changes the moment that field is deselected.
  • Does giving the field an explicit resolver fix the problem on its own?
    No. It relocates the fetch into visible, reviewable code but the count still scales with the number of parents. It is worth doing because it makes the work explicit and gives you one place to preload or batch from, but the count only drops when the data is fetched together rather than per parent.
  • What does this imply for how you instrument fields?
    Instrument every field, not only the ones with resolvers. Teams commonly filter resolver-less fields out of tracing to cut span volume, which is precisely the class of field that hides I/O. If span volume is the concern, sample requests rather than excluding a category of field.
  • Why does code review rarely catch this?
    Because the diff that triggers it is usually a document or a schema change, not a resolver change - two fields added to a query, or a field exposed on a type. The offending behaviour lives in the record class the schema layer was handed, which was written long before and looks correct in isolation.

It looks like reading a number off a page, but the page is a window: every glance sends a runner to the archive. The code shows a glance; the archive shows a thousand trips.

saying these in an interview costs you the question

  • Says a field with no resolver does no work
  • Assumes the parent-loading fetch materialised everything
  • Blames GraphQL rather than the lazy property read
  • Measures only whole-request latency, never per field
  • Excludes resolver-less fields from tracing as noise
  • Adds a cache before locating the repeated fetch

context