skip to content

When does GraphQL selection-set projection cost more than the over-fetch it saves?

level: seniorimportance: should knowfreq 36%

answer

  1. Optimizations are argued with numbers
  2. Every shape is another statement
  3. Plan caches reward repetition, not variety
  4. Untested shapes fail as quiet nulls
  5. Derive the mapping; fail at startup

basics

~20 s

When the saved bytes are small and the mapping from selected fields to columns is hand-written. You trade a millisecond for many distinct statements, weaker plan reuse, and silent nulls when a field goes unmapped.

solid answer

~50 s

Projection driven by lookahead is worth it when one wide read or one expensive join dominates a measured trace — dropping a case-file resolver from 1,830 ms to 118 ms by eliding a join is unarguable. It stops being worth it when the win is small and the costs are structural. Every distinct client shape becomes a distinct backend statement, so plan and statement caches see variety instead of repetition and slow-query analysis fragments across shapes. Resolver behaviour now depends on document shape, so a shape no test exercises is a code path no test exercises. And the failure mode is silent: a field whose column was never mapped resolves to null rather than raising. The safe form is a **total, schema-derived mapping** that fails loudly on an unmapped field, plus a fixed set of always-fetched columns; the unsafe form is a hand-written switch maintained by a small team.

code

pseudocode · 22 lines
pseudocode
# Hand-written and partial: the fallthrough is the bug
COLUMN_FOR = {
    FieldKind.FILE_NUMBER: "file_number",
    FieldKind.STATUS:      "status",
    FieldKind.FILED_AT:    "filed_at",
    # FieldKind.SEAL_REASON added to the enum last release; never mapped here
}

function projection(selectedFieldNames):
    columns = ["id"]
    for name in selectedFieldNames:
        column = COLUMN_FOR.get(fieldKindOf(name))
        if column != null:
            columns.add(column)          # unmapped fields vanish silently
    return columns

# Total and derived: an unmapped field cannot reach production
function buildProjector(schemaType, columnMap):
    for field in schemaType.fields:
        if columnMap.missing(field.name) and not hasDedicatedResolver(field):
            fail("no column mapping for " + schemaType.name + "." + field.name)
    return projectorFrom(columnMap)

go deeper

for a junior

Understand the tradeoff in one line: fetching only the selected fields is faster, but the resolver now has to be right about what was selected, and being wrong shows up as missing data.

for a middle

Be able to name the concrete costs — more distinct statements, weaker plan-cache reuse, code paths that only some document shapes reach — and explain why a missing column becomes a null rather than an error.

for a senior

Lead with the measurement and the elision-versus-column-trimming split, then show how you keep it safe: a total mapping checked at startup, an always-fetched floor for non-null fields, projection-level tests and a flag to fall back to the full read.

for a principal

Own the standard rather than the instance. Decide whether this optimization is allowed to be hand-written at all on a small team, given that its failure mode is invisible, and require that any resolver whose behaviour depends on document shape ships with generated-shape tests.

## Start from the measurement, not the technique Lookahead projection is an optimization, and optimizations are argued with numbers. The honest question is what fraction of the field's latency is the over-fetch. On a legal case-file graph, a `case(id:)` resolver that eagerly joined parties, filings and hearings measured 1,830 ms at the 95th percentile; eliding the joins nobody selected brought it to 118 ms. That is a fix worth structural cost. Trimming a 47-column row to 9 columns on a primary-key read, and watching 47 ms become 41 ms, is not — the row was already in the buffer cache and you have bought six milliseconds with a permanent maintenance liability. So the first split is **elision versus projection**. Skipping a join, an aggregate or a whole downstream call because that branch was not selected is usually a large, stable win. Narrowing the column list of a read you were making anyway is usually a small one. Interviewers ask this question to see whether you reach for the technique or for the trace. ## Cost one: shape explosion Without projection a resolver issues one statement. With it, it issues one per distinct selection shape. A production graph with a handful of screens can easily produce two dozen shapes against one table, and each is a separate statement text. Prepared-statement and plan caches reward repetition, and now they see variety. Index coverage varies by shape: the nine-column shape is covered by an index, the twelve-column one is not, and only that shape is slow. Slow-query analysis fragments the same logical read across many entries, so the read never rises to the top of any list even when it is collectively the most expensive thing you do. None of this is fatal. It is the reason to prefer a **small, bounded set of projections** — for example a narrow and a wide variant chosen by whether any expensive branch was selected — over a bespoke column list per request. ## Cost two: behaviour that depends on document shape Once the resolver reads the document, the document is an input to the resolver. A shape no test sends is a code path no test runs, and the space of shapes is combinatorial. This is where the characteristic incident comes from. A hand-written mapping turned selected field names into columns through a switch over a small enum of known fields. A release added a `SEALED` value to the case status enum and, alongside it, a `sealReason` field — plus a new member of that internal field enum. The mapping had a fallthrough branch that quietly returned no column. The projection therefore never selected `seal_reason`, the default resolution found no such property on the loaded row, and the field resolved to `null` on every request. Because `sealReason` is nullable, nothing errored. The suite stayed green because its documents predated the field. A four-person platform team found it three weeks later, from a support ticket rather than from a monitor. That is the shape of every projection bug: **wrong projection surfaces as absent data, not as a failure.** Compare it with the over-fetching version of the same resolver, which cannot have this bug at all because it always loads everything. ## Making projection safe when you do want it * **Make the mapping total and derived.** Build the field-to-column map from the schema at startup and fail fast if any field of the type has no mapping and no dedicated resolver. An unmapped field must be a startup error, never a runtime fallthrough. * **Always fetch a floor.** Include the identifier and every column the type's non-null fields depend on, regardless of selection. Non-null fields turn a missing column into a propagating error, which is the worst version of the failure. * **Prefer eliding branches to trimming columns.** Fewer decision points, larger win, and a missed elision costs latency rather than correctness. * **Test projection, not just results.** Assert the generated projection for documents that use an alias, a named fragment and a conditional directive. Generating shapes from the schema catches the unmapped-field case that a fixed suite never will. * **Watch the interaction with per-request caching.** If a resolver caches loaded objects by identifier, a partially projected object can be handed to a later resolver that needs a column the first fetch never asked for. Either include the projection in that cache key or never cache partial objects. The batching layer itself is a separate subject; the hazard here is specifically that projection makes a cached object non-interchangeable with itself. * **Have an off switch.** A flag that reverts to the full read lets you neutralise a projection bug in minutes instead of shipping a fix. ## The judgement to state out loud Projection buys latency with coupling. Buy it where a trace shows the over-fetch dominates, keep the mapping mechanical rather than hand-maintained, and remember that the failure mode is a quiet null. On a small team, an optimization whose bugs are invisible needs a stronger justification than one whose bugs shout.

  • How would you catch an unmapped field before it reaches production rather than three weeks after?
    Make the mapping total and check it at startup: walk the type's fields and fail to boot if any field has neither a column mapping nor a dedicated resolver. Back that with tests that assert the generated projection for schema-derived document shapes rather than a fixed handful, so a newly added field is exercised the moment it exists.
  • Which is usually the bigger win, trimming columns or skipping a join, and why?
    Skipping a join, an aggregate or a whole downstream call, by a wide margin. A row you were already reading costs little extra per column, especially when it is already cached, whereas an unnecessary join multiplies work and can dominate the field's latency. Elision also has fewer decision points, so it is easier to keep correct.
  • Why is a wrong projection more dangerous than an over-fetching resolver?
    Because it fails silently. A column the projection dropped is simply absent, so a nullable field resolves to null and nothing errors; a non-null field raises and propagates, which at least is visible. An over-fetching resolver cannot produce either symptom — it is slow, and slow is measurable. Optimizations whose bugs are invisible need stronger justification.
  • When would you decline projection entirely and solve the over-fetch another way?
    When the wide read is unavoidable but cheap, or when the expensive part is one field rather than the row. Then give that field its own resolver so its cost is explicit and only paid when selected, or move the aggregate behind a materialized view. Both keep the parent resolver's behaviour independent of document shape.

saying these in an interview costs you the question

  • Projects from the selection set without measuring the over-fetch first
  • Hand-maintains a field-to-column switch with a silent fallthrough
  • Ignores that many query shapes hurt plan and statement cache reuse
  • Assumes tests cover shapes no test document ever sends
  • Caches partially projected objects under an identifier-only key
  • Treats a dropped column as an error case rather than a silent null

context