For a nested GraphQL connection, why can't one lookup with a single row cap serve first: 10 per parent?
answer
- Whose limit is being applied
- Skew decides who gets the rows
- Empty pages that look like data
- Rank within each parent, not overall
- Fetch one extra row per parent
basics
~10 sA cap on the whole lookup is a global cap: its rows can all belong to one parent, leaving the rest with empty pages and hasNextPage false. Each parent needs its own top-N slice.
solid answer
~50 sThe requirement is twenty-three independent slices, and one capped lookup produces one. Ask the store for the scan events of twenty-three shipments "limit ten" and you get the ten earliest overall; for a waybill carrying 41,900 events those ten are all its own, and the other twenty-two shipments come back as empty connections with `hasNextPage: false` — well-formed, plausible, and wrong. Three honest strategies exist. One lookup per parent with its own limit: always correct, one round trip per parent. One lookup that fetches **all** children of all parents and trims to ten each in memory: correct only if nothing caps it, so its cost is the total child count — fine when fan-out is small and bounded, ruinous under skew. Or a top-N-per-parent lookup the store evaluates itself: correct and bounded in one round trip. Whichever you choose, fetch eleven per parent so `hasNextPage` is answerable.
code
pseudocode · 14 lines# nested connection resolver: first N per parent, N+1 fetched
resolveScanEvents(parents, first, afterByParent):
slices = store.topNPerParent(
parentIds = parents.map(p -> p.id),
order = [scannedAt ASC, eventId ASC],
startAfter= afterByParent, # a start key per parent, not one
limit = first + 1) # the extra row answers hasNextPage
for p in parents:
rows = slices[p.id]
more = rows.size > first
emit connection(
edges = rows.take(first).map(r -> edge(cursor(r), r)),
hasNextPage = more)go deeper
Understand first that the nested list belongs to one parent, so a limit applied to a shared lookup is not the limit the document asked for. Expect a lopsided parent to swallow the whole allowance.
Be able to name the three fetch shapes — one lookup per parent, over-fetch and trim, and a top-N-per-parent lookup — and say what each bounds, plus why an extra row per parent is needed to answer hasNextPage.
Argue it as correctness, not performance: describe the silent empty pages, why test fixtures hide them, and how you would detect the pattern in production from rows-read versus rows-emitted and from parents reporting no children.
Decide where the ranking belongs. Pushing top-N-per-parent into the store bounds rows and round trips but couples the schema to store capability; keeping it in the service is portable and pays in transferred rows. Say which you would standardise on and why.
## The bug: a global cap masquerading as a per-parent slice A nested connection field is executed once per parent, but a naive implementation reaches for one round trip: gather the twenty-three shipment ids from the outer page, ask the store for their scan events "with a limit of ten", and hand each parent its slice. The limit in that sentence is the problem. A cap applied to the whole lookup is a **global** cap. Ordered by time, the ten rows it returns can all belong to one shipment — and in freight data they usually do, because a single long-haul waybill can carry 41,900 scan events while a domestic one carries nine. What the client sees is not an error. It is twenty-two shipments whose `scanEvents` connection is an empty `edges` array with `hasNextPage: false` — a perfectly well-formed statement that those shipments have no scan events. The data is wrong and it looks right. Nothing in the response body flags it, no resolver threw, and monitoring stays green. ## Why one flat range predicate cannot express it The requirement is "the first ten children **of each** parent" — twenty-three independent ordered slices with twenty-three independent start points. A single flat lookup expresses one ordered slice. If the inner field also carries `after:` cursors that differ per parent, the twenty-three start keys differ too, and no single range predicate covers them. The per-parent-ness is intrinsic to what a nested connection means, so the fetch strategy has to reproduce it. ## The three honest strategies **One lookup per parent.** Twenty-three lookups, each with its own start key and its own limit. Always correct, trivially reasonable, and the round-trip count grows with the outer page. Collapsing those repeated lookups into one round trip is a distinct concern with its own machinery; correctness does not depend on it. **One lookup, over-fetch then trim.** Fetch every child of all twenty-three parents ordered by parent and then by the sort key, and trim in memory to ten per parent. This is correct **only if you truly fetch them all** — the moment you add a safety cap to bound the fetch you are back to the first bug. Its cost is therefore the total child count, not the requested one: 41,900 rows crossing the wire to produce ten. It is a legitimate choice exactly when the fan-out is known to be small and bounded — a shipment has at most a dozen legs, so over-fetching legs and trimming is fine, and nobody should be paging them at all. **A top-N-per-parent lookup.** One round trip in which the store itself ranks children within each parent and returns at most eleven per parent. Correct and bounded at the same time; the mechanism the engine uses to evaluate it is the database's own concern. This is the strategy that scales, and it is the one an interviewer is usually listening for after the other two have been dismissed. ## Fetch N+1, not N `hasNextPage` on a parent's connection is the claim "there is an eleventh row for this parent". You cannot answer it from ten rows. Every strategy above therefore requests `first + 1` per parent and discards the extra after using it to set the flag. A second count query per parent is the wrong answer: it doubles the lookups to learn one boolean. ## Why over-fetch-then-trim keeps getting shipped anyway It is the shape that works in the test fixture. Seed data gives every shipment three or four scan events, so trimming to ten in memory is exact, fast, and passes review. The failure only appears against production skew, and it appears as a performance cliff (one shipment dragging 41,900 rows into memory) or, if someone "fixes" that with a cap, as silent data loss. If a change adds a nested connection, the test that matters is the one with a deliberately lopsided parent. ## How to detect it in a running system Two signals. First, the ratio of rows fetched to rows returned per field — a nested field whose resolver reads far more than it emits is over-fetching and trimming. Second, empty nested connections on parents known to have children; a spot check against the store finds the global-cap bug immediately. Both are per-field measurements, which is why field-level instrumentation is worth more on nested connections than anywhere else in a schema. ## The judgement being tested The interviewer wants to hear that "the slice is per parent" is a **correctness** property, not an optimisation, and that you can rank the strategies by what each one bounds: per-parent lookups bound rows but not round trips, over-fetch bounds round trips but not rows, and a top-N-per-parent lookup bounds both at the cost of pushing the ranking into the store. Ordering also has to be total per parent for any of them to resume correctly, which is a prerequisite rather than a strategy.
- When is fetching every child and trimming in memory actually the right call?When the fan-out is known to be small and bounded by the domain — a shipment has at most a dozen legs, so reading them all and trimming is exact and cheap, and arguably the field should not be paged at all. It stops being right the moment one parent can hold tens of thousands of children, because the cost is the total child count rather than the requested one.
- How would you catch this bug in a running service rather than in review?Two per-field signals. Compare rows read against rows emitted for the nested field: a resolver reading far more than it returns is over-fetching and trimming. And alert on nested connections that come back empty for parents known to have children — the global-cap bug shows up immediately as a population of parents reporting no children at all.
- Why fetch first + 1 rows per parent instead of running a count?Because hasNextPage is one boolean — "is there an eleventh row for this parent" — and one extra row answers it exactly. A count query per parent doubles the lookups, is more expensive than the slice itself on a large child set, and can disagree with the slice if rows change between the two reads.
- What must be true of the child ordering for any of these strategies to resume correctly?The order must be total, not merely specified: a sort key with ties gives no single resume point, so a cursor built from it can skip or repeat rows on the next page. Appending a unique tie-breaker to the sort makes every position unambiguous, which is a prerequisite for per-parent slicing rather than a strategy of its own.
Asking a depot for "the ten newest parcels across these twenty-three routes" is not the same request as asking for "the ten newest on each route", and only the second is what a nested connection means.
saying these in an interview costs you the question
- Applies one limit across all parents' children
- Adds a safety cap to an over-fetching lookup
- Reports hasNextPage false from an empty slice
- Validates nested paging only on uniform seed data
- Runs a separate count query per parent
- Calls per-parent slicing an optimisation, not correctness