Why can computing hasNextPage make a GraphQL connection field slow, and how do you avoid it?
answer
- Two statements where one would do
- One query is bounded, the other is not
- Ask for one more than you need
- Drop the extra row before minting cursors
- Only the count grows with the table
basics
~20 sBecause the naive implementation answers it with a second query that counts or scans every row matching the filter, which grows with the data while the page query stays bounded. Fetch one row more than the page size instead, and drop it before building edges.
solid answer
~50 sThe page-info booleans are the only part of a connection response that describes data the client did not ask for, so they are where a paged field quietly acquires a second, unbounded query. Counting all matching rows to decide whether at least one more exists is invisible on a small dataset and dominates the request on a large one — the page read is capped by the page size, the count is capped by nothing. The standard fix is **limit plus one**: ask the store for one row more than the page size, using the same filter and the same total order. If the extra row arrives, `hasNextPage` is true. Then **drop that row before building edges and cursors** — if `endCursor` is taken from a row the client never received, every following page starts one object too late. Backward paging is the mirror, fetching one more than the backward page size to set `hasPreviousPage`.
code
pseudocode · 21 linesfunction artworksPage(filter, first, afterCursor):
rows = store.query(
where = filter,
after = afterCursor,
orderBy = totalOrder,
limit = first + 1)
hasNext = rows.length > first
if hasNext:
rows = rows.slice(0, first)
edges = rows.map(r -> { node: r, cursor: cursorFor(r) })
return {
edges: edges,
pageInfo: {
hasNextPage: hasNext,
startCursor: edges.isEmpty() ? null : edges.first().cursor,
endCursor: edges.isEmpty() ? null : edges.last().cursor
}
}go deeper
Know that a connection's hasNextPage has to be computed from something, and that asking the store for one row more than the page needs is the usual way to get it. Tuning this is not expected of you yet.
Walk the limit-plus-one approach end to end, including dropping the extra row before edges and cursors are built, and the mirrored version for backward paging.
Diagnose it rather than just fix it: name the per-field timing and statement-level evidence that separates the bounded page read from the unbounded count, and explain why a small lower-environment dataset hides the difference completely.
Own the guardrail, not the patch. A maximum page size, a review rule that no connection resolver issues an unbounded aggregate, and a lower environment whose data volume and selectivity resemble production are what stop the next team from rediscovering this the same way.
## The failure, concretely A 4-person platform team runs a museum collection graph. Two connection fields matter: a gallery's objects, paged 24 at a time, and a catalogue-wide search connection over the same objects. Staging carries a seeded slice of 9,412 objects; production holds 1,284,663 catalogued objects across 214 galleries. The connection resolver does two things per request. It runs a keyset page read — everything after the client's cursor, in the total order, bounded by the page size — and then, to fill `hasNextPage`, it runs a second statement with the same predicate that counts the matches. In staging both return inside 40 ms and nobody notices. In production the search connection's p99 climbs to 8.3 seconds and the field begins timing out against the request budget, but only for broad filters: "oil on canvas, acquired before 1965" matches 214,880 rows, and the count walks all of them. The page read is flat at 25 rows whatever the filter. **The bounded query was never the problem; the unbounded one was, and only production had enough rows to reveal it.** ## Why a count is the wrong shape for this question `hasNextPage` asks *is there at least one more element after this position*. A count answers *how many elements match in total*, which is strictly more information and strictly more work — a global aggregate paying for a local yes-or-no. The cost of the correct question does not grow with the collection; the cost of the count does. ## Limit plus one Ask for one row more than you intend to return. If it comes back, something follows the page. ```pseudocode function artworksPage(filter, first, afterCursor): rows = store.query( where = filter, after = afterCursor, orderBy = totalOrder, # same order the cursors were minted in limit = first + 1) # one more than the client asked for hasNext = rows.length > first if hasNext: rows = rows.slice(0, first) # DROP the probe row before anything else edges = rows.map(r -> { node: r, cursor: cursorFor(r) }) return { edges: edges, pageInfo: { hasNextPage: hasNext, hasPreviousPage: afterCursor == null ? false : store.existsBefore(filter, afterCursor), startCursor: edges.isEmpty() ? null : edges.first().cursor, endCursor: edges.isEmpty() ? null : edges.last().cursor } } ``` One extra row per page is a rounding error beside the page itself. A count is not. ## Three bugs this trick invites 1. **Not trimming at all.** The client asked for 24 and receives 25 edges. Any page-size cap the schema advertises is now off by one, and a client that renders fixed-size pages shows a stray item. 2. **Trimming the edges but taking `endCursor` from the untrimmed rows.** This is the expensive one because it is silent. The client receives 24 objects, sends back a cursor pointing at the 25th, and its next page begins after an object it never saw. One object is lost per page; crawling 1,284,663 objects in pages of 24 means roughly 53,527 pages and roughly that many missing objects, discovered weeks later by whoever reconciles the export. 3. **Fetching the probe row differently.** If the extra row is read with a relaxed filter or a different sort than the page, `hasNextPage` can be true when the genuine next page would come back empty, and the client makes a pointless round trip that returns nothing. ## Why staging never caught it Not a code difference — a data-shape difference. On 9,412 rows every plan is a cheap scan and every count returns instantly, so the two statements look identical in cost. Two signals separate them once you look at production: **per-field timing** that attributes latency to the connection field rather than to the operation as a whole, and the storage layer's statement log showing **two statements per connection field where the design called for one**. The durable fix for the blind spot is a lower environment whose row counts and selectivity resemble production, or at minimum running the count's plan against production statistics before shipping it. ## When you cannot over-fetch Sometimes the source returns exactly what you asked for and no more — a remote service that pages one way with its own opaque token. Note carefully what the convention allows here. The best-effort escape hatch covers only the boolean in the direction you are **not** travelling. On a forward page, `hasNextPage` is the exact one, so "I could not tell, so I said false" is not compliant. You need the source's own more-results flag or next-page token, or you page one item deeper than you display. Guessing false there is how a client stops crawling at an arbitrary point and reports a partial catalogue as complete. ## A note on scope If the schema also exposes a count of the whole result set, that is a separate field with its own separate cost and its own separate decision. It is not what fills these two booleans, and reaching for it to answer `hasNextPage` is what created the incident in the first place.
- What exactly breaks if endCursor is taken before the extra row is dropped?Every subsequent page starts one object too late. The client receives the page size it asked for but a cursor pointing at the row beyond it, so continuing from that cursor skips the object that should have led the next page. It loses exactly one object per page, with no error and no gap the client can detect, which is why it usually surfaces in a reconciliation rather than in monitoring.
- What is the equivalent when a client pages backward?The same trick mirrored: read one row more than the backward page size in the reverse order, and if the extra row arrives, `hasPreviousPage` is true. Drop it before building edges, and remember the results still have to be returned in the collection's forward order — the extra row is trimmed from the far end, which is the end the client would otherwise treat as the page boundary.
- The upstream source cannot over-fetch and you are paging forward. What are your options?Use whatever more-results signal the source provides — many return a next-page token or an explicit flag, and either answers the question directly. Failing that, read one item deeper than you display and hold it back. What you cannot do is return false because you did not check: on a forward page that boolean is the exact one, and a wrong false makes clients stop crawling early and treat a partial result as the whole collection.
You do not count everyone in the queue to find out whether someone is behind you. You glance at the next spot — the limit-plus-one row is that glance.
saying these in an interview costs you the question
- Filling hasNextPage with a count over all matching rows
- Saying it was fast in staging so the query is fine
- Fetching the whole result set and slicing it in memory
- Setting hasNextPage true whenever the page came back full
- Returning false for hasNextPage when the count is too slow
- Reading the probe row with a looser filter or a different sort