skip to content

Users paging a GraphQL list field by offset report seeing one row twice and missing another. What causes it?

level: seniorimportance: should knowfreq 46%

answer

  1. A number, not a landmark
  2. Two requests, two independent orderings
  3. Something moved — or nothing did
  4. Ties have no promised relative order
  5. A tie-breaker fixes only half of it

basics

~20 s

Offset names a position, not a row. Between the two requests the ordering shifted — an insert before the window, or tied rows swapping places — so one row landed in both pages and another landed in neither.

solid answer

~50 s

An offset is a position in an ordering the server recomputes on every request, so two page requests only tile the set if that ordering is identical both times. Two independent things break it. **Concurrent writes**: on a livestock pedigree graph, six calves registered while the user reads `animals(offset: 4318, limit: 37)` push everything six positions later, so the next page repeats six animals; a deletion produces a silent gap instead. **A non-total order**: sorting by `birthDate` alone leaves large tie groups, and nothing constrains the relative order of tied rows between two executions — so pages can disagree with no writes at all. Separate them by re-running the sequence with writes frozen. The fixes differ: appending a unique tie-breaker to every sort kills the second cause, while only a row-anchored resume point, or a snapshotted paging session, kills the first.

code

graphql · 15 lines
graphql
query PedigreePage3 {
  animals(offset: 4318, limit: 37, orderBy: BIRTH_DATE_ASC) {
    id
    name
    birthDate
  }
}

query PedigreePage4 {
  animals(offset: 4355, limit: 37, orderBy: BIRTH_DATE_ASC) {
    id
    name
    birthDate
  }
}

go deeper

for a junior

Know that offset counts positions rather than naming rows, so a page can shift when data changes underneath. The first thing to ask about any paged list is what it is sorted by.

for a middle

Be ready to explain both mechanisms — the set changing and the ordering not being total — and why a total count or a bigger page size fixes neither.

for a senior

An interviewer expects a diagnosis: freeze the writes to separate the two causes, diff page ids against an unpaged read, then choose between a tie-breaker, a row-anchored resume and a snapshotted session — and say which guarantee you now offer.

for a principal

Own the contract across every paged field. Decide whether numbered pages are worth server-side session state and a TTL, keep one stability guarantee rather than a different one per team, and get it written into the schema's documented behaviour.

## What offset means, precisely `offset: 4318` means "run the query, order the results, throw away the first 4,318, return the next `limit`". Every request is independent. Nothing is carried from one page to the next except a number, and that number is a **position in an ordering that is recomputed from scratch each time**. If the ordering that request two computes is not the same ordering request one computed, the two windows do not tile the set — they overlap, or they leave a hole. There are two independent ways the ordering can differ, and a strong answer names both. Most candidates name only the first. ## Cause one: the set changed between the requests A user browsing a livestock pedigree graph is at `animals(offset: 4318, limit: 37)`. While they read that page, a calving-season registration batch inserts six calves that sort *before* position 4,318. Every row from there on shifts six positions later. The user clicks next, the client sends `offset: 4355`, and the first six animals of that page are six animals they have already seen. A deletion does the mirror image: rows shift earlier, and the rows that crossed the boundary are never rendered at all — a silent gap, which is the worse of the two because nobody reports it. Write rate decides whether you ever see this. A reference table written once a week never drifts; a live registration feed drifts on nearly every page turn. ## Cause two: the ordering was never total This is the one that surprises people, because it needs **no writes at all**. Sorting by `birthDate` alone does not define an order — it defines a partial order with tie groups. A calving-season intake can register several hundred calves on the same date, and within that tie group nothing constrains the relative order of rows. A data store is free to return tied rows differently between two executions: a different plan, a parallel scan, a different replica, a different cache state. Consecutive offset windows then cut the tie group at a different place each time, and the same animal lands in both pages while another lands in neither. Neither GraphQL nor the cursor connections convention says a word about ordering guarantees. Whatever stability you have comes entirely from your data source and the sort you actually wrote. ## Telling the two apart * **Freeze the writes.** Re-run the page sequence against a pinned as-of read, a repeatable-read transaction, or a quiet window. If the duplicates persist, the ordering is not total. If they vanish, it is concurrent mutation. * **Diff the multiset.** Log the ids returned for every page of one session and compare against a single unpaged read of the same predicate. You will see exactly which ids are doubled and which are missing, which also tells you roughly where the boundary moved. * **Read the sort.** Is the sort key unique? Is there any tie-breaker appended? If not, you have cause two whether or not you also have cause one. ## Fixes, strongest first 1. **Make the ordering total.** Append a unique column — the primary key will do — to every sort: `ORDER BY birth_date ASC, id ASC`. This removes cause two outright, costs nothing conceptually, and is a prerequisite for any cursor scheme you might adopt later. 2. **Move the resume point from a position to a row.** "Give me the 37 animals that come after (birth date 2026-03-11, id 9147)" never counts anything, so an insert elsewhere in the set cannot shift it. This is the whole idea the connection convention encodes, and it fixes cause one — which a tie-breaker alone does not. 3. **Snapshot the paging session.** If numbered pages are a hard product requirement and offsets must stay, pin the session to a consistent read: a repeatable-read transaction held open, an as-of timestamp threaded through the arguments, or a materialised result set keyed by a session token with a TTL. This buys correctness with server state and an expiry policy, so decide deliberately. 4. **Make the client tolerant.** De-duplicate by a stable id as pages arrive. This hides repeats; it does nothing about gaps, so treat it as a cosmetic patch on top of one of the above, never as the fix. Note what does *not* help: adding a `totalCount`, increasing the page size, or re-fetching page one. And under a tight latency budget — say a 340 ms p99 on the list field — option 3 is the expensive one, because it adds state and a second failure mode; option 2 usually costs nothing extra. ## What to say in the interview Name both mechanisms, say which fix addresses which, and be explicit that this is not a GraphQL defect: `offset` and `limit` are ordinary arguments, and the guarantee a client gets is whatever your resolver's ordering and consistency story actually provides. Then state the contract you intend to publish — "pages are stable within a session" or "page boundaries are approximate under concurrent writes" — because an undocumented guarantee is the thing that generated the bug report.

  • Does adding a unique tie-breaker to the sort stop rows repeating across pages?
    It stops the half caused by unstable tie ordering, which is the half that happens with no writes. It does nothing about the other half: an insert before the window still shifts every later position, so the boundary still moves. Only a row-anchored resume point, or pinning the paging session to a consistent snapshot, addresses that.
  • The product needs numbered pages, so offset has to stay. What do you offer?
    Pin the paging session to a consistent read — a repeatable-read transaction, an as-of timestamp threaded through the arguments, or a materialised result set keyed by a session token with a TTL — and have the client de-duplicate by id as a cosmetic backstop. Then document the guarantee on the field: "page boundaries are approximate under concurrent writes" is a better contract than silence.
  • How would you prove the duplicates are paging drift rather than a resolver bug?
    Log the ids returned for every page of one session and diff that multiset against a single unpaged read of the same predicate; doubled and missing ids show where the boundary moved. Then re-run the sequence against a pinned as-of read: if the duplicates survive frozen writes, the ordering is not total rather than the data being in motion.

saying these in an interview costs you the question

  • Says GraphQL's pagination is broken
  • Names only concurrent writes and misses non-unique sort keys
  • Assumes a sort on a non-unique column returns a stable order
  • Adds a total count and calls the problem solved
  • Believes the connections convention guarantees a stable ordering
  • Suggests re-fetching page one on every navigation

context