skip to content

A GraphQL connection is sorted only by a non-unique timestamp. What breaks when a client pages through rows that tie on it?

level: middleimportance: must knowfreq 58%

answer

  1. The resume is a comparison, not a count
  2. Two predicates, both wrong
  3. One skips the block, one repeats it
  4. Compare the pair, not the key
  5. Direction flips both halves

basics

~20 s

Rows that tie on the timestamp have no defined order among themselves, so a cursor inside that block names no single position. Depending on the resume comparison, the client sees the tied rows twice or misses them entirely.

solid answer

~50 s

The server resumes by turning the cursor's values into a comparison. With only a timestamp to compare, both possible comparisons are wrong. **Strictly greater** drops every row still tied at the cursor's timestamp, so the client silently loses them. **Greater-or-equal** re-selects the whole tied block on every request, so rows repeat and, if the block is bigger than the page size, the walk never advances — the accumulated list grows without a bound. Worse, the sequence *within* the tied block is not fixed, so consecutive executions can arrange those rows differently. The fix is a total order and a lexicographic resume: sort by `(scheduledAt, id)` and select rows where `scheduledAt > cursor.scheduledAt OR (scheduledAt = cursor.scheduledAt AND id > cursor.id)`. Both comparison halves must follow the sort direction — for a descending sort, both become `<`.

code

pseudocode · 11 lines
pseudocode
# A: strictly greater -> the 8,350 rows still tied at the cursor's
# timestamp are skipped and never delivered
fetch(where: scheduledAt > cursor.scheduledAt,
      orderBy: [scheduledAt ASC],
      limit: 50)

# B: greater-or-equal -> the whole tied block is re-selected every
# request; hasNextPage stays true and the walk never advances
fetch(where: scheduledAt >= cursor.scheduledAt,
      orderBy: [scheduledAt ASC],
      limit: 50)

go deeper

for a junior

Recall that a cursor is resumed by comparing values, not by counting rows, and that equal sort-key values leave the server no way to say which rows come next. Being able to name the duplicate-or-skip symptom is the bar here.

for a middle

Write both broken predicates and the lexicographic fix on a whiteboard, and say which failure each one produces. Be precise that the comparison direction must match the sort direction on the tie-breaker half as well.

for a senior

Diagnose from symptoms: duplicate ids across pages, a walk that never sets hasNextPage false, rows that vanish only at certain page sizes. Explain why bulk imports and date-only sources create the large tied blocks that expose a defect present since day one.

for a principal

Make it a class of bug the organisation cannot ship: a shared paging primitive that builds the comparison from the declared order, plus a contract test that seeds more ties than the page size. Argue why per-team hand-rolled resume logic is where this keeps coming back.

## The setup A clinic's appointment connection is ordered by `scheduledAt` ascending. A client walks it with `first: 50, after: <cursor>`. Server-side, the cursor is decoded back into the ordering values it carries and turned into a fetch that says, in effect, "rows that sort after this one". If the only ordering value is a timestamp, and timestamps repeat, that fetch cannot be written correctly. There are two candidate predicates and both are defective. ## Predicate A: strictly greater — the client silently loses rows ``` where scheduledAt > cursor.scheduledAt order by scheduledAt asc limit 50 ``` Suppose the last row of the previous page had `scheduledAt = 2026-03-02T09:00:00Z`, and 8,400 rows share that exact value because a migration stamped a whole recall batch with one timestamp. The previous page delivered 50 of them. This predicate now excludes *every* row at that timestamp — including the 8,350 not yet delivered. The client jumps straight to 09:05 and never learns that 8,350 appointments exist. Nothing errors; the page is well-formed and `hasNextPage` is honest about the rows the predicate can still see. ## Predicate B: greater-or-equal — the list grows without a bound ``` where scheduledAt >= cursor.scheduledAt order by scheduledAt asc limit 50 ``` Now the tied block is never excluded. Every request re-selects rows from the same 8,400 and the storage layer, having no further ordering instruction, is free to return them in a different arrangement each time. `hasNextPage` remains true forever, the client keeps asking, and the accumulated result list grows without a bound — the same appointment id appearing dozens of times — until the process runs out of heap or an operator kills it. This is the classic production incident behind this question: an unbounded paging loop that is actually a tie-defect, not a client bug. ## Why "just order by scheduledAt" is not enough even without the predicate problem Even ignoring the resume predicate, the arrangement *within* the tied block is not defined. Two executions of the same fetch can return the 8,400 rows in different sequences — different plans, a parallel scan, a different replica, a different cache state. So even a server that somehow tracked an offset into the block would be indexing into a sequence that is not the same sequence next time. ## The fix: compare the pair, not the key Make the order total by appending a unique, immutable column — the row id — and resume with a **lexicographic** comparison over the pair: ``` where scheduledAt > cursor.scheduledAt or (scheduledAt = cursor.scheduledAt and id > cursor.id) order by scheduledAt asc, id asc limit 50 ``` Read it as: strictly later timestamps always qualify; at the *same* timestamp, only rows whose id sorts after the cursor's id qualify. That predicate has exactly one correct answer for every row, so the 8,400 tied rows are traversed in a fixed internal sequence, 50 at a time, each row exactly once, and the walk terminates. The cursor therefore has to carry both ordering values, not just the visible sort key — the server cannot reconstruct the second half of the comparison from a timestamp alone. Note also that the comparison direction has to track the sort direction: for a descending sort both halves become `<`, including the tie-breaker half. Mixing them (descending timestamp, ascending id) is a real and easy bug, and it produces exactly the same duplicate-and-skip symptoms inside every tied block. ## Detecting it The symptoms are recognisable and worth being able to name in an interview: - **Duplicate ids across pages.** Collect ids while walking and assert the set size equals the row count. Repeats mean predicate B. - **A page walk that never ends,** or that returns far more rows than the collection contains. - **Missing rows that reappear when the page size changes.** With predicate A, whether a row is skipped depends on where the page boundary happens to land inside the tied block, so changing `first` from 50 to 100 changes which rows vanish. Anything whose correctness depends on page size is a tie defect. - **A defect that only shows up in production.** Test fixtures usually have distinct timestamps; bulk imports, batch jobs, coarse date-only sources and low-resolution clocks are what create large tied blocks. ## Client-side workarounds, and why they are not the fix Deduplicating by id on the client hides predicate B's repeats but not predicate A's omissions, and it does nothing about the unbounded loop — the backend is still being asked for pages forever. Capping the number of pages a client will fetch is a sensible seatbelt, not a fix. The defect is in the ordering, so the fix belongs in the ordering: every sort the field supports ends with a unique column, and every resume comparison is lexicographic over the whole tuple.

  • How would you detect this from the client side, in production?
    Collect node ids while walking a connection and assert the set size matches the number of rows received; repeats mean the greater-or-equal defect. Also alarm on walks that exceed a sane page count, and compare results across two page sizes — a row that disappears when `first` changes from 50 to 100 is a tie defect, because correctness is depending on where the page boundary lands.
  • Can the client just deduplicate by id and move on?
    No. Deduplication hides the repeated rows but not the skipped ones, and it does nothing about the unbounded walk — the backend is still serving the same tied block forever. A page-count cap is a reasonable seatbelt against runaway loops, but the defect lives in the server's ordering and has to be fixed there.
  • Why do these defects almost never show up in tests?
    Fixtures usually give every row a distinct timestamp, so no tied block exists to page across. Ties are created by bulk imports, batch jobs, date-only source data and coarse clocks — all production phenomena. A useful regression test seeds more rows sharing one sort-key value than the page size and asserts the walk returns each id exactly once.

saying these in an interview costs you the question

  • Says the server tracks an offset into the tied block
  • Uses greater-or-equal and deduplicates on the client
  • Assumes tied rows come back in a stable sequence
  • Flips the timestamp comparison but not the tie-breaker's
  • Blames the client for an unbounded paging loop
  • Claims a larger page size makes the problem go away

context