Compare paginating a REST collection with `page`/`size` (or `offset`/`limit`) parameters against paginating with an opaque cursor token such as Stripe's `starting_after` or Google's `pageToken`. What does each give up, and when would you choose each?
answer
- offset = position; cursor = anchor on last seen row
- insert before your page → offset repeats/skips rows
- deep offset cost grows; cursor cost flat
- cursor: no page numbers, no totals, needs ending_before for back
- opaque + signed token = freedom to change implementation
basics
~20 sOffset paging is simple and supports jumping to any page, but gets slower the deeper you go and duplicates or skips rows when the collection changes between requests. Cursor paging anchors on the last item seen, so it stays cheap and stable, but only supports next/prev — no page numbers, no totals.
solid answer
~50 s**Offset** (`?page=3&size=50` or `?offset=100&limit=50`) is stateless, trivially understood, allows jumping straight to page 47, and supports "page 3 of 214" UI. Two costs: the server must skip N rows before returning any, so deep pages get progressively more expensive; and the window is defined by *position*, so an insert or delete before your position shifts everything — page 2 then repeats a row you already saw, or skips one you never will. **Cursor** (`?starting_after=o_123&limit=50` or `?page_token=...`) encodes a position in the sort order — "the rows after this anchor". Cost is independent of depth, and inserts before your anchor cannot shift your window, so no duplicates or skips for rows already passed. Give up: page numbers, jump-to-page, usually exact totals, and going backwards unless you explicitly support `ending_before`. Choose offset for small, bounded, mostly-static collections with page-number UI (admin tables). Choose cursors for large, high-write, or programmatically-iterated collections — feeds, event logs, export jobs, anything an SDK loops over. Cursor is the right default for a public API.
code
http · 11 linesGET /v1/orders?offset=100&limit=50&sort=-created_at HTTP/1.1
GET /v1/orders?limit=50&starting_after=o_9f3&sort=-created_at HTTP/1.1
HTTP/1.1 200 OK
Content-Type: application/json
{
"data": [ ... 50 orders ... ],
"meta": { "has_more": true, "next_cursor": "eyJjIjoiMjAyNC0wMS0wOFQxMDoxMiIsImlkIjoib185ZjMifQ" }
}go deeper
Describe both parameter styles and give the concrete duplicate/skip example that offset paging suffers when rows are inserted.
Cover deep-offset cost, cursor stability and exactly what it does and does not guarantee, and the loss of page numbers, totals and backwards paging.
Discuss opacity and signing, cursor-versus-sort/filter validation, capping max offset on legacy endpoints, and running both schemes during a migration.
Set cursors as the estate default for unbounded collections, weigh the product cost of losing page-number UI, and own the deprecation path for offset across many clients you cannot see.
## Offset pagination: the contract `GET /orders?page=3&size=50`, or equivalently `?offset=100&limit=50`. The window is defined by an absolute position in the ordered result: "skip 100, take 50". Nothing is remembered between requests, which is its main appeal — any page is reachable from a URL, links are bookmarkable, and page-number UI falls out for free. ### Cost that grows with depth To return rows 100 000–100 049 the server must first pass over 100 000 rows in the ordering. The work is proportional to the offset, so the last page of a large collection is dramatically more expensive than the first. Users rarely browse to page 4000, but crawlers, scripted exports and "loop until empty" clients do, and they turn a fast endpoint into a slow one at exactly the moment they are iterating hardest. The API-contract consequence — which is the part that belongs in this discussion — is that offset makes tail latency a function of a client-supplied number, so you end up capping the maximum offset, which is itself a contract wart. (The storage-engine reasons for the cost belong to database indexing, not the API contract.) ### Instability under concurrent writes This is the sharper problem. Suppose the list is newest-first and a client reads page 1 (items 1–50), then page 2 (items 51–100). If three new items are inserted in between, everything shifts down by three: items 48, 49, 50 — already shown on page 1 — reappear at the top of page 2. Deletions produce the mirror bug: rows slide up and a client never sees the ones that crossed the boundary while it was reading. Nothing is broken in the storage layer; the contract itself is defined by position in a set that is moving. For a human clicking through an admin table this is a curiosity. For a nightly export loop it is silent data loss, and it is invisible in testing because test data does not change under you. ## Cursor pagination: the contract `GET /orders?limit=50&starting_after=o_9f3` (Stripe) or `?page_size=50&page_token=CiAKGjBpNDd...` (Google AIP style). The client sends back an opaque token that the server issued; the server decodes it into an anchor position in the sort order and returns the rows *after* that anchor. ### Why it is stable and cheap Because the window is defined relative to a value the client has already seen rather than a count of rows skipped, inserting rows earlier in the ordering does not shift it: rows already returned stay behind the anchor. Depth costs nothing extra — fetching the page after anchor X is the same amount of work whether X is the 50th or the five-millionth row. Note the honest scope of the guarantee: cursors prevent *duplicates and skips caused by shifting positions*. They do not give you a snapshot. A row you already passed that is later modified or deleted will not be reflected, and a newly inserted row that sorts before your anchor will never appear in this iteration. That is usually exactly what an iterating client wants, but say it precisely rather than claiming "consistent pagination". ### What you give up - **Random access.** There is no "page 47". Reaching a distant point means walking there. This kills page-number UI, which is the real reason product teams resist cursors. - **Totals.** Cursor APIs usually expose `has_more` rather than a count, because the whole point was to avoid whole-set work. - **Backwards.** Needs an explicit reverse cursor (`ending_before`) or an explicit `previous_page_token`. - **Simplicity.** The token must encode the sort keys of the anchor, and usually the sort order and filters, so the server can detect a client mixing a cursor with different query parameters and reject it. ### Opacity as contract Cursors must be documented as opaque: base64 of an internal structure, ideally signed or otherwise tamper-evident. The reason is freedom — if clients cannot parse the token, you may change the encoding, add fields, or switch strategy without breaking anyone. If it is a readable `id=1234`, someone *will* construct one by hand, and it becomes contract. Stripe's `starting_after` is the notable exception — it is a plain object id — which is friendly but permanently ties the cursor to the id and the default sort. ## Choosing | | Offset | Cursor | |---|---|---| | Jump to arbitrary page | yes | no | | Cost at depth | grows | flat | | Stable under concurrent inserts | no | yes for passed rows | | Exact total natural | yes (at a price) | no | | Bookmarkable page URL | yes | token may expire | Offset is fine for admin screens over bounded, low-churn data where humans want page numbers. Cursors are the right default for public APIs, anything an SDK iterates, event/audit logs, feeds, and any collection that grows without bound. A common pragmatic hybrid: cursor-based iteration as the supported contract, plus a capped offset (`max_offset`) retained for a legacy UI, with the cap documented. Another: offset paging for the first N pages, cursors beyond — clever but two contracts, so justify it. ## Migration Because the parameters differ, you can add cursor support alongside offset without breaking clients: accept both, return a `next` link that uses a cursor, and let clients that follow links migrate for free. Then measure who still sends `offset`, deprecate, and remove. Being able to describe that path — rather than "we'd switch to cursors" — is what a senior answer sounds like.
- Does cursor pagination give the client a consistent snapshot of the collection?No. It guarantees that rows already passed will not be re-served and that rows after the anchor are not skipped because of shifting positions. Rows inserted before the anchor after you passed it will never appear in that iteration, and rows you already returned may since have changed or been deleted. If a true point-in-time view is required, that needs a snapshot or an export job, not pagination.
- How do you add cursor pagination to an API that already ships offset pagination?Accept both: keep `offset`/`limit` working, add a cursor parameter, and make every response's `next` link a cursor URL. Clients that follow links migrate with no code change. Then measure which API keys still send `offset`, cap the maximum offset to bound the damage, publish a deprecation notice with a date, and finally reject the parameter with a 400 pointing at the cursor.
Offset paging is "start reading at line 500" in a book someone is still inserting pages into; cursor paging is a bookmark placed after the last line you actually read.
saying these in an interview costs you the question
- Claiming cursors give a consistent snapshot of the whole collection
- Saying offset pagination is 'just slower' without mentioning duplicated and skipped rows under concurrent writes
- Exposing a cursor that is plainly a decodable offset or row id, then treating it as opaque in the docs
- Assuming cursors can support jump-to-page or an exact total 'if you add a count'
- Applying a cursor issued under one sort or filter to a request with different sort or filter parameters