What does it mean for an HTTP request to be self-contained, and how would you redesign a paginated endpoint that currently relies on the server remembering which page each client last received?
answer
- Cold instance + this request = correct answer
- Position travels in the URL, not the server
- Opaque signed cursor over server-held cursor
- Offset: deep pages slow, drift under inserts
- ETag + If-Match = precondition in the request
basics
~20 sSelf-contained means the request carries everything needed to process it: credentials, target, and position. For pagination, put the position in the request - an offset or an opaque cursor the server returns to the client - instead of a server-held cursor.
solid answer
~50 sA self-contained request carries its full context: who is calling (a credential in the request), what is being acted on (the URL), how (method and headers), and any position or filter needed (query parameters or body). Nothing is implied by prior requests. For pagination, the server-remembered variant (`GET /search/next` advancing a per-client cursor) breaks on any instance change, cannot be retried safely, cannot be bookmarked or shared, and produces wrong results if two tabs interleave. The fix is to return the position to the client and let it send it back. Either offset-based (`?page=3&size=20`) or, better for large or changing datasets, an opaque cursor encoding the last key seen, returned in the response as a `next` link. The server decodes it per request and holds nothing. As a bonus the request becomes cacheable, retryable and shareable, since the URL alone determines the response.
code
http · 9 linesGET /items?limit=2&after=eyJpZCI6MTAyfQ HTTP/1.1
Host: api.example.com
Authorization: Bearer <token>
HTTP/1.1 200 OK
Content-Type: application/json
Link: </items?limit=2&after=eyJpZCI6MTA0fQ>; rel="next"
{"items":[{"id":103},{"id":104}],"next":"eyJpZCI6MTA0fQ"}go deeper
Say the request must carry credentials, target and page position, and show the fix as adding page or cursor parameters to the URL.
Contrast offset with keyset cursors on performance and stability, and explain why the cursor should be opaque and returned by the server.
Cover retry semantics, concurrent clients, tampering and cursor expiry, and extend the principle to ETag/If-Match preconditions and job resources.
Discuss it as a contract decision - what clients may construct versus only echo - and its effect on cacheability, evolvability of the query engine, and multi-tenant isolation.
## What self-contained includes A request is self-contained when a server that has never seen this client can process it correctly using only the request plus shared storage. Concretely it must carry: - **Identity and authorisation** - a credential in the request (an `Authorization` header or a cookie), not an implicit "this connection is logged in". - **Target** - the resource, in the URL. - **Intent** - the HTTP method and any content-negotiation headers. - **Parameters** - filters, sort order, page size, and crucially **position**. - **Preconditions** - anything the server must check, such as a validator sent in an `If-Match` header, rather than "the version I fetched earlier". The common failure is position and preconditions: teams remember to send credentials but let the server track where the client is. ## Why server-held cursors fail A design like `POST /search` (start) then repeated `GET /search/next`: - **Breaks on rerouting.** A different instance has no cursor, so sticky sessions become mandatory. - **Breaks on restart.** Deploy or crash mid-pagination and the client cannot continue. - **Is not retryable.** If `next` times out after the server advanced but before the response arrived, a retry skips a page, and the client cannot distinguish that from a lost request. - **Is not concurrent-safe.** Two tabs or two workers share one cursor and interleave results. - **Is not addressable.** You cannot bookmark, share, log or replay page 3; caches and proxies cannot help, because the response depends on invisible server state. - **Costs memory.** Open cursors or held result sets consume server resources proportional to concurrent clients, and need timeouts and eviction. ## The stateless redesign **Offset pagination.** `GET /items?offset=40&limit=20`. Simple, allows jumping to a page, and every request is fully described. Weaknesses: deep offsets get slow because the database must skip rows, and inserts or deletes between requests cause items to be skipped or repeated. **Cursor (keyset) pagination.** The response includes a `next` link containing an opaque token that encodes the sort key of the last returned row. The server decodes it into a predicate such as `WHERE (sort_key) > (value)`. Position lives in the token, so the server holds nothing; performance is stable at any depth because there is no skipping; and results are stable under concurrent inserts. Design notes for cursors: make the token **opaque** to the client (encoded, and signed if it must not be tampered with) so you can change its internals later; include the sort key and a fingerprint of the filters so a cursor cannot be replayed against a different query; give it an expiry if it references a snapshot; and return it as a link in the response body or a `Link` header so the client never constructs one itself. ## The same principle beyond pagination - **Concurrency control**: instead of remembering the version a client fetched, return a validator in an `ETag` and require the client to send it back in `If-Match` on write, so the precondition travels with the request. - **Multi-step flows**: instead of a server-held wizard, create a resource per flow and let the client hold its id. - **Bulk operations**: instead of a server-side job handle tied to one instance, create a job resource in shared storage and return its URL. In every case the shape is the same: whatever the server was tempted to remember becomes either a value the client echoes back or a row in shared storage with an address. ## Costs to acknowledge Self-contained requests are larger, repeat authentication work per call, and can leak information if you put sensitive data in a client-visible token - hence opaque and signed cursors. Those are real, and a good answer names them rather than pretending the pattern is free.
- When is offset pagination still the right choice over cursor pagination?When the dataset is small or slow-changing, when users genuinely need to jump to an arbitrary page number, and when a total count is part of the contract. Offsets are simpler and support random access; cursors win once the result set is large enough that deep offsets get slow, or volatile enough that shifting rows would skip or duplicate items.
- Why should a pagination cursor be opaque and signed?Opaque keeps its internals a private implementation detail, so you can change the sort key or encoding without breaking clients that treat it as a blob. Signing prevents a client from crafting a cursor that shifts the query - widening a filter, or paging into another tenant's rows - since the server would otherwise trust decoded values it never issued.
saying these in an interview costs you the question
- Keeping an open database cursor or result set per client on the server
- Exposing raw internal ids or filter fields inside a cursor and trusting them unvalidated
- Assuming offset pagination is stable while rows are being inserted or deleted
- Treating self-contained as meaning the whole entity must be sent on every request
- Putting the position in a custom header the client cannot bookmark rather than in the URL