skip to content

Why is an unpaginated LIST expensive for the Kubernetes API server, and how do the limit and continue parameters let a client page through results?

level: middleimportance: should knowfreq 44%

answer

  1. cost grows with object count
  2. seats estimated before dispatch
  3. limit, then follow the token
  4. short page is not the end
  5. compaction expires the token: 410

basics

~20 s

A full LIST makes kube-apiserver load, convert and serialize every matching object in one response, and Priority and Fairness charges it several seats. With limit, the server returns pages plus a continue token for the next request until the token comes back empty.

solid answer

~50 s

A LIST without `limit` makes kube-apiserver read, filter, convert and serialize the whole collection and keep the result in memory until it has been sent. The cost grows with the number and size of objects. API Priority and Fairness estimates that work in advance and charges the LIST several seats, so a few big LISTs can fill a priority level. Paging splits the work: `?limit=500` returns at most 500 items plus `metadata.continue`, and the client repeats the request with `continue=<token>` until the token is empty. All the pages together form one consistent snapshot. A page can hold fewer items than `limit`, or none, and still carry a token, so clients must follow the token rather than stop on a short page. If etcd has compacted that snapshot's revision, the server returns `410 Gone`, and the client starts the list over. `kubectl get` pages by default with `--chunk-size=500`.

code

bash · 2 lines
bash
TOKEN=$(kubectl get --raw '/apis/batch/v1/namespaces/doc-ocr/jobs?limit=250' | jq -r '.metadata.continue')
kubectl get --raw "/apis/batch/v1/namespaces/doc-ocr/jobs?limit=250&continue=${TOKEN}" | jq '.items | length'

go deeper

for a junior

Remember that big lists should be fetched in pages: limit sets the page size, and continue carries the token for the next page.

for a middle

Explain what the server does for a full LIST, how APF estimates its seats, and the short-page and 410 Gone rules of the continue token.

for a senior

Spot unpaginated LISTs as the cause of API server memory spikes and 429s for neighbours, and push periodic sweeps to page while controllers use a watch cache.

for a principal

Treat large collections as a platform risk: set retention for completed objects, publish client guidelines on paging, and budget APF seats for bulk readers.

## What a LIST costs A **LIST** request, such as `GET /apis/batch/v1/jobs`, asks kube-apiserver for every matching object in a single response. To answer it, the server has to: - read the objects, either from its **watch cache** or from etcd, depending on the request's options; - decode them, filter them by label or field selector, and convert them to the requested API version; - serialize the whole result and keep it in memory until it has been written to the client. That work grows with the number and size of the objects, not with the number of requests. One unpaginated LIST over a large collection can allocate a lot of memory and run for seconds. Several such LISTs running at once are a classic cause of kube-apiserver memory spikes. ## How API Priority and Fairness charges for it **API Priority and Fairness (APF)** limits how many requests run at once in each priority level. It counts this capacity in **seats**. A GET or a small write takes one seat. For a LIST, the server **estimates** the work before running it. The estimate uses how many objects of that resource exist, their average size, whether a selector is set, whether a `limit` is set, and whether the cache can serve the request. The LIST is then charged several seats, up to a cap for each level. So one expensive LIST can tie up as much capacity as many ordinary requests, and everything behind it in that priority level waits longer or gets `429 Too Many Requests`. Asking for smaller pages lowers the estimate for each request. ## Paging with limit and continue The list API supports paging directly: 1. The client sends `?limit=500`. The server returns at most 500 items. If more remain, it sets `metadata.continue` to an opaque token, and it may also set `metadata.remainingItemCount` to an estimate of how many are left. 2. The client sends the same request again with `&continue=<token>`, and repeats until a response comes back with an empty `continue`. 3. All the pages together form a **consistent snapshot** at the resourceVersion of the first page. The list does not change partway through. Some rules catch people out: - The server may return **fewer than `limit` items, even zero, and still set `continue`**. This happens especially when a selector filters out most objects. A client must keep going until `continue` is empty, and must never stop just because a page is short. - A token is valid only while etcd still keeps that old revision. After compaction, the server answers **410 Gone** with "The provided continue parameter is too old to display a consistent list result". The client then has to start the list again from the first page. Some versions of this error also offer a token for continuing without the consistency guarantee. - Send the same selectors and namespace on every page as on the first one. The token only makes sense for the same query. ## Tooling that already pages | Client | How it pages | |---|---| | `kubectl get` | `--chunk-size`, default 500; `--chunk-size=0` turns paging off | | client-go | the `k8s.io/client-go/tools/pager` package follows `continue` tokens for you | | raw HTTP | `kubectl get --raw` with `limit` and `continue` query parameters | ## Pagination is not the whole answer Paging makes each request cheaper. It does not make repeated full LISTs cheap. A controller that pages through 38,640 Jobs on every reconcile still reads all 38,640 objects each time. The usual fix for that is a list-then-watch cache, which is a separate mechanism. Paging is the right tool for one-off or periodic sweeps, such as audits, exports and cleanup scripts, and for the initial list itself. ## A worked example A document-OCR pipeline runs on a 140-node cluster shared by 22 teams. It creates one Job per batch of pages, and completed Jobs have piled up to 38,640 objects. Its nightly report script ran `GET /apis/batch/v1/namespaces/doc-ocr/jobs` with no `limit`. The response took several seconds, the script's pod ran close to its 2.6 GiB memory limit, and other ServiceAccounts in the same priority level got 429s during the run. With `limit=250`, the same read becomes 155 smaller requests (38,640 / 250 = 154.56, rounded up to 155 pages). Each one is estimated at a fraction of the original seat cost, and the script only ever holds one page in memory.

  • A paged LIST returns an empty items array but a non-empty continue token. Is the list finished?
    No. The server can return short or even empty pages when a label or field selector filters out the objects it scanned for that page. Only an empty `metadata.continue` means the list is complete. A client that stops on an empty page silently misses objects further on.
  • What should a client do when a continue request fails with 410 Gone?
    The snapshot behind the token has been compacted away in etcd, so the server can no longer continue that consistent list. The safe choice is to start the list again without `continue`. For sweeps that can tolerate it, the error response may offer a token to continue without the consistency guarantee. Objects changed in the meantime can then appear or be missing.

saying these in an interview costs you the question

  • Pagination is unnecessary because the API server serves every LIST from cache for free.
  • A page with fewer items than limit means the list is complete.
  • A continue token stays valid forever, so a paused export can resume days later.
  • Adding a label selector makes a LIST cheap because the server only reads matching objects.
  • Paging a LIST on every reconcile loop fixes a controller's API load.