How do you configure pagination in an Airbyte low-code connector manifest?
answer
- It belongs to the retriever, next to the requester
- Three strategies: cursor, offset, page number
- A RequestOption says where the token is injected
- Say explicitly when the API has run out
- The default hazard is looping forever
basics
~20 sAttach a DefaultPaginator to the retriever, pick a strategy — CursorPagination, OffsetIncrement or PageIncrement — and use page_token_option and page_size_option to say where the token and size are injected into the request. Always set a stop condition.
solid answer
~50 sPagination lives on the retriever as a `DefaultPaginator` with three parts. The **`pagination_strategy`** decides what the next token *is*: `CursorPagination` reads it out of the response with an interpolation such as `{{ response.next_cursor }}` or a link header, `OffsetIncrement` adds the page size to a running offset, and `PageIncrement` counts pages. The **`page_token_option`** is a `RequestOption` saying where that token goes — `inject_into` a `request_parameter`, a `header`, `body_json`, or the `path` when the API hands back a full next URL. The **`page_size_option`** does the same for the page size. The part candidates forget is `stop_condition`: with a cursor strategy you must tell the paginator when the API has run out, e.g. `{{ not response.has_more }}`. Without it, an API that echoes the last cursor or returns an empty page forever will loop until the sync is killed.
code
yaml · 17 linesretriever:
type: SimpleRetriever
paginator:
type: DefaultPaginator
page_size_option:
type: RequestOption
inject_into: request_parameter
field_name: per_page
pagination_strategy:
type: CursorPagination
page_size: 100
cursor_value: "{{ response.next_cursor }}"
stop_condition: "{{ not response.get('has_more', False) }}"
page_token_option:
type: RequestOption
inject_into: request_parameter
field_name: cursorgo deeper
Know that a declarative Airbyte stream needs a paginator, and that offset, page-number and cursor styles exist. Being able to point at the right one for a given API is enough here.
Explain the paginator's three parts — strategy, token injection, size injection — and why a cursor strategy needs an explicit stop condition while an offset strategy usually ends on a short page.
Diagnose from symptoms: a sync that never finishes, one that returns page one repeatedly, or one that quietly drops rows under concurrent writes. Tie each to a specific misconfiguration and its fix.
Frame paging as a cost and correctness decision across a connector fleet: request volume against vendor rate limits, page size versus checkpoint granularity, and preferring server-managed cursors over offsets in anything you standardise on.
## Where the paginator sits In a declarative Airbyte stream, the retriever is composed of a requester (how to make one call), a record selector (where records are in the response) and a paginator (how to make the *next* call). The paginator is the component that turns one endpoint into a stream: the framework calls the requester, extracts records, asks the paginator for a next token, and repeats until the paginator says stop or a page yields nothing. Because the loop belongs to the framework, all you supply is a description. ## The three strategies **`CursorPagination`** is for APIs that hand you an opaque token or a next URL. Its `cursor_value` is an interpolation over the response — typically `{{ response.next_cursor }}`, `{{ response['meta']['next'] }}`, or a value pulled from a `Link` header. This is the strategy to use whenever the API's own answer tells you where to go next, which is also the safest kind of paging because the server manages the position. **`OffsetIncrement`** is for `?limit=100&offset=200` APIs. You give it a `page_size` and it adds that to the offset each round. It is stateless on the server, which means it is vulnerable to rows shifting underneath you: if rows are inserted or deleted in the source while you page, offset paging can skip or duplicate records. Mitigate by requesting a stable sort order in `request_parameters` — ordering by an immutable key rather than by an updated timestamp. **`PageIncrement`** is for `?page=1&per_page=100` APIs and has the same skew hazard as offsets, since a page number is just an offset in disguise. ## Injecting the token A strategy produces a value; `page_token_option` says where to put it. It is a `RequestOption` with an `inject_into` of `request_parameter`, `header`, `body_json` or `body_data`, plus the `field_name` to use. There is one special case worth knowing: when the API returns a complete next URL, you inject into the `path`, and the framework requests that URL directly instead of appending a parameter to the base path. `page_size_option` is the same mechanism for the size, since most APIs want the size on every request, not just after the first. Inside the requester you can also refer to `{{ next_page_token }}` in interpolations when you need the value somewhere unusual, but the request options are the idiomatic route. ## Stop conditions are the load-bearing part The framework stops paging when the paginator returns no token, when a page returns zero records, or when `stop_condition` evaluates truthy. With offset and page strategies, a short or empty page naturally ends the loop. With cursor strategies, you cannot rely on that: many APIs return the *same* cursor on the last page, or a null buried in a structure your interpolation renders as a string. Both cases produce an infinite loop that re-reads the last page until the sync times out or the destination fills with duplicates. So write the condition explicitly against whatever the API actually signals — `{{ not response.get('has_more', False) }}`, `{{ response.next_cursor == none }}`, `{{ not response['data'] }}`. ## Interaction with incremental reads and slicing Pagination composes with the incremental cursor and the partition router rather than replacing them. If the stream has a `DatetimeBasedCursor` with a `step`, the framework reads one time window at a time and pages *within* each window; state advances per window, not per page. If a `SubstreamPartitionRouter` fans out over parent records, each partition is paged independently. Understanding the nesting matters when you diagnose a slow sync: pages inside windows inside partitions multiply, and a page size of 25 on a five-year backfill with daily windows is a very large number of requests. ## Failure modes to be able to name - **Infinite loop** from a missing or wrong `stop_condition` — the single most common declarative-pagination bug. - **Wrong injection target**: the token goes in as a query parameter when the API expects a header, so every page is page one and the sync returns the same rows forever, or terminates after one page. - **Records extracted from the wrong key**, so pages come back but the record selector emits nothing and the loop ends immediately on an "empty" page. - **Skew under offset paging**, producing missing or duplicated rows on a busy table. - **Page size larger than the API's cap**, silently clamped, which does not break correctness but makes your slice-count arithmetic wrong. ## In the Python CDK The hand-written equivalent is a stream that implements `next_page_token(response)` — returning `None` to stop — and folds the token into `request_params` or `request_headers`. Same loop, same hazards; the declarative components exist so you stop rewriting it per connector.
- An API returns a complete next-page URL. How do you follow it declaratively?Use `CursorPagination` with a `cursor_value` reading the URL out of the response or the Link header, and set the `page_token_option` to inject into the `path`. The framework then requests that URL directly instead of appending a query parameter to the configured base path.
- Why can offset-based paging lose rows on a busy source table?An offset is a position in a result set the server recomputes for every request. If rows are inserted or deleted between pages, records shift across page boundaries and some are skipped or repeated. Requesting a stable sort on an immutable key reduces it; a server-managed cursor avoids it.
- How does pagination interact with a DatetimeBasedCursor that has a step configured?The cursor slices the read into time windows and pagination happens inside each window. State advances once a window completes, not once a page does. That nesting is why a small page size plus narrow windows can turn a backfill into an enormous number of requests.
saying these in an interview costs you the question
- Assuming the framework detects the last page automatically
- Injecting the page token as a parameter when it belongs in a header
- Confusing page size with the number of pages to fetch
- Believing offset paging is safe on a table being written to
- Thinking pagination replaces the incremental cursor