When should you build an Airbyte connector with the low-code YAML manifest instead of the Python CDK?
answer
- Describe the API vs program the API
- JSON REST with normal auth and paging
- Where a path extractor stops being enough
- Custom components let you mix the two
- Fleet maintenance, not just first build
basics
~20 sUse the low-code declarative manifest when the source is a plain REST API with JSON responses, standard auth and standard pagination. Drop to the Python CDK when the source is not HTTP, the response needs real parsing, or the request logic is stateful.
solid answer
~50 sThe declarative manifest is a `manifest.yaml` assembled from named components — an `HttpRequester` for the call, a `RecordSelector` with a `DpathExtractor` to pull records out of the response body, a `DefaultPaginator`, an authenticator, and a `DatetimeBasedCursor` for incremental reads. If the API fits that shape, YAML is strictly better: less code, no test harness to maintain, and the Connector Builder gives you a live test-read while you edit. You go to Python when the source is a database, a file store or an SDK-only API; when responses need decoding rather than path extraction (XML, CSV inside a zip, deeply reshaped payloads); when requests depend on prior responses in ways the components don't express; or when signing is custom. It is not either/or — a manifest can reference custom Python components by `class_name`, so you keep the declarative skeleton and write Python only where it hurts.
code
yaml · 19 linesstreams:
- type: DeclarativeStream
name: users
primary_key: ["id"]
retriever:
type: SimpleRetriever
requester:
type: HttpRequester
url_base: "https://api.example.com/v1"
path: "/users"
http_method: GET
authenticator:
type: BearerAuthenticator
api_token: "{{ config['api_key'] }}"
record_selector:
type: RecordSelector
extractor:
type: DpathExtractor
field_path: ["data"]go deeper
Know that Airbyte lets you build a connector by describing a REST API in YAML instead of writing Python, and that the Connector Builder is the UI for doing it.
Be able to name the pieces of a declarative stream — requester, record selector with a path extractor, paginator, authenticator, cursor — and give concrete cases where each stops being sufficient.
Show you have hit the boundary in practice: async report APIs, non-JSON payloads, custom signing, and how you kept the manifest by plugging in a custom component instead of rewriting the connector.
Own the fleet argument: standardising on declarative connectors lowers upgrade cost and spreads ownership beyond Python developers, at the price of debuggability and component coverage. Say when you would accept that trade and when you would not.
## Two ways to write the same connector Airbyte's CDK offers two authoring styles that both produce a connector satisfying the same four-command protocol. The **low-code declarative** style is a `manifest.yaml` file interpreted at runtime by a declarative source: you describe *what* the API looks like and the framework does the calling, paging, retrying, extracting and checkpointing. The **Python CDK** style is code: you subclass stream classes and implement methods. The Connector Builder in the Airbyte UI is a front end for the first one — it edits a manifest and gives you a test-read panel that shows the request it sent, the raw response and the records extracted, which shortens the loop enormously compared with rebuilding a Docker image. ## What the manifest covers well A declarative stream is composed of a retriever holding three parts. The **requester** owns `url_base`, `path`, `http_method`, `request_parameters` and the authenticator. The **record selector** owns the extractor, normally a `DpathExtractor` with a `field_path` like `["data", "items"]`, which is how you say "the records are under this key". The **paginator** owns how to get the next page. On top of that sit an authenticator (bearer token, API key in a header or query parameter, OAuth2 with a refresh token), an incremental cursor (`DatetimeBasedCursor` with `cursor_field`, `datetime_format`, `start_datetime` and a `step` for windowed reads), and a partition router for streams that must be read per parent record. Interpolation strings like `{{ config['api_key'] }}`, `{{ stream_partition.id }}` and `{{ response.next_cursor }}` glue the pieces to runtime values, and `$parameters` plus YAML anchors let you factor the repetition out across twenty near-identical streams. That covers a startling share of SaaS APIs. Most REST APIs are: authenticate with a header, GET a collection, read records from one JSON key, follow a cursor or offset, filter by an updated-at parameter. For those, YAML removes an entire category of work — you are not writing, testing or upgrading Python. ## Where the manifest runs out The honest boundaries are: - **Not HTTP.** Databases, object stores, message queues and SDK-only vendors are Python territory; the declarative components assume an HTTP request/response. - **Response needs decoding, not extraction.** A path extractor navigates JSON. XML, CSV, NDJSON inside a gzip, or a payload where records must be reshaped or joined before emission wants code. - **Stateful or multi-step request flows.** Report APIs that make you POST a job, poll for completion, then download a file are awkward to express as a simple retriever, and custom request signing (an HMAC over the body, a rotating nonce) has no declarative component. - **Genuinely odd pagination or error semantics.** APIs that page by mutating a filter, or that signal end-of-data with a 404, will fight the standard strategies. A useful smell test: if you find yourself writing a Jinja expression with real logic in it, you have outgrown the manifest. ## You do not have to choose globally The manifest supports custom components: a component block with a `class_name` pointing at a Python class in your connector package, used for a custom requester, extractor, paginator strategy or authenticator. That means the common case is "declarative manifest with one custom extractor", not "rewrite everything in Python". Preserving the declarative skeleton keeps most of the maintenance benefit. ## The maintenance argument, which is the real one Interviewers usually want the second-order answer. A YAML connector has a smaller blast radius when the CDK upgrades: you are consuming versioned components rather than subclassing implementation. It is readable by someone who is not a Python developer — an analyst can see which endpoint feeds which stream. And it is uniform: a fleet of thirty declarative connectors looks like one thing with thirty configurations, whereas thirty hand-written Python connectors are thirty codebases with thirty sets of quirks and thirty flavours of retry logic. Against that, YAML debugging is worse when it goes wrong — a failing interpolation gives you a less direct error than a stack trace — and you are limited to the components that exist. ## How to answer the decision question Start declarative. Prototype in the Connector Builder against the real API, because the test-read tells you within minutes whether the response shape fits an extractor and whether pagination terminates. If it fits, ship YAML. If one part resists, add a custom component. Reach for a fully hand-written Python connector when the source is not an HTTP API at all, or when the extraction is a genuine program rather than a description.
- What does the Connector Builder give you that editing manifest YAML in an editor does not?A live test-read against the real API while you edit: it shows the exact request sent, the raw response, and the records the extractor pulled out, plus the pages it followed and the state it emitted. That collapses the build-image-and-run-a-sync loop into seconds and is the fastest way to find a wrong field path or a paginator that never terminates.
- How would you handle an API that requires a custom request signature inside a declarative manifest?Write the signing logic as a Python class and reference it from the manifest as a custom authenticator or requester component via its `class_name`. The rest of the manifest — extraction, pagination, cursor — stays declarative, so you keep the maintenance benefit while writing code only where the vendor is unusual.
- Why might a team standardise on the declarative manifest even for connectors that a Python developer could write quickly?Uniformity across a fleet. Thirty manifests share one retry, pagination and checkpoint implementation and upgrade together with the CDK, while thirty bespoke connectors accumulate thirty sets of quirks. Manifests are also legible to non-Python colleagues, so ownership can sit with the team that consumes the data.
saying these in an interview costs you the question
- Claiming the low-code manifest only handles trivial toy APIs
- Assuming YAML connectors cannot do incremental syncs
- Believing you must rewrite in Python for one awkward endpoint
- Treating the Connector Builder as a different product from the manifest
- Ignoring that the source may not be HTTP at all