skip to content

Airbyte

An open-source EL platform with a large connector catalog and an SDK for writing your own, self-hostable or managed. It is the usual reference point when a team weighs building connectors against buying a managed ELT service.

on this pageshow

explore

questions

12

Which four commands must an Airbyte source connector implement?

level: juniorimportance: must knowfreq 58%

answer

  1. Four verbs, one Docker entrypoint
  2. Config form, credential test, stream list, data
  3. JSON messages on stdout, one per line
  4. spec / check / discover / read

basics

~20 s

An Airbyte source implements spec, check, discover and read. spec returns the JSON-Schema config form, check validates credentials, discover returns the catalog of streams and their schemas, and read emits records and state messages on stdout.

solid answer

~40 s

A source connector is a Docker image whose entrypoint accepts four commands. **`spec`** returns a `ConnectorSpecification` whose `connectionSpecification` is a JSON Schema — the platform renders it as the setup form, and fields marked `airbyte_secret` are masked. **`check`** takes that config and answers one question: can I connect and authenticate? It returns a connection status of succeeded or failed. **`discover`** takes the config and returns an `AirbyteCatalog`: every stream the source can offer, each with a JSON Schema and its supported sync modes. **`read`** takes the config, a `ConfiguredAirbyteCatalog` (the subset of streams the user actually selected, with the chosen sync mode) and an optional state blob, then writes newline-delimited `AirbyteMessage` JSON to stdout — `RECORD` messages, `STATE` checkpoints, plus `LOG` and `TRACE`. The CDK implements all four for you; you supply the streams.

code

text · 5 lines
text
docker run --rm my-source spec
docker run --rm -v $(pwd):/s my-source check --config /s/config.json
docker run --rm -v $(pwd):/s my-source discover --config /s/config.json
docker run --rm -v $(pwd):/s my-source read \
  --config /s/config.json --catalog /s/configured_catalog.json --state /s/state.json

go deeper

for a junior

Memorise the four commands and what each one returns, and be able to say that a connector is a Docker image talking JSON over stdout. That alone answers the screening version of this question.

for a middle

Be ready to explain the message envelope: RECORD, STATE, LOG, TRACE and the command-specific SPEC, CONNECTION_STATUS and CATALOG, and why the configured catalog differs from the discovered one.

for a senior

Show that you debug with these commands — running the image by hand with a one-stream configured catalog to isolate a failing stream, and distinguishing configuration errors from system errors in TRACE output.

for a principal

Talk about the protocol as a stable boundary: it is why connectors can be written in any language, versioned as images, and swapped without touching the platform, and what that costs in schema and spec compatibility discipline.

## The protocol is the whole contract Airbyte does not link your connector into its runtime. A connector is a Docker image, and the platform interacts with it entirely through a command-line entrypoint plus newline-delimited JSON on stdout. That is the Airbyte protocol, and it is deliberately small: four commands and a handful of message types. Everything else — the Python CDK, the low-code declarative manifest, the Connector Builder UI — is convenience layered on top of these four commands. Knowing them is what lets you debug a connector by running the image by hand instead of guessing at the UI. ## `spec` `spec` takes no arguments and returns a `ConnectorSpecification`. Its important field is `connectionSpecification`, a JSON Schema describing the configuration the connector needs: API key, subdomain, start date, and so on. The platform renders that schema as the source setup form, so the schema's `title`, `description`, `order` and `required` fields are user-facing copy. Any property that holds a credential must be marked `airbyte_secret: true` so the value is masked in the UI and redacted in logs. The spec may also declare OAuth wiring so the platform can run the consent flow for the user. Because the spec is user-facing, changing it is a compatibility event, not a private refactor. ## `check` `check` takes `--config` and answers one narrow question: with these credentials, can the connector reach the source? It returns a `CONNECTION_STATUS` message of `SUCCEEDED` or `FAILED` with a message. The discipline is to make `check` cheap but honest: hit an endpoint that actually requires the credential (not an unauthenticated health route), and don't list a million rows to prove connectivity. A `check` that passes while `read` immediately 403s is the classic bad implementation — usually because it called a public endpoint or ignored a scope the real streams need. ## `discover` `discover` takes `--config` and returns an `AirbyteCatalog`: a list of streams, each with a `name`, a `json_schema` describing its records, `supported_sync_modes`, and optionally `source_defined_cursor`, `default_cursor_field` and `source_defined_primary_key`. For an API source the schemas are usually static JSON files shipped with the connector; for a database source they are derived by inspecting the information schema at runtime. The catalog is what the user picks streams from, and it is what downstream typing is built on, so a schema that lies about a field's type causes destination errors far from the connector. ## `read` `read` takes `--config`, `--catalog` and optionally `--state`. Note that the catalog passed to `read` is a `ConfiguredAirbyteCatalog`, not the raw discovered one: it is the subset of streams the user selected, each annotated with the chosen `sync_mode` and `destination_sync_mode` and, where relevant, the cursor and primary-key fields. The connector's job is to emit `RECORD` messages for those streams and to emit `STATE` messages representing progress that the platform will hand back as `--state` on the next run. Everything goes to stdout as one JSON object per line; anything you print that is not a valid `AirbyteMessage` corrupts the stream, which is why connectors must log through the CDK logger rather than using bare `print`. ## The message types The envelope is `AirbyteMessage`, discriminated by `type`. `SPEC`, `CONNECTION_STATUS` and `CATALOG` are the outputs of the first three commands. `RECORD` carries `stream`, `data` and `emitted_at`. `STATE` carries a checkpoint. `LOG` carries a level and message. `TRACE` carries structured errors and status, which is how a connector reports a configuration error distinctly from a transient system error so the UI can tell the user something useful. ## Why the split matters in practice The separation is what makes the platform generic: the same four commands work for a REST API, a database, a file store or a destination, so the scheduler, the UI and the destination side never need connector-specific code. It also gives you a debugging ladder. Run `spec` to prove the image is sane, `check` to isolate credentials, `discover` to see whether the stream you expect even exists, and `read` with a hand-written configured catalog containing a single stream when you want to reproduce one stream's failure without waiting for a full sync. Most connector bugs get localised in minutes this way. ## Common mistakes Writing non-JSON to stdout; making `check` weaker than `read`; hardcoding a schema in `discover` that no longer matches what the API returns; and forgetting that `read` receives the *configured* catalog, so a connector that ignores it and reads every stream burns quota on data the user never asked for.

  • What is the difference between the catalog returned by discover and the one passed to read?
    `discover` returns an `AirbyteCatalog` — everything the source *could* offer. `read` receives a `ConfiguredAirbyteCatalog`: only the streams the user selected, each annotated with the chosen sync mode, destination sync mode, and cursor or primary-key fields. A connector must honour that selection rather than reading everything.
  • Why must a connector never use a bare print statement?
    stdout is the protocol channel. The platform parses each line as an `AirbyteMessage`, so stray text corrupts or truncates the record stream. Use the CDK's logger, which wraps output in `LOG` messages, or write to stderr.
  • How do you tell the platform that a failure is the user's misconfiguration rather than a bug?
    Emit a `TRACE` message with an error whose failure type marks it as a configuration error. The UI then surfaces the message to the user instead of presenting an internal stack trace, and the platform will not treat it as a retryable system fault.

saying these in an interview costs you the question

  • Thinking the connector writes to the destination itself
  • Assuming check runs the same queries as read
  • Confusing the discovered catalog with the configured one
  • Printing debug output to stdout during read
  • Believing the platform imports the connector as a library

context

open as a page

What does each Airbyte sync mode do to the destination table?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Airbyte pairs a read mode with a write mode. Full Refresh Overwrite replaces the table each sync; Full Refresh Append re-appends everything; Incremental Append adds only records past the cursor; Incremental Append + Deduped keeps one current row per primary key.

open as a page

How does a custom Airbyte source emit cursor state so the next sync resumes correctly?

level: middleimportance: must knowfreq 52%

basics

~20 s

The stream declares a cursor field, then emits STATE messages during read holding the highest cursor value whose records have already been emitted. The platform persists the last STATE it saw and hands it back as the next run's input, so the connector resumes from there.

open as a page

What does Airbyte's Incremental | Append + Deduped sync mode require from a stream?

level: middleimportance: must knowfreq 68%

basics

~20 s

It needs a cursor field to decide what is new and a primary key to collapse it. Airbyte appends the incremental records to a raw table, then rebuilds a final table holding the highest-cursor row per primary key.

open as a page

When should you build an Airbyte connector with the low-code YAML manifest instead of the Python CDK?

level: middleimportance: should knowfreq 50%

basics

~20 s

Use the low-code declarative manifest when the source is a plain REST API with JSON responses, standard auth and standard pagination. Drop to the Python CDK when the source is not HTTP, the response needs real parsing, or the request logic is stateful.

open as a page

How do you configure pagination in an Airbyte low-code connector manifest?

level: middleimportance: should knowfreq 44%

basics

~20 s

Attach a DefaultPaginator to the retriever, pick a strategy — CursorPagination, OffsetIncrement or PageIncrement — and use page_token_option and page_size_option to say where the token and size are injected into the request. Always set a stop condition.

open as a page

Why can an Airbyte incremental sync on an updated_at cursor silently miss rows?

level: middleimportance: should knowfreq 58%

basics

~20 s

An incremental sync only returns records whose cursor value is at or past the stored one. Rows changed without bumping updated_at, rows with a null cursor, rows committed late with an older timestamp, and hard deletes are therefore never read.

open as a page

Your custom Airbyte source keeps failing on HTTP 429 — how do you make it back off?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Handle 429 in the connector's retry layer: mark it retryable, derive the wait from the Retry-After header when the API sends one and fall back to exponential backoff, cap the attempts, and checkpoint state often so a final failure loses little work.

open as a page

When would you choose Full Refresh | Overwrite over incremental for an Airbyte stream?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Choose overwrite when the stream is small enough to re-read and correctness matters more than cost: no trustworthy cursor, rows mutated without bumping one, or hard deletes that must disappear. Overwrite self-heals every run; incremental never revisits what it already passed.

open as a page

If an Airbyte sync fails mid-stream, does the next attempt resume or start over?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It depends on whether progress was checkpointed. Airbyte only commits state for records the destination has accepted, so a connector that emits state during the read resumes from the last committed point; one that emits state only at the end restarts the stream.

open as a page

What does Airbyte write into its raw destination tables before typing and deduping?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

Each record lands as a JSON payload plus Airbyte metadata — a generated record id, an extraction timestamp and a loaded timestamp — in a raw table in a separate internal schema. A later step parses it into typed columns in the final table.

open as a page

What do Airbyte's connector acceptance tests check before a connector ships?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

They run the connector image against a real source and assert protocol conformance: the spec is valid and marks secrets, check passes on good config and fails on bad, discover returns a valid catalog, read returns records matching declared schemas, and incremental reads honour state.

open as a page