Which four commands must an Airbyte source connector implement?
answer
- Four verbs, one Docker entrypoint
- Config form, credential test, stream list, data
- JSON messages on stdout, one per line
- spec / check / discover / read
basics
~20 sAn Airbyte source implements spec, check, discover and read. spec returns the JSON-Schema config form, check validates credentials, discover returns the catalog of streams and their schemas, and read emits records and state messages on stdout.
solid answer
~40 sA source connector is a Docker image whose entrypoint accepts four commands. **`spec`** returns a `ConnectorSpecification` whose `connectionSpecification` is a JSON Schema — the platform renders it as the setup form, and fields marked `airbyte_secret` are masked. **`check`** takes that config and answers one question: can I connect and authenticate? It returns a connection status of succeeded or failed. **`discover`** takes the config and returns an `AirbyteCatalog`: every stream the source can offer, each with a JSON Schema and its supported sync modes. **`read`** takes the config, a `ConfiguredAirbyteCatalog` (the subset of streams the user actually selected, with the chosen sync mode) and an optional state blob, then writes newline-delimited `AirbyteMessage` JSON to stdout — `RECORD` messages, `STATE` checkpoints, plus `LOG` and `TRACE`. The CDK implements all four for you; you supply the streams.
code
text · 5 linesdocker run --rm my-source spec
docker run --rm -v $(pwd):/s my-source check --config /s/config.json
docker run --rm -v $(pwd):/s my-source discover --config /s/config.json
docker run --rm -v $(pwd):/s my-source read \
--config /s/config.json --catalog /s/configured_catalog.json --state /s/state.jsongo deeper
Memorise the four commands and what each one returns, and be able to say that a connector is a Docker image talking JSON over stdout. That alone answers the screening version of this question.
Be ready to explain the message envelope: RECORD, STATE, LOG, TRACE and the command-specific SPEC, CONNECTION_STATUS and CATALOG, and why the configured catalog differs from the discovered one.
Show that you debug with these commands — running the image by hand with a one-stream configured catalog to isolate a failing stream, and distinguishing configuration errors from system errors in TRACE output.
Talk about the protocol as a stable boundary: it is why connectors can be written in any language, versioned as images, and swapped without touching the platform, and what that costs in schema and spec compatibility discipline.
## The protocol is the whole contract Airbyte does not link your connector into its runtime. A connector is a Docker image, and the platform interacts with it entirely through a command-line entrypoint plus newline-delimited JSON on stdout. That is the Airbyte protocol, and it is deliberately small: four commands and a handful of message types. Everything else — the Python CDK, the low-code declarative manifest, the Connector Builder UI — is convenience layered on top of these four commands. Knowing them is what lets you debug a connector by running the image by hand instead of guessing at the UI. ## `spec` `spec` takes no arguments and returns a `ConnectorSpecification`. Its important field is `connectionSpecification`, a JSON Schema describing the configuration the connector needs: API key, subdomain, start date, and so on. The platform renders that schema as the source setup form, so the schema's `title`, `description`, `order` and `required` fields are user-facing copy. Any property that holds a credential must be marked `airbyte_secret: true` so the value is masked in the UI and redacted in logs. The spec may also declare OAuth wiring so the platform can run the consent flow for the user. Because the spec is user-facing, changing it is a compatibility event, not a private refactor. ## `check` `check` takes `--config` and answers one narrow question: with these credentials, can the connector reach the source? It returns a `CONNECTION_STATUS` message of `SUCCEEDED` or `FAILED` with a message. The discipline is to make `check` cheap but honest: hit an endpoint that actually requires the credential (not an unauthenticated health route), and don't list a million rows to prove connectivity. A `check` that passes while `read` immediately 403s is the classic bad implementation — usually because it called a public endpoint or ignored a scope the real streams need. ## `discover` `discover` takes `--config` and returns an `AirbyteCatalog`: a list of streams, each with a `name`, a `json_schema` describing its records, `supported_sync_modes`, and optionally `source_defined_cursor`, `default_cursor_field` and `source_defined_primary_key`. For an API source the schemas are usually static JSON files shipped with the connector; for a database source they are derived by inspecting the information schema at runtime. The catalog is what the user picks streams from, and it is what downstream typing is built on, so a schema that lies about a field's type causes destination errors far from the connector. ## `read` `read` takes `--config`, `--catalog` and optionally `--state`. Note that the catalog passed to `read` is a `ConfiguredAirbyteCatalog`, not the raw discovered one: it is the subset of streams the user selected, each annotated with the chosen `sync_mode` and `destination_sync_mode` and, where relevant, the cursor and primary-key fields. The connector's job is to emit `RECORD` messages for those streams and to emit `STATE` messages representing progress that the platform will hand back as `--state` on the next run. Everything goes to stdout as one JSON object per line; anything you print that is not a valid `AirbyteMessage` corrupts the stream, which is why connectors must log through the CDK logger rather than using bare `print`. ## The message types The envelope is `AirbyteMessage`, discriminated by `type`. `SPEC`, `CONNECTION_STATUS` and `CATALOG` are the outputs of the first three commands. `RECORD` carries `stream`, `data` and `emitted_at`. `STATE` carries a checkpoint. `LOG` carries a level and message. `TRACE` carries structured errors and status, which is how a connector reports a configuration error distinctly from a transient system error so the UI can tell the user something useful. ## Why the split matters in practice The separation is what makes the platform generic: the same four commands work for a REST API, a database, a file store or a destination, so the scheduler, the UI and the destination side never need connector-specific code. It also gives you a debugging ladder. Run `spec` to prove the image is sane, `check` to isolate credentials, `discover` to see whether the stream you expect even exists, and `read` with a hand-written configured catalog containing a single stream when you want to reproduce one stream's failure without waiting for a full sync. Most connector bugs get localised in minutes this way. ## Common mistakes Writing non-JSON to stdout; making `check` weaker than `read`; hardcoding a schema in `discover` that no longer matches what the API returns; and forgetting that `read` receives the *configured* catalog, so a connector that ignores it and reads every stream burns quota on data the user never asked for.
- What is the difference between the catalog returned by discover and the one passed to read?`discover` returns an `AirbyteCatalog` — everything the source *could* offer. `read` receives a `ConfiguredAirbyteCatalog`: only the streams the user selected, each annotated with the chosen sync mode, destination sync mode, and cursor or primary-key fields. A connector must honour that selection rather than reading everything.
- Why must a connector never use a bare print statement?stdout is the protocol channel. The platform parses each line as an `AirbyteMessage`, so stray text corrupts or truncates the record stream. Use the CDK's logger, which wraps output in `LOG` messages, or write to stderr.
- How do you tell the platform that a failure is the user's misconfiguration rather than a bug?Emit a `TRACE` message with an error whose failure type marks it as a configuration error. The UI then surfaces the message to the user instead of presenting an internal stack trace, and the platform will not treat it as a retryable system fault.
saying these in an interview costs you the question
- Thinking the connector writes to the destination itself
- Assuming check runs the same queries as read
- Confusing the discovered catalog with the configured one
- Printing debug output to stdout during read
- Believing the platform imports the connector as a library