How does Kafka Connect validate a connector configuration before accepting it, and how does PUT /connector-plugins/{class}/config/validate help you build safe deployment tooling?
answer
- PUT /connector-plugins/{class}/config/validate
- non-mutating pre-flight
- error_count + configs[].value.errors
- ConfigDef + Connector.validate()
- same validation -> 400 on POST/PUT
basics
~20 sConnect runs each connector's Validator over the proposed config. PUT /connector-plugins/{class}/config/validate returns a structured per-field result (definitions, current values, recommended values, and per-field errors) without creating anything, so tooling can pre-flight a config and surface errors before deploying.
solid answer
~50 sBefore a connector is created/updated, Connect invokes the connector plugin's config validation: each ConfigDef property is checked and the connector's validate() method can add cross-field and connectivity checks. The REST surface for this is PUT /connector-plugins/{className}/config/validate, whose body is the flat config map. It returns a JSON document with name, error_count, groups, and a configs array; each entry has a definition (key, type, importance, required), a value block (current value, recommended_values, visible, and an errors list of messages), so you see exactly which fields are wrong and why. This is non-mutating — nothing is created. Tooling uses it as a pre-flight: validate first, fail the pipeline on error_count > 0 with the field-level messages, and only POST/PUT the config when clean. The same validation runs implicitly inside POST /connectors and PUT /config, which return 400 with this structure when a config is invalid.
go deeper
Know that Connect checks a config before accepting it and rejects bad ones with 400.
Know the validate endpoint exists and returns per-field errors without creating anything.
Use the non-mutating validate endpoint as a CI pre-flight; parse error_count and field errors; understand ConfigDef + validate().
Architect a config-as-code pipeline: lint/validate at PR time, gate deploys on error_count, drive UIs from definitions/recommendations, and separate connectivity vs schema failures.
**Why validation exists.** A connector's behavior is driven entirely by its string config map, and a typo or a missing required property would otherwise only surface as a runtime task failure. Kafka Connect therefore validates configs up front using two layers: 1. **ConfigDef** — every connector plugin declares its properties via a `ConfigDef` (key, type, default, importance, validators, recommenders). Connect checks types, required-ness, and any per-field `Validator` (e.g., range, enum). 2. **Connector.validate(Map)** — the plugin can override `validate()` to add **cross-field** logic and even **live connectivity checks** (e.g., a JDBC connector can try the connection.url, a sink can verify the target exists). Recommenders can also populate dynamic `recommended_values` (e.g., list available tables). **The dedicated endpoint — PUT /connector-plugins/{className}/config/validate.** Body is the flat config map (you must include `connector.class`). It is **non-mutating**: it validates and returns results without creating or changing any connector. The response shape: ``` { "name": "io.confluent.connect.jdbc.JdbcSourceConnector", "error_count": 1, "groups": ["Common", "Database"], "configs": [ { "definition": { "name": "connection.url", "type": "STRING", "required": true, "importance": "HIGH" }, "value": { "name": "connection.url", "value": null, "recommended_values": [], "errors": ["Missing required configuration"], "visible": true } }, ... ] } ``` Key fields: `error_count` (zero means valid), and per-config `value.errors` (the human-readable messages) plus `value.recommended_values` (UI/autocomplete hints). **How create/update use it.** POST /connectors and PUT /connectors/{name}/config run the **same** validation internally. If it fails they return **400 Bad Request** carrying this exact structure, so the field-level errors are available on the create path too. Some implementations let you submit and accept config that has warnings but not hard errors. **Building safe tooling (the payoff).** Because the validate endpoint is side-effect-free, deployment pipelines can: - **Pre-flight every config**: PUT to validate, parse `error_count`; if > 0, fail the CI job and print each `definition.name` + its `errors`. - **Lint at PR time** without a running production connector being touched. - **Drive UIs**: render forms from `definition` (type/importance/required) and offer dropdowns from `recommended_values`. - **Catch environment problems early** when the connector implements connectivity checks in validate() (wrong URL, bad credentials) before the connector ever runs and fails a task. **Edge cases / gotchas:** - Validation that performs connectivity checks can be **slow** or can **fail for transient reasons**; pipelines should treat connectivity errors distinctly from schema errors. - The endpoint validates against the plugin **on that worker** — the class must be installed on the Connect cluster you call. - A clean validate does not guarantee runtime success (data-level issues still surface at task time), but it eliminates the entire class of config-shape and missing-property failures before deploy.
- Is the validate endpoint side-effecting? Why does that matter for CI?No — it validates without creating or changing any connector, so a CI pipeline can pre-flight configs safely on a real cluster without risk of partially deploying a bad connector.
- What status and body do POST /connectors and PUT /config return for an invalid config?400 Bad Request with the same validation structure (per-field errors and error_count), because they run the identical validation internally before accepting the config.
saying these in an interview costs you the question
- Saying the validate endpoint creates or modifies the connector — it is non-mutating.
- Thinking validation only checks types — connectors can add cross-field and connectivity checks via validate().
- Assuming a clean validate guarantees the connector won't fail at runtime — data-level errors still occur.
- Forgetting the plugin class must be installed on the worker you call.