What are the essential worker-level configuration properties needed to bootstrap a distributed Connect worker, and what does each control?
answer
- bootstrap.servers = broker seed list
- group.id = cluster identity
- 3 internal topics + RF
- key/value converters
- plugin.path + REST listeners
basics
~20 sYou need bootstrap.servers (Kafka brokers), group.id (cluster identity), the three internal topic names with replication factors, key/value converters, plugin.path, and the REST listener. These let the worker connect to Kafka, join its cluster, and load plugins.
solid answer
~40 sA distributed worker's `worker.properties` must define: `bootstrap.servers` — the Kafka brokers the worker connects to for all internal topics and for producing/consuming data; `group.id` — the identity that groups workers into one Connect cluster (must be unique per cluster); `config.storage.topic`, `offset.storage.topic`, `status.storage.topic` plus their `*.replication.factor` (and partition counts) for cluster state; `key.converter` and `value.converter` (e.g., JSON/Avro) for serialization defaults; `plugin.path` so the worker can discover connector/converter/SMT plugins; and the REST API binding via `listeners`/`rest.advertised.host.name` for management and inter-worker communication. `bootstrap.servers` is just a seed list — the worker discovers the full broker set from it. Getting `group.id` or internal-topic names wrong is the classic mistake that makes a worker silently form/join the wrong cluster.
go deeper
Know the worker needs bootstrap.servers and where it finds plugins.
List the full essential set: bootstrap.servers, group.id, internal topics, converters, plugin.path, REST.
Explain seed-list discovery, advertised REST for inter-worker forwarding, and per-client overrides.
Standardize worker config templates incl. security prefixes, RF policy, and cluster isolation conventions.
## Bootstrapping a distributed worker A Connect **worker** reads a `worker.properties` file at startup (`connect-distributed.sh worker.properties`). The essential keys: ### Connectivity - **`bootstrap.servers`** — comma-separated `host:port` seed list of Kafka brokers, e.g. `broker1:9092,broker2:9092`. The worker connects here to read/write the internal topics and to produce (sources) / consume (sinks) data. It is only a **seed**: from any reachable broker the worker fetches cluster metadata and learns the full broker set. List a few brokers for resilience if a seed is down. ### Cluster identity - **`group.id`** — the cluster's name. All workers sharing this id form one Connect cluster and rebalance work among themselves. Two clusters on the same Kafka **must** use different `group.id` values (and different internal topic names). ### State topics - **`config.storage.topic`**, **`offset.storage.topic`**, **`status.storage.topic`** — names of the internal topics (covered separately). Each pairs with a `*.replication.factor` (use 3 in prod) and offset/status also take partition counts. ### Serialization - **`key.converter`** / **`value.converter`** — default converters for record keys/values, e.g. `org.apache.kafka.connect.json.JsonConverter` or an Avro converter. Connectors can override per-connector. Also `*.converter.schemas.enable` toggles embedded schemas for JSON. ### Plugins - **`plugin.path`** — directories of plugin JARs the worker loads with classloader isolation. ### REST / inter-worker - **`listeners`** (e.g., `http://0.0.0.0:8083`) and **`rest.advertised.host.name`**/`rest.advertised.port` — the worker exposes a REST API for managing connectors; in distributed mode workers also **forward REST requests to the leader** over this interface, so advertised host/port must be reachable by peers. ### Security (when applicable) - `security.protocol`, SASL/SSL settings, and the `producer.*`, `consumer.*`, `admin.*` prefixes to override client configs for the internal producers/consumers/admin client. ### Common bootstrap mistakes - Reusing another cluster's `group.id` or internal topic names → state collision. - Wrong/unreachable `bootstrap.servers` → worker can't start (can't reach internal topics). - `rest.advertised.host.name` set to an unreachable address → inter-worker request forwarding fails, so non-leader workers can't apply config changes. - Replication factor 1 on internal topics in prod → state loss on broker failure.
- Is bootstrap.servers the full list of brokers the worker will ever use?No. It's a seed list; the worker contacts a reachable seed, fetches cluster metadata, and discovers all brokers. You list a few seeds for resilience, not the whole cluster.
- Why does rest.advertised.host.name matter in a multi-worker cluster?Non-leader workers forward connector management requests to the leader over the REST interface. If the advertised address isn't reachable by peers, config changes routed through a follower fail.
saying these in an interview costs you the question
- Saying bootstrap.servers must list every broker (it's a seed list).
- Omitting group.id or reusing another cluster's group.id/internal topics.
- Forgetting the REST listener is also used for inter-worker request forwarding.