skip to content

Connect Framework and Runtime

The Connect worker model: standalone versus distributed mode, the internal config/offset/status topics, and plugin loading. Interviewers ask because distributed mode's coordination is what makes Connect fault-tolerant.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is Kafka Connect, and what problem does it solve compared to writing your own producer/consumer applications?

level: juniorimportance: must knowfreq 70%

answer

  1. framework, not custom code
  2. source = into Kafka; sink = out of Kafka
  3. connector → tasks → workers
  4. converters handle format
  5. not a processing engine

basics

~10 s

Kafka Connect is a framework for streaming data between Kafka and external systems (databases, files, S3) using ready-made connectors, so you configure plugins instead of writing custom producer/consumer code.

solid answer

~40 s

Kafka Connect is a runtime and framework, shipped with Apache Kafka, for moving data between Kafka and external systems without writing bespoke code. Source connectors pull data from systems like databases into Kafka topics; sink connectors push from topics into systems like Elasticsearch or S3. You deploy reusable connector plugins and configure them with JSON/properties — declaring topics, converters, and connection details. Connect handles the hard cross-cutting concerns for you: scalability via worker clusters, fault tolerance and rebalancing, offset tracking so work resumes after a restart, at-least-once (or exactly-once for sources) delivery, schema/format conversion via converters, and a REST API for management. Writing your own producers/consumers means re-implementing all of that per integration; Connect standardizes it.

go deeper

for a junior

Know the one-line definition and the source-vs-sink direction.

for a middle

Explain connector/task/worker/converter roles and why Connect beats hand-rolled producers.

for a senior

Discuss delivery guarantees, SMT limits vs. Streams, and when Connect is the wrong tool.

for a principal

Frame Connect's place in a data platform: standardized ingest/egress layer, schema governance via converters/registry, and ops model across teams.

## What Kafka Connect is **Apache Kafka** is a distributed log/messaging system. **Kafka Connect** is a separate component bundled with Kafka whose only job is to move data *between Kafka and other systems* reliably and at scale, without you writing integration code. ### The two directions - **Source connector**: reads from an external system (e.g., a PostgreSQL table, a file, a CDC stream) and *produces* records into Kafka topics. - **Sink connector**: *consumes* from Kafka topics and writes to an external system (e.g., S3, Elasticsearch, JDBC). ### Key vocabulary - **Connector**: a logical job, configured by you, that defines what to copy and how. A connector instance splits work into **tasks**. - **Task**: the unit of parallelism that actually moves data. A connector with `tasks.max=4` may run up to 4 tasks across the cluster. - **Worker**: a JVM process running the Connect runtime; it hosts and executes tasks. - **Converter**: serializes/deserializes record keys and values to/from a wire format (JSON, Avro, Protobuf) — set via `key.converter`/`value.converter`. - **Plugin**: the packaged connector/converter/transform JARs the worker loads from `plugin.path`. ### Why not just write a producer/consumer? A hand-rolled integration forces you to re-solve: parallelism and load balancing, restart/resume (offset bookkeeping), failure handling and retries, scaling out, schema conversion, and operational control (start/stop/reconfigure). Connect provides all of these as a **framework**: you write configuration, not plumbing. For common systems a maintained connector already exists, so often you write *zero* code. ### Edge cases / nuances - Connect is **not** a stream-processing engine — for joins/aggregations/transformations beyond simple per-record **Single Message Transforms (SMTs)**, use Kafka Streams or ksqlDB. - Source connectors can achieve **exactly-once** (KIP-618, `exactly.once.source.support`); sinks are typically **at-least-once** and need idempotent writes downstream.

  • What is the difference between a source connector and a sink connector?
    A source connector imports data from an external system into Kafka topics (it produces); a sink connector exports data from Kafka topics into an external system (it consumes).
  • When would you NOT use Kafka Connect?
    When you need real stream processing — joins, windowed aggregations, enrichment across streams. Connect only does ingest/egress plus light per-record SMTs; use Kafka Streams or ksqlDB for processing.

saying these in an interview costs you the question

  • Calling Connect a stream-processing engine like Kafka Streams.
  • Saying you must write Java code for every integration (maintained connectors are configured, not coded).
  • Confusing source/sink direction (source = into Kafka).

context

open as a page

Compare standalone mode and distributed mode in Kafka Connect. When would you choose each, and how does state/configuration storage differ?

level: middleimportance: must knowfreq 75%

basics

~20 s

Standalone runs a single worker and stores offsets in a local file — simple, no fault tolerance, good for dev. Distributed runs multiple coordinated workers that store config, offsets, and status in Kafka topics, giving scalability and fault tolerance.

open as a page

Describe the three internal topics a distributed Connect cluster uses (config, offset, status). What does each store, and what configuration do they require?

level: seniorimportance: must knowfreq 60%

basics

~10 s

config.storage.topic holds connector/task configs (single partition, compacted). offset.storage.topic holds source-connector read positions (many partitions, compacted). status.storage.topic holds connector/task/worker status (many partitions, compacted). All need high replication.

open as a page

What are the essential worker-level configuration properties needed to bootstrap a distributed Connect worker, and what does each control?

level: middleimportance: should knowfreq 45%

basics

~20 s

You need bootstrap.servers (Kafka brokers), group.id (cluster identity), the three internal topic names with replication factors, key/value converters, plugin.path, and the REST listener. These let the worker connect to Kafka, join its cluster, and load plugins.

open as a page

What is plugin.path in Kafka Connect, and how does the runtime isolate connector plugins? What problems does misconfiguring it cause?

level: seniorimportance: should knowfreq 50%

basics

~10 s

plugin.path is a list of directories where the worker finds connector/converter/transform JARs. Connect loads each plugin in an isolated classloader to avoid dependency conflicts. Misconfiguring it causes 'class not found' errors or dependency clashes.

open as a page

How do workers in a distributed Connect cluster coordinate work assignment, and how did incremental cooperative rebalancing (KIP-415) improve the original protocol?

level: principalimportance: should knowfreq 35%

basics

~20 s

Workers sharing a group.id use Kafka's group-membership protocol to elect a leader that assigns connectors and tasks. The original protocol stopped all work on every change (stop-the-world); incremental cooperative rebalancing (KIP-415) only reassigns the affected tasks, avoiding global pauses.

open as a page