skip to content

Explain the relationship between a Connector and its Tasks. What is the role of taskConfigs() and tasks.max?

level: middleimportance: must knowfreq 72%

answer

  1. Connector coordinates, Task executes
  2. taskConfigs(maxTasks) -> list size = task count
  3. tasks.max = ceiling, not guarantee
  4. Sink parallelism capped by partition count
  5. requestTaskReconfiguration on partition change

basics

~20 s

A Connector is a single coordinator instance that splits work into Tasks, the things that actually move data. taskConfigs(maxTasks) returns the config for each task; tasks.max caps how many tasks Connect will run for that connector.

solid answer

~50 s

A Connector instance is a coordinator: there is exactly one per connector config, and it doesn't move data. Its job is to figure out HOW to divide the workload and hand back a list of per-task configurations via taskConfigs(int maxTasks). Connect then instantiates that many Task objects (SourceTask or SinkTask) and distributes them across worker JVMs in the cluster; the tasks do the actual polling/putting. tasks.max is the upper bound you set in config — Connect will create at most tasks.max tasks, and the connector's taskConfigs() should return no more than that. The connector may return fewer if it can't usefully parallelize (e.g. a JDBC source with one table returns one task config even if tasks.max=10). The Connector also calls context.requestTaskReconfiguration() when the external partitioning changes (e.g. a new table appears), prompting Connect to re-invoke taskConfigs() and rebalance.

go deeper

for a junior

Know there's one connector that creates several tasks which do the work.

for a middle

Explain taskConfigs() returning per-task configs and tasks.max as a ceiling.

for a senior

Discuss why sink parallelism is bounded by partition count and how reconfiguration triggers rebalance.

for a principal

Design partition-aware task splitting, capacity-plan tasks.max vs source/sink parallelism, and reason about failure isolation per task.

## Two layers: Connector and Task Kafka Connect deliberately separates *coordination* from *execution*: - **Connector** (`SourceConnector` or `SinkConnector` subclass): a single logical instance per connector configuration. It runs on exactly one worker (the leader for that connector). It is a control-plane object — it validates configuration, decides how to break the job into parallel units, and reacts to changes in the external system. **It moves zero data.** - **Task** (`SourceTask` or `SinkTask` subclass): the data-plane worker. Connect creates 1..N task instances and spreads them across the worker JVMs in the cluster. Each task runs its own loop: a SourceTask repeatedly calls `poll()`; a SinkTask receives batches via `put()`. ## taskConfigs(int maxTasks) When a connector starts (or requests reconfiguration), Connect calls: ``` List<Map<String,String>> taskConfigs(int maxTasks) ``` The connector returns a **list of property maps**, one per task. The size of that list is the number of tasks Connect will launch. `maxTasks` is the value of `tasks.max`. The connector partitions its external work and bakes the partition assignment into each map. Example: a JDBC source with 4 tables and `tasks.max=2` might return 2 configs, each assigned 2 tables via a property like `tables=t1,t2`. Key rule: **the connector returns at most maxTasks configs, but may return fewer** when more parallelism wouldn't help. A FileStreamSource reading one file always returns exactly one task config regardless of tasks.max. ## tasks.max `tasks.max` is a connector-level config property (default 1). It is the ceiling on task count. It does NOT guarantee that many tasks run — the connector decides the actual number, bounded above by this value. Setting tasks.max higher than the available parallelism (e.g. more than the number of topic partitions for a sink, or tables for a source) wastes the extra slots; some tasks simply get no work. ## Sink connectors are special For sink connectors, parallelism is ultimately bounded by **the number of partitions across the subscribed topics**, because each Kafka partition is assigned to exactly one consumer (task) at a time. Even if tasks.max=20 and taskConfigs returns 20, if the topics have only 6 partitions total, 14 tasks will be idle. For sinks the connector typically just returns maxTasks identical configs and lets the consumer group's partition assignor distribute partitions. ## Reconfiguration The Connector holds a `ConnectorContext`. When the external partitioning changes — a new database table, a topic gaining partitions — the connector calls `context.requestTaskReconfiguration()`. Connect then re-invokes `taskConfigs()` and performs a rebalance so the new task set picks up the changed assignment. ## Failure isolation Tasks fail independently. A single task hitting an unrecoverable error transitions to FAILED while sibling tasks keep running; you restart individual tasks via the REST API (`POST /connectors/{name}/tasks/{id}/restart`).

  • If I set tasks.max=10 on a sink connector subscribed to a topic with 3 partitions, how many tasks do real work?
    At most 3 — each partition is assigned to one task, so 7 tasks sit idle. Partition count is the real upper bound for sink parallelism.
  • How does a connector tell Connect to re-split the work when the external system changes?
    It calls context.requestTaskReconfiguration(), which makes Connect re-invoke taskConfigs() and rebalance the new task set.

saying these in an interview costs you the question

  • Saying tasks.max guarantees exactly that many tasks run.
  • Claiming the Connector itself reads/writes data.
  • Thinking sink parallelism can exceed total partition count.
  • Believing one task config map maps to multiple tasks.

context