What are the three core connectors that make up MirrorMaker 2, and what does each one do?
answer
- Source = data, Checkpoint = offsets, Heartbeat = liveness
- Built on Kafka Connect (KIP-382)
- One flow = source->target = a set of connectors
- Heartbeats measure end-to-end lag
- Connectors split into tasks for parallelism
basics
~10 sMirrorMaker 2 has three connectors: MirrorSourceConnector copies topic data from the source cluster to the target; MirrorCheckpointConnector copies consumer-group offset progress; MirrorHeartbeatConnector emits periodic heartbeats to confirm the replication path is alive.
solid answer
~40 sMM2 runs as a set of Kafka Connect connectors. MirrorSourceConnector is the workhorse: it consumes records from source-cluster topics and produces them onto the target cluster, also propagating ACLs and (optionally) topic configs. MirrorCheckpointConnector reads committed consumer-group offsets on the source, translates them to equivalent target offsets, and writes them to a checkpoints topic so consumers can fail over. MirrorHeartbeatConnector periodically produces small heartbeat records into a heartbeats topic on both clusters; these prove the end-to-end path works and provide a clock for measuring replication lag. Each connector spawns multiple tasks for parallelism — MirrorSourceConnector tasks are partitioned across the source topic-partitions it owns. Together they form one directional 'flow' (source->target); a bidirectional setup runs two flows.
go deeper
Memorize the three connector names and one-line role for each: source=data, checkpoint=offsets, heartbeat=liveness.
Know each connector splits into tasks, that they run on Kafka Connect, and which config flags enable/disable each.
Explain how connectors map to a directional flow, when you'd disable checkpoints, and how heartbeats give an end-to-end latency signal.
Reason about running MM2 on a dedicated vs shared Connect cluster, task-level parallelism limits, and the operational trade-offs of each connector for a DR vs active-active design.
MirrorMaker 2 (MM2) is Kafka's cross-cluster replication tool, introduced in KIP-382. Unlike the legacy MirrorMaker 1 (a standalone consumer+producer loop), MM2 is built **on top of Kafka Connect**, so it inherits Connect's distributed runtime, REST API, offset/config/status topics, scaling via tasks, and fault tolerance. A **connector** in Kafka Connect is a plugin that defines how to move data; it splits its work into **tasks** that run on **workers** (JVM processes). MM2 ships three connector classes, each with a distinct job: 1. **MirrorSourceConnector** — the data-plane connector. It discovers which source-cluster topics match the configured allow/deny filters, then consumes their records and produces them to correspondingly-named topics on the **target** cluster (the remote topic name is prefixed, e.g. `us-west.orders`). It also creates target topics if missing, replicates ACLs, and (when `sync.topic.configs.enabled=true`) keeps target topic configs in sync. It parallelizes by assigning source topic-partitions across its tasks. 2. **MirrorCheckpointConnector** — replicates **consumer progress**, not data. It periodically reads committed offsets of source-cluster consumer groups (via `__consumer_offsets`), maps each source offset to the equivalent offset of the mirrored record on the target, and emits **checkpoint** records to a `<source>.checkpoints.internal` topic on the target. This lets a consumer that fails over to the target resume near where it left off. (The exact offset-translation math is owned by a sibling topic and out of scope here.) 3. **MirrorHeartbeatConnector** — a liveness/observability connector. It produces tiny **heartbeat** records (a timestamp and the source/target alias) into a `heartbeats` topic at a fixed interval (`emit.heartbeats.interval.seconds`, default 1s). Because heartbeats are themselves mirrored, you can observe them arriving on the target and measure end-to-end replication latency, and confirm the path is healthy even when no business data is flowing. **Key edge cases:** Each connector is independently enable-able (`MirrorSourceConnector` via `source->target.enabled`, plus `emit.checkpoints.enabled` and `emit.heartbeats.enabled` flags). Heartbeats are cheap and almost always on. You can run the source connector alone if you only need data and don't care about consumer failover. In a dedicated MM2 cluster (the `connect-mirror-maker.sh` driver), all three are auto-configured from a single `mm2.properties`; in a generic Connect cluster you POST each connector's config to the REST API yourself.
- Why is MM2 built on Kafka Connect instead of being a standalone loop like MirrorMaker 1?To inherit Connect's distributed scaling (tasks/workers), REST management, automatic offset/config/status persistence in internal topics, rebalancing, and fault tolerance — MM1 was a single-process consumer+producer with no built-in HA, scaling, or offset translation.
- If you only need to replicate data for disaster recovery and don't care about per-consumer failover, which connectors can you skip?You can disable the checkpoint connector (set emit.checkpoints.enabled=false). Keep MirrorSourceConnector for data; heartbeats are cheap and worth keeping for health monitoring, but are also optional.
saying these in an interview costs you the question
- Saying MM2 is a single process/loop like MirrorMaker 1 — it runs on Kafka Connect with tasks and workers.
- Confusing the connectors' roles, e.g. claiming MirrorSourceConnector translates consumer offsets (that's MirrorCheckpointConnector).
- Claiming heartbeats carry business data — they are tiny liveness records only.
- Thinking you must run all three connectors; each is independently enable-able.