skip to content

MirrorMaker 2 Replication Flows

How MirrorMaker 2 runs on Connect with source, checkpoint and heartbeat connectors to copy topics between clusters. It is the default answer to 'how do you replicate Kafka', so interviewers expect its moving parts.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What are the three core connectors that make up MirrorMaker 2, and what does each one do?

level: juniorimportance: must knowfreq 70%

answer

  1. Source = data, Checkpoint = offsets, Heartbeat = liveness
  2. Built on Kafka Connect (KIP-382)
  3. One flow = source->target = a set of connectors
  4. Heartbeats measure end-to-end lag
  5. Connectors split into tasks for parallelism

basics

~10 s

MirrorMaker 2 has three connectors: MirrorSourceConnector copies topic data from the source cluster to the target; MirrorCheckpointConnector copies consumer-group offset progress; MirrorHeartbeatConnector emits periodic heartbeats to confirm the replication path is alive.

solid answer

~40 s

MM2 runs as a set of Kafka Connect connectors. MirrorSourceConnector is the workhorse: it consumes records from source-cluster topics and produces them onto the target cluster, also propagating ACLs and (optionally) topic configs. MirrorCheckpointConnector reads committed consumer-group offsets on the source, translates them to equivalent target offsets, and writes them to a checkpoints topic so consumers can fail over. MirrorHeartbeatConnector periodically produces small heartbeat records into a heartbeats topic on both clusters; these prove the end-to-end path works and provide a clock for measuring replication lag. Each connector spawns multiple tasks for parallelism — MirrorSourceConnector tasks are partitioned across the source topic-partitions it owns. Together they form one directional 'flow' (source->target); a bidirectional setup runs two flows.

go deeper

for a junior

Memorize the three connector names and one-line role for each: source=data, checkpoint=offsets, heartbeat=liveness.

for a middle

Know each connector splits into tasks, that they run on Kafka Connect, and which config flags enable/disable each.

for a senior

Explain how connectors map to a directional flow, when you'd disable checkpoints, and how heartbeats give an end-to-end latency signal.

for a principal

Reason about running MM2 on a dedicated vs shared Connect cluster, task-level parallelism limits, and the operational trade-offs of each connector for a DR vs active-active design.

MirrorMaker 2 (MM2) is Kafka's cross-cluster replication tool, introduced in KIP-382. Unlike the legacy MirrorMaker 1 (a standalone consumer+producer loop), MM2 is built **on top of Kafka Connect**, so it inherits Connect's distributed runtime, REST API, offset/config/status topics, scaling via tasks, and fault tolerance. A **connector** in Kafka Connect is a plugin that defines how to move data; it splits its work into **tasks** that run on **workers** (JVM processes). MM2 ships three connector classes, each with a distinct job: 1. **MirrorSourceConnector** — the data-plane connector. It discovers which source-cluster topics match the configured allow/deny filters, then consumes their records and produces them to correspondingly-named topics on the **target** cluster (the remote topic name is prefixed, e.g. `us-west.orders`). It also creates target topics if missing, replicates ACLs, and (when `sync.topic.configs.enabled=true`) keeps target topic configs in sync. It parallelizes by assigning source topic-partitions across its tasks. 2. **MirrorCheckpointConnector** — replicates **consumer progress**, not data. It periodically reads committed offsets of source-cluster consumer groups (via `__consumer_offsets`), maps each source offset to the equivalent offset of the mirrored record on the target, and emits **checkpoint** records to a `<source>.checkpoints.internal` topic on the target. This lets a consumer that fails over to the target resume near where it left off. (The exact offset-translation math is owned by a sibling topic and out of scope here.) 3. **MirrorHeartbeatConnector** — a liveness/observability connector. It produces tiny **heartbeat** records (a timestamp and the source/target alias) into a `heartbeats` topic at a fixed interval (`emit.heartbeats.interval.seconds`, default 1s). Because heartbeats are themselves mirrored, you can observe them arriving on the target and measure end-to-end replication latency, and confirm the path is healthy even when no business data is flowing. **Key edge cases:** Each connector is independently enable-able (`MirrorSourceConnector` via `source->target.enabled`, plus `emit.checkpoints.enabled` and `emit.heartbeats.enabled` flags). Heartbeats are cheap and almost always on. You can run the source connector alone if you only need data and don't care about consumer failover. In a dedicated MM2 cluster (the `connect-mirror-maker.sh` driver), all three are auto-configured from a single `mm2.properties`; in a generic Connect cluster you POST each connector's config to the REST API yourself.

  • Why is MM2 built on Kafka Connect instead of being a standalone loop like MirrorMaker 1?
    To inherit Connect's distributed scaling (tasks/workers), REST management, automatic offset/config/status persistence in internal topics, rebalancing, and fault tolerance — MM1 was a single-process consumer+producer with no built-in HA, scaling, or offset translation.
  • If you only need to replicate data for disaster recovery and don't care about per-consumer failover, which connectors can you skip?
    You can disable the checkpoint connector (set emit.checkpoints.enabled=false). Keep MirrorSourceConnector for data; heartbeats are cheap and worth keeping for health monitoring, but are also optional.

saying these in an interview costs you the question

  • Saying MM2 is a single process/loop like MirrorMaker 1 — it runs on Kafka Connect with tasks and workers.
  • Confusing the connectors' roles, e.g. claiming MirrorSourceConnector translates consumer offsets (that's MirrorCheckpointConnector).
  • Claiming heartbeats carry business data — they are tiny liveness records only.
  • Thinking you must run all three connectors; each is independently enable-able.

context

open as a page

How do you control which topics and consumer groups MM2 replicates, using allow/deny filters?

level: middleimportance: must knowfreq 60%

basics

~10 s

MM2 uses topics/topics.exclude and groups/groups.exclude regex lists (per flow). The 'topics'/'groups' allowlist says what to replicate; the 'exclude' denylist removes matches. By default internal and MM2 bookkeeping topics/groups are excluded.

open as a page

What does sync.topic.configs.enabled do, and which topic configurations does MM2 keep in sync on the target?

level: middleimportance: should knowfreq 45%

basics

~10 s

When sync.topic.configs.enabled=true (the default), MM2 periodically copies the source topic's configuration (like cleanup.policy, retention, etc.) onto the mirrored target topic so they stay consistent. A property filter controls which config keys are synced.

open as a page

How does MirrorSourceConnector achieve parallelism, and how do tasks.max and source partitions affect replication throughput?

level: seniorimportance: should knowfreq 40%

basics

~20 s

MirrorSourceConnector divides the source topic-partitions it must replicate across tasks. Parallelism is bounded by tasks.max and by the number of source partitions — you never get more useful tasks than partitions, since each partition is handled by one task.

open as a page

How would you design an active-active bidirectional MM2 topology, and how do replication flows avoid infinite loops?

level: principalimportance: should knowfreq 35%

basics

~20 s

You define two flows (A->B and B->A), each running its own set of MM2 connectors. Loops are avoided because mirrored topics are prefixed with the source cluster alias, and MM2's default filters exclude already-remote topics, so B never re-mirrors A's data back to A.

open as a page