Why does cross-cluster replication with MirrorMaker 2 not provide exactly-once semantics, and what are the consequences for consumers on the target cluster?
answer
- MM2 = Connect consume A, produce B
- EOS is cluster-local → can't span A+B
- duplicates on restart/rebalance
- offsets differ; offset-sync + approximate translation
- KIP-656 narrows but ≠ end-to-end EOS
basics
~20 sMirrorMaker 2 is a Kafka Connect consume-from-source, produce-to-target pipeline across two clusters. Because the source consume and target produce span different clusters with no shared transaction, it's at-least-once: failures cause duplicate records on the target. Offsets and producer state don't carry over, so target consumers must tolerate duplicates and remapped offsets.
solid answer
~50 sMirrorMaker 2 (MM2) replicates by running a Kafka Connect MirrorSourceConnector that consumes from the source cluster and produces to the target cluster. EOS is cluster-local, but here the read and the write are on two separate clusters, so the transaction coordinator can't bind them — replication is at-least-once. On connector restart or rebalance, records since the last committed source offset are re-produced, creating duplicates on the target. Producer PID/epoch and the source's transaction markers don't transfer, so a transactional read-committed view doesn't replicate cleanly, and offsets differ between clusters (MM2 maintains an offset-sync topic and translates consumer-group offsets approximately, not exactly). Consequences: target consumers must be idempotent or dedupe, must not assume identical offsets, and failover (consumer migration) can replay or skip near the boundary. Newer work (KIP-656 / exactly-once support for MirrorMaker via Connect EOS) narrows duplicates but full cross-cluster exactly-once end-to-end is still not a guarantee.
go deeper
Know MM2 copies data between clusters and may produce duplicates (at-least-once).
Explain that it's consume-on-A/produce-on-B with no shared transaction, so restarts duplicate; offsets differ across clusters.
Detail offset-sync translation being approximate, producer-state not transferring, and the consumer-side implications for DR failover.
Design DR/active-active topologies around at-least-once replication: idempotent consumers, offset translation strategy, and explicit replay-vs-skip failover semantics.
## What MM2 is **MirrorMaker 2 (MM2)** is Kafka's cross-cluster replication tool built on **Kafka Connect**. Its core connector, **MirrorSourceConnector**, is effectively a consumer on the **source** cluster and a producer to the **target** cluster, copying topic records. Companion connectors replicate consumer-group offsets (**MirrorCheckpointConnector**) and cluster topology (**MirrorHeartbeatConnector**). ## Why it isn't exactly-once Kafka's EOS guarantee is **cluster-local**: a transaction can atomically bind produced records and the consumer offset commit *only when they're in the same cluster*. In MM2, the **consume happens on cluster A** and the **produce happens on cluster B** — two independent clusters, two independent transaction coordinators. There is no single transaction spanning both. So MM2 inherits the same **two-non-atomic-steps** problem as any sink: - It produces records to the target, then records progress (source offset) in a Connect offsets topic. - A crash/rebalance between those steps causes **re-consumption from the last committed source offset** and **re-production** to the target → **duplicates**. Hence replication is **at-least-once** by default. ## Offset and producer-state caveats - **Offsets don't match across clusters.** The same logical record gets a *different* offset on the target. MM2 keeps an **offset-sync topic** and the checkpoint connector **translates** committed consumer-group offsets from source to target, but translation is **approximate** (best-effort, can land slightly before the exact record), so a migrated consumer may **replay a few records or, near edges, skip**. - **Producer idempotence/epoch doesn't carry over.** The source's PID, sequence numbers, and transaction commit/abort markers are source-cluster constructs; MM2 re-produces with its *own* producer, so transactional boundaries and `read_committed` semantics aren't faithfully reproduced on the target. - **Transactions don't replicate transparently.** Aborted records and exactly-once read semantics on the source don't translate into the same guarantee on the target. ## Consequences for target consumers 1. **Tolerate duplicates** — be idempotent or dedupe by a business key / message id. 2. **Don't assume identical offsets** — never hardcode cross-cluster offset equality; use offset translation or business keys for positioning. 3. **Plan for failover replay/skip** — during disaster-recovery consumer migration, expect a small replay (preferred) and design for it; verify whether your config risks skipping. 4. **Ordering is per-partition** and preserved within a partition, but global ordering/transaction atomicity is not. ## Improvements (still not full EOS) - **KIP-656** and Kafka Connect's exactly-once source support (KIP-618) allow MM2's source connector to use **Connect exactly-once**, reducing duplicates by making the produce + Connect-offset-commit atomic *on the target*. This tightens the guarantee but does **not** deliver true end-to-end cross-cluster exactly-once across arbitrary failover, and requires enabling/configuring Connect EOS. ## One-liner MM2 is a cross-cluster consume-produce; because EOS is cluster-local, replication is at-least-once with approximate offset translation — target consumers must be idempotent and offset-agnostic.
- A consumer group fails over from the source to the target cluster during DR. Why might it replay some messages?MM2 translates the group's committed offsets from source to target using the offset-sync topic, but the translation is approximate and conservative — it typically maps to a position at or slightly before the exact record, so the resumed consumer re-reads a few already-processed records. Idempotent processing absorbs the replay.
- Does enabling Kafka transactions on the source cluster make MM2 replication exactly-once?No. Source transactions are cluster-local; MM2 re-produces with its own producer on the target, so source PID/epoch and commit markers don't transfer. KIP-656/Connect EOS can reduce duplicates on the target side, but cross-cluster end-to-end exactly-once across failover is still not guaranteed.
saying these in an interview costs you the question
- Claiming MM2 replicates with exactly-once or that source offsets are identical on the target
- Assuming transactional/read-committed semantics replicate transparently across clusters
- Designing target consumers that hardcode cross-cluster offset equality
- Believing KIP-656 delivers full end-to-end cross-cluster EOS