skip to content

Explain how MirrorMaker 2 enables consumer failover across clusters, including offset translation and the role of MirrorCheckpointConnector.

level: seniorimportance: must knowfreq 45%

answer

  1. offsets are per-cluster — not portable
  2. 3 connectors: Source, Checkpoint, Heartbeat
  3. checkpoints topic = upstream→downstream offset map
  4. RemoteClusterUtils.translateOffsets / sync.group.offsets.enabled
  5. translation approximate → replay → need idempotency
  6. Cluster Linking preserves offsets (no translation)

basics

~20 s

Source and target clusters have different offsets for the same record, so you can't reuse raw offsets after failover. MM2's MirrorCheckpointConnector tracks the mapping and writes checkpoints; consumers use RemoteClusterUtils (or automatic sync to __consumer_offsets) to resume at the equivalent position on the target.

solid answer

~40 s

When MirrorMaker 2 replicates a topic, each record gets a brand-new offset on the target cluster, because offsets are per-partition append positions and the target's log started independently. So a consumer that was at offset 5000 on the source cannot simply seek to 5000 on the target. MM2 solves this with three connectors: MirrorSourceConnector copies records, MirrorCheckpointConnector reads the source's __consumer_offsets and emits checkpoints (source-offset → target-offset mappings per consumer group) into a checkpoints topic, and MirrorHeartbeatConnector measures lag/liveness. On failover, a consumer calls RemoteClusterUtils.translateOffsets() to convert its last committed source offsets to target offsets, or you enable sync.group.offsets.enabled=true so MM2 writes translated offsets directly into the target's __consumer_offsets. Translation is approximate (mapped to the nearest known checkpoint), so consumers may replay a few records — designs must be idempotent.

go deeper

for a junior

Know that mirrored clusters have different offsets and MM2 helps consumers find their place after failover.

for a middle

Name the three MM2 connectors and that the checkpoint connector maps source offsets to target offsets.

for a senior

Explain RemoteClusterUtils vs sync.group.offsets, the at-least-once replay implication, and prefixing.

for a principal

Design failover topology, choose MM2 vs offset-preserving Cluster Linking, and mandate idempotent consumers.

## The core problem: offsets aren't portable A Kafka **offset** is just the sequential position of a record within a single partition's log on one cluster. When MirrorMaker 2 (MM2) copies partition `orders-0` from cluster A to cluster B, cluster B writes those records into its **own** log starting from its own next offset. Because of independent log start, prior topic creation, retention deletions, and MM2 batching, **record R that is at offset 5000 on A might be at offset 13278 on B**. Therefore a consumer that committed offset 5000 on A cannot blindly `seek(5000)` on B after failover — it would land at the wrong record. ## MM2 connector trio (runs on Kafka Connect) 1. **MirrorSourceConnector** — the data plane. Consumes from remote topics and produces them to the target, by default with a prefix (e.g. `A.orders` on B under `DefaultReplicationPolicy`; configurable via `replication.policy` — IdentityReplicationPolicy keeps original names). 2. **MirrorCheckpointConnector** — the offset-translation plane. It reads the **source** cluster's `__consumer_offsets` to learn where each consumer group is, correlates that with the source→target offset mapping it has observed, and emits **checkpoint** records `{group, topic, partition, upstreamOffset, downstreamOffset}` into a `<source>.checkpoints.internal` topic on the target. 3. **MirrorHeartbeatConnector** — emits heartbeats to a `heartbeats` topic to measure end-to-end replication lag and prove the link is alive. ## How a consumer resumes on the target Two mechanisms: - **Pull/manual**: the application calls `RemoteClusterUtils.translateOffsets(props, targetCluster, group, timeout)`. This reads the checkpoints topic and returns the **target** offsets corresponding to the group's last committed **source** offsets. The app `seek()`s there before consuming. - **Automatic**: set `sync.group.offsets.enabled=true` (with `sync.group.offsets.interval.seconds`). MM2 periodically writes translated offsets straight into the **target's** `__consumer_offsets` for the mirrored group, so on failover the group just starts and finds committed offsets already present. ## Approximation and edge cases - Translation maps to the **nearest checkpoint at or before** the source offset, so a consumer may **replay** some records after failover. Downstream processing must be **idempotent** / tolerant of duplicates (at-least-once across the failover boundary). - If a consumer group was inactive or its source offsets weren't committed, there may be no checkpoint to translate. - `emit.checkpoints.interval.seconds` controls checkpoint freshness — too coarse increases potential replay. - Topic-name prefixing matters: with the default policy the target topic is `A.orders`, so translation and client configs must use the prefixed name (or use IdentityReplicationPolicy for active-passive simplicity, watching for loop hazards in active-active). - Cluster Linking (Confluent) offers an alternative where offsets are **byte-for-byte preserved** (offset-preserving), removing the need for translation — a key contrast to call out. ## Summary MM2 doesn't make offsets identical; it makes them **translatable**. MirrorCheckpointConnector is the bookkeeper, RemoteClusterUtils/`sync.group.offsets` is the applier, and idempotent consumers absorb the small replay.

  • Why can't a consumer reuse its source offset directly on the target cluster?
    Offsets are per-partition log positions on a specific cluster. The target's log started independently, so the same record sits at a different offset; raw offset reuse would seek to the wrong place.
  • What guarantee do you get after offset-translated failover, and what design implication follows?
    At-least-once: translation maps to the nearest checkpoint at-or-before the source offset, so consumers may replay a few records. Consumers must be idempotent or deduplicate.
  • How does Cluster Linking differ from MM2 on offsets?
    Cluster Linking is offset-preserving — mirrored partitions keep identical offsets — so consumers can resume without offset translation, simplifying failover.

saying these in an interview costs you the question

  • Saying offsets are identical across mirrored clusters by default with MM2.
  • Forgetting MirrorCheckpointConnector and how checkpoints are produced.
  • Promising exactly-once across failover (it's at-least-once with replay).
  • Ignoring topic prefixing under DefaultReplicationPolicy.

context