skip to content

What is Confluent Cluster Linking, and how does it differ at a high level from a Connect-based replicator?

level: juniorimportance: must knowfreq 70%

answer

  1. Broker-native, no Connect cluster
  2. Read-only mirror topics
  3. Offset-preserving / byte-for-byte
  4. Link object lives on destination
  5. Confluent feature, not vanilla Kafka

basics

~20 s

Cluster Linking is a broker-built-in feature that copies topics from a source Kafka cluster to a destination cluster over a direct link, with no separate Connect cluster. The destination gets read-only mirror topics that keep the same message offsets.

solid answer

~40 s

Cluster Linking is a feature built directly into the Kafka broker (Confluent Platform / Confluent Cloud) that replicates data between two clusters. You create a 'link' on the destination cluster pointing at the source, then create mirror topics that pull data over that link. Unlike MirrorMaker 2 or a Connect-based replicator, there is no separate Connect cluster, no connector tasks, and no consumers/producers to operate — the destination brokers fetch from the source brokers natively, much like a follower replica. Crucially, mirror data is byte-for-byte and offset-preserving: offset 100 on the source is offset 100 on the destination. Mirror topics are read-only until promoted. This makes it ideal for disaster recovery, cluster migration, and geo-replication where offset consistency matters.

go deeper

for a junior

Know it's broker-built-in replication with read-only, offset-preserving mirror topics — no Connect.

for a middle

Explain the link/mirror-topic model and contrast with MM2 (offsets, operational overhead).

for a senior

Discuss when to choose it (DR, migration, offset consistency) and its read-only-until-promoted lifecycle.

for a principal

Position it in an overall replication architecture, including trade-offs vs MM2 and topology constraints (unidirectional, Confluent-only).

## The problem Kafka clusters often need a copy of their data in another cluster — for disaster recovery (DR), migrating to a new cluster, or serving readers in another region. Historically this was done with **MirrorMaker 2 (MM2)**, which runs on **Kafka Connect**: a separate cluster of worker processes that consume from the source and produce to the destination. That works, but it has moving parts (a Connect cluster, connector tasks, consumer groups) and — importantly — it does **not** preserve offsets, because re-producing assigns brand-new offsets on the destination. ## What Cluster Linking is **Cluster Linking** moves replication *into the broker itself*. Instead of an external process, the **destination cluster's brokers connect directly to the source cluster's brokers** and fetch data, behaving much like a follower replica fetches from a leader. The named connection between the two clusters is called a **cluster link**. Key concepts: - **Cluster link**: a configured, named connection object created on the **destination** cluster that points at the source (bootstrap servers, security config). One link can carry many topics. - **Mirror topic**: a topic on the destination that is bound to a source topic via the link. It is **read-only** — producers cannot write to it — and it continuously pulls new records from the source. - **Byte-for-byte / offset-preserving**: record at offset N on the source becomes the record at offset N on the destination. Partition count, ordering, and offsets all match. ## Why offset preservation matters Because offsets match exactly, a consumer that was at offset 5000 on the source can resume at offset 5000 on the destination and read the *same* message. This is what makes Cluster Linking suited for DR failover and live migration — consumer position is meaningful across clusters. MM2 cannot do this natively (it needs offset-translation via a checkpoint topic, which is approximate). ## How it differs from Connect-based replication | Aspect | Cluster Linking | MM2 / Connect replicator | |---|---|---| | Runtime | Built into the broker | Separate Connect cluster | | Offsets | Preserved exactly | Re-assigned (translation needed) | | Mirror topic | Read-only until promoted | Writable normal topic | | Operational overhead | Low (no extra cluster) | Higher (workers, tasks) | ## Edge cases / caveats - A mirror topic **cannot be written to** while it's mirroring; you must **promote** (or fail over) it to make it writable. - The source topic's config (partitions, etc.) is mirrored; you don't manage partitions independently on the destination. - Cluster Linking is a **Confluent** feature (Confluent Platform / Confluent Cloud), not part of vanilla Apache Kafka.

  • On which cluster do you create the cluster link — source or destination?
    On the destination cluster. The destination's brokers initiate the connection and fetch (pull) from the source, similar to how a follower replica fetches from a leader.
  • Can an application produce to a mirror topic?
    No. Mirror topics are read-only while mirroring. You must promote or fail over the mirror topic to make it a normal, writable topic first.

saying these in an interview costs you the question

  • Saying Cluster Linking runs on Kafka Connect or needs a Connect cluster
  • Claiming it re-assigns offsets like MM2 (it preserves them)
  • Thinking the link is created on the source cluster
  • Saying mirror topics are writable by default

context