What is Confluent Cluster Linking, and how does it differ at a high level from a Connect-based replicator?
answer
- Broker-native, no Connect cluster
- Read-only mirror topics
- Offset-preserving / byte-for-byte
- Link object lives on destination
- Confluent feature, not vanilla Kafka
basics
~20 sCluster Linking is a broker-built-in feature that copies topics from a source Kafka cluster to a destination cluster over a direct link, with no separate Connect cluster. The destination gets read-only mirror topics that keep the same message offsets.
solid answer
~40 sCluster Linking is a feature built directly into the Kafka broker (Confluent Platform / Confluent Cloud) that replicates data between two clusters. You create a 'link' on the destination cluster pointing at the source, then create mirror topics that pull data over that link. Unlike MirrorMaker 2 or a Connect-based replicator, there is no separate Connect cluster, no connector tasks, and no consumers/producers to operate — the destination brokers fetch from the source brokers natively, much like a follower replica. Crucially, mirror data is byte-for-byte and offset-preserving: offset 100 on the source is offset 100 on the destination. Mirror topics are read-only until promoted. This makes it ideal for disaster recovery, cluster migration, and geo-replication where offset consistency matters.
go deeper
Know it's broker-built-in replication with read-only, offset-preserving mirror topics — no Connect.
Explain the link/mirror-topic model and contrast with MM2 (offsets, operational overhead).
Discuss when to choose it (DR, migration, offset consistency) and its read-only-until-promoted lifecycle.
Position it in an overall replication architecture, including trade-offs vs MM2 and topology constraints (unidirectional, Confluent-only).
## The problem Kafka clusters often need a copy of their data in another cluster — for disaster recovery (DR), migrating to a new cluster, or serving readers in another region. Historically this was done with **MirrorMaker 2 (MM2)**, which runs on **Kafka Connect**: a separate cluster of worker processes that consume from the source and produce to the destination. That works, but it has moving parts (a Connect cluster, connector tasks, consumer groups) and — importantly — it does **not** preserve offsets, because re-producing assigns brand-new offsets on the destination. ## What Cluster Linking is **Cluster Linking** moves replication *into the broker itself*. Instead of an external process, the **destination cluster's brokers connect directly to the source cluster's brokers** and fetch data, behaving much like a follower replica fetches from a leader. The named connection between the two clusters is called a **cluster link**. Key concepts: - **Cluster link**: a configured, named connection object created on the **destination** cluster that points at the source (bootstrap servers, security config). One link can carry many topics. - **Mirror topic**: a topic on the destination that is bound to a source topic via the link. It is **read-only** — producers cannot write to it — and it continuously pulls new records from the source. - **Byte-for-byte / offset-preserving**: record at offset N on the source becomes the record at offset N on the destination. Partition count, ordering, and offsets all match. ## Why offset preservation matters Because offsets match exactly, a consumer that was at offset 5000 on the source can resume at offset 5000 on the destination and read the *same* message. This is what makes Cluster Linking suited for DR failover and live migration — consumer position is meaningful across clusters. MM2 cannot do this natively (it needs offset-translation via a checkpoint topic, which is approximate). ## How it differs from Connect-based replication | Aspect | Cluster Linking | MM2 / Connect replicator | |---|---|---| | Runtime | Built into the broker | Separate Connect cluster | | Offsets | Preserved exactly | Re-assigned (translation needed) | | Mirror topic | Read-only until promoted | Writable normal topic | | Operational overhead | Low (no extra cluster) | Higher (workers, tasks) | ## Edge cases / caveats - A mirror topic **cannot be written to** while it's mirroring; you must **promote** (or fail over) it to make it writable. - The source topic's config (partitions, etc.) is mirrored; you don't manage partitions independently on the destination. - Cluster Linking is a **Confluent** feature (Confluent Platform / Confluent Cloud), not part of vanilla Apache Kafka.
- On which cluster do you create the cluster link — source or destination?On the destination cluster. The destination's brokers initiate the connection and fetch (pull) from the source, similar to how a follower replica fetches from a leader.
- Can an application produce to a mirror topic?No. Mirror topics are read-only while mirroring. You must promote or fail over the mirror topic to make it a normal, writable topic first.
saying these in an interview costs you the question
- Saying Cluster Linking runs on Kafka Connect or needs a Connect cluster
- Claiming it re-assigns offsets like MM2 (it preserves them)
- Thinking the link is created on the source cluster
- Saying mirror topics are writable by default