How does Cluster Linking keep consumer group offsets and ACLs in sync across clusters, and why does it matter for failover?
answer
- consumer.offset.sync.enable + group.filters
- Offset-preserving ⇒ no translation (vs MM2 checkpoints)
- acl.sync.enable + acl.filters (JSON)
- Periodic, not real-time
- Don't run a group on both clusters at once
basics
~20 sThe link can be configured to periodically copy committed consumer-group offsets and ACLs from the source to the destination. Because offsets are byte-for-byte identical across clusters, the copied commit positions are directly valid, so consumers can resume at the same place after failover.
solid answer
~50 sA cluster link supports two optional sync features. **Consumer offset sync** periodically copies committed offsets for selected consumer groups from the source's __consumer_offsets to the destination, controlled by configs like `consumer.offset.sync.enable`, a `consumer.offset.group.filters` JSON allow/deny list, and `consumer.offset.sync.ms` for cadence. Because Cluster Linking preserves offsets byte-for-byte, a copied commit (group X at offset 5000) means exactly the same record on the destination — no offset translation needed, unlike MM2. **ACL sync** copies Kafka ACLs from source to destination via `acl.sync.enable` plus an `acl.filters` JSON spec, so authorization rules survive the move. Together they mean that after promote/failover, your consumers and their permissions are already configured on the destination, so applications resume at the correct position with the right access. This is what makes Cluster Linking a low-friction DR/migration tool: the data, the read positions, and the security policy all travel over one link.
go deeper
Know the link can copy consumer offsets and ACLs, not just data.
Name the enable flags and explain why synced offsets are directly valid (offset preservation).
Discuss filters, periodicity/lag, the don't-run-on-both-clusters caveat, and the MM2-translation contrast.
Architect DR so positions+ACLs+data all flow over one link; reason about RPO on positions and bidirectional/active-active pitfalls.
## Why sync matters Replicating *data* alone isn't enough for a real failover. When apps cut over to the destination cluster they need two more things: (1) to **resume reading at the right position** (their committed consumer offsets), and (2) to still be **authorized** (their ACLs). Cluster Linking can carry both over the same link. ## Consumer offset sync Kafka consumers commit their progress as offsets stored in the internal **`__consumer_offsets`** topic, keyed by (group, topic, partition). Cluster Linking can periodically copy these committed offsets from source to destination. Relevant link configs (set on the link or per mirror): - **`consumer.offset.sync.enable`** — turn the feature on. - **`consumer.offset.group.filters`** — a JSON include/exclude spec selecting *which* groups to sync (you rarely want all of them). - **`consumer.offset.sync.ms`** — how often to sync. ### Why no offset translation is needed The key insight: Cluster Linking is **offset-preserving**. Offset 5000 on the source *is* offset 5000 on the destination — same record. So a copied commit of 'group `payments` at offset 5000' is immediately, exactly correct on the destination. Contrast **MirrorMaker 2**, which re-produces records and therefore assigns *different* offsets; MM2 must maintain a **checkpoint/offset-translation** mapping that is only approximate. Cluster Linking sidesteps that whole problem. ### Important caveat If a consumer is *also actively running and committing on the destination*, syncing from the source can move its position backward or forward unexpectedly. Best practice: keep consumers running on **one** cluster at a time, and let offset sync prime the destination so they resume cleanly after cutover. There can also be a small lag between the latest source commit and what's been synced. ## ACL sync Kafka **ACLs** (access control lists) define which principals may perform which operations on which resources. After failover, the destination must enforce the same rules. Cluster Linking can copy ACLs: - **`acl.sync.enable`** — turn it on. - **`acl.filters`** — a JSON spec selecting which ACLs to mirror. This keeps authorization consistent so apps that were allowed to read/write on the source remain allowed on the destination. ## Putting it together for failover With data + offset sync + ACL sync flowing over the link, a failover/promote yields a destination where: topics exist with identical offsets, consumer groups have valid committed positions, and the security policy is in place. Apps repoint their bootstrap servers and continue with minimal manual setup — the essence of Cluster Linking's DR value. ## Edge cases - Offset sync is **periodic**, not real-time, so a tiny RPO on *positions* can exist even when data is fully replicated. - Filters are JSON and must be valid or the sync won't behave as expected. - Syncing offsets for groups active on both clusters simultaneously is an anti-pattern.
- Why doesn't Cluster Linking need MM2-style offset translation?Because it preserves offsets byte-for-byte. A committed offset on the source points to the identical record at the same offset on the destination, so the copied value is directly usable — no checkpoint/translation mapping required.
- What goes wrong if you sync offsets for a group that's actively committing on the destination?The synced source offsets can overwrite the destination's live progress, moving the group's position backward or forward unexpectedly and causing reprocessing or skipped records. Run a group on one cluster at a time.
saying these in an interview costs you the question
- Saying Cluster Linking needs offset translation like MM2
- Claiming offset/ACL sync is real-time/transactional (it's periodic)
- Forgetting filters select which groups/ACLs are synced
- Assuming all consumer groups are synced by default