skip to content

Walk through how you'd fail a consumer group over to a target cluster using RemoteClusterUtils.translateOffsets, and when you'd prefer sync.group.offsets.enabled instead.

level: seniorimportance: should knowfreq 45%

answer

  1. stop source group -> translate -> alterConsumerGroupOffsets -> start on target
  2. translateOffsets returns target TP -> OffsetAndMetadata
  3. sync.group.offsets.enabled = hands-off DR
  4. won't overwrite a live target group
  5. subscribe to RENAMED replicated topics

basics

~20 s

Call RemoteClusterUtils.translateOffsets to get translated target offsets for the group, commit them to the group on the target, then start consumers there pointed at the replicated topics. Prefer sync.group.offsets.enabled when you want MM2 to keep the target offsets continuously synced so failover needs no extra tooling.

solid answer

~40 s

Manual path: with the source group stopped, call RemoteClusterUtils.translateOffsets(props, targetAlias, groupId, timeout) which reads <source>.checkpoints.internal and returns Map<TopicPartition, OffsetAndMetadata> of translated target offsets. You then use AdminClient.alterConsumerGroupOffsets (or a KafkaConsumer commit while the group is empty) to write those offsets into the target __consumer_offsets, remap topic names to the replicated names (e.g. orders -> source.orders), and start consumers on the target. Automated path: set sync.group.offsets.enabled=true so the MirrorCheckpointConnector continuously commits translated offsets into the target group (every sync.group.offsets.interval.seconds), skipping groups actively consuming on the target. Prefer the automated path for warm-standby/DR where you want minimal manual steps at failover time; prefer manual translateOffsets when you need precise control over timing, when the group is active on both sides, or in tooling/scripts that orchestrate a controlled cutover.

go deeper

for a junior

Know that a helper translates offsets and you commit them to the target group before starting consumers.

for a middle

Sequence the manual steps: stop, translate, alter offsets, start; know sync.group.offsets.enabled automates it.

for a senior

Compare manual vs automated, handle topic renaming and inactive-group requirement, reason about staleness.

for a principal

Design DR runbooks/automation, decide per-topology (active-active vs warm standby), and handle fail-back via reverse-flow checkpoints.

## Goal Move a consumer group from the **source** cluster (failed/failing) to the **target** cluster so it resumes near the right position on the replicated topics, with minimal duplicates and no skipped data. ## Manual failover with RemoteClusterUtils 1. **Stop / fence the group on the source** so its committed offsets stop moving. Translation is only meaningful for a quiesced position. 2. **Translate**: call ``` Map<TopicPartition, OffsetAndMetadata> offsets = RemoteClusterUtils.translateOffsets(props, targetClusterAlias, consumerGroupId, timeout); ``` This reads the latest **checkpoints** from `<source>.checkpoints.internal` on the target and returns translated **target** offsets keyed by **target** topic-partitions (already renamed per the replication policy, e.g. `primary.orders-0`). 3. **Apply** the offsets to the target group with `AdminClient.alterConsumerGroupOffsets(groupId, offsets)` while the group is empty/inactive on the target. 4. **Start consumers** on the target, subscribed to the replicated topic names. They begin from the committed translated offsets. ## Automated path: sync.group.offsets.enabled Setting `sync.group.offsets.enabled=true` (with `emit.checkpoints.enabled=true`) makes the **MirrorCheckpointConnector** itself commit the translated offsets into the target `__consumer_offsets` every `sync.group.offsets.interval.seconds` (default 60). At failover you simply **start the consumers on the target** — their group already has up-to-date translated offsets. The connector **will not overwrite** the offset of a group that is **actively consuming** on the target, so it's safe in active/active and avoids clobbering live progress. ## When to choose which - **Prefer `sync.group.offsets.enabled`** for warm-standby DR and for many groups, where you want hands-off failover and acceptable ~interval staleness. - **Prefer manual `translateOffsets`** when you need a precisely-timed, scripted cutover; when you only fail over selected groups; or when you want to inspect/validate the translated offsets before committing. It's also the natural API inside custom failover automation. ## Pitfalls - **Topic renaming**: translated keys use the **replicated** topic names. Consumers on the target must subscribe to those names (or use IdentityReplicationPolicy if you replicate without prefixes). - **Stale checkpoints / lag**: the translated position is as fresh as the last emitted checkpoint (`emit.checkpoints.interval.seconds`). Expect a small replay window. - **Group must be inactive when you alterConsumerGroupOffsets** — the Admin call is rejected for a group with active members. - **Active-active fail-back**: when failing back, the reverse flow's checkpoints handle the return trip; ensure both flows run so offsets exist in both directions.

  • Why must the consumer group be inactive on the target before you call alterConsumerGroupOffsets?
    AdminClient rejects alterConsumerGroupOffsets for a group with active members to prevent corrupting a live group's committed positions; you commit only while the group is empty/stopped.
  • What does sync.group.offsets.enabled do when a group is already actively consuming on the target?
    It skips overwriting that group's offsets, so live consumption on the target is not clobbered by translated source offsets — important for active-active topologies.

saying these in an interview costs you the question

  • Forgetting that translated offsets key on the renamed (prefixed) target topics.
  • Trying to alterConsumerGroupOffsets while the target group has active members.
  • Assuming sync.group.offsets.enabled blindly overwrites live target offsets (it skips active groups).

context