skip to content

A consumer group on the source has been committing offsets for hours, but RemoteClusterUtils.translateOffsets returns nothing for it on the target. What would you check?

level: middleimportance: should knowfreq 40%

answer

  1. emit.checkpoints.enabled + interval elapsed?
  2. groups.exclude defaults: console-consumer/connect/__
  3. is the Checkpoint connector task FAILING?
  4. offset syncs cover the committed offset?
  5. right alias + RENAMED topic in query?

basics

~20 s

Check that emit.checkpoints is enabled, the group isn't excluded by groups/groups.exclude, the <source>.checkpoints.internal topic exists with data, the offset-syncs topic has syncs covering the group's partitions, and that you're querying the right target alias and renamed topics.

solid answer

~40 s

First confirm checkpoints are even being produced: emit.checkpoints.enabled must be true and emit.checkpoints.interval.seconds elapsed. Verify the group is in scope — groups (default .*) minus groups.exclude (which by default excludes console-consumer-, connect-, __ groups); a default-excluded or unmatched group never gets checkpoints. Inspect the target's <source-alias>.checkpoints.internal topic: if it's empty, the MirrorCheckpointConnector task may be failing (check Connect logs). The connector also needs offset syncs in mm2-offset-syncs.<target>.internal that cover the group's committed offsets — if the topic only just started replicating, far-back committed offsets may lack syncs. Also confirm you passed the correct targetClusterAlias and that you're looking at the renamed replicated topic-partitions. Finally, the source group must actually have committed offsets (an assign-without-commit consumer produces nothing to checkpoint).

go deeper

for a junior

Know to check that checkpoints are enabled and the topic has data.

for a middle

Run the full checklist: scope/exclusions, connector health, sync coverage, correct alias and renamed topics.

for a senior

Diagnose offset-sync coverage gaps and connector ACL/task failures from logs.

for a principal

Build observability (heartbeats, checkpoint-lag metrics) so these gaps surface before a failover, not during one.

## Systematic checklist ### 1. Are checkpoints enabled and due? - `emit.checkpoints.enabled` must be **true** (it is by default, but a custom config can disable it). - At least one `emit.checkpoints.interval.seconds` (default 60) must have elapsed since the connector started for the first checkpoints to appear. ### 2. Is the group in scope? - The checkpoint connector only handles groups matching **`groups`** (default `.*`) and **not** matching **`groups.exclude`**. The default exclusion pattern filters out internal/utility groups like `console-consumer-.*`, `connect-.*`, and `__.*`. A group named to match an exclusion (or not matching `groups`) is silently skipped. - `refresh.groups.interval.seconds` controls how often new groups are discovered — a brand-new group may not be picked up until the next refresh. ### 3. Does the checkpoints topic have data? - On the **target** cluster, inspect **`<source-alias>.checkpoints.internal`**. If empty: - The **MirrorCheckpointConnector** task may be **failing** — check Kafka Connect worker logs / `GET /connectors/<name>/status`. - The connector may lack ACLs to read source `__consumer_offsets` or write the checkpoints topic. ### 4. Do offset syncs cover the group's offsets? - The checkpoint connector needs entries in **`mm2-offset-syncs.<target>.internal`** that map the group's committed **source** offsets. If replication started recently, the group's committed offset may predate any recorded sync, so translation yields nothing (or only translates newer positions). The historical fix here is the logarithmic-spread sync retention; very old offsets can still be untranslatable if no covering sync exists. ### 5. Did the source group actually commit? - A consumer using **manual assignment without committing** (or `enable.auto.commit=false` and never committing) leaves nothing in source `__consumer_offsets` to checkpoint. ### 6. Are you querying correctly? - `RemoteClusterUtils.translateOffsets(props, targetClusterAlias, groupId, timeout)` — wrong **targetClusterAlias** or wrong **groupId** returns empty. - Results key on **renamed** topic-partitions (e.g. `primary.orders-0`); make sure you're not looking for the un-prefixed name (unless using `IdentityReplicationPolicy`). ### 7. Replication policy mismatch - If the consumer queries with a different `replication.policy` understanding than MM2 used to write, topic-name expectations diverge. ## Quick triage order 1. Connector status/logs (is the task alive?) 2. Does the checkpoints topic have records at all? 3. Is the group excluded by `groups.exclude`? 4. Do offset syncs cover the group's committed offsets? 5. Right alias/group/renamed-topic in the query?

  • Which default exclusion might silently skip a group you expect to see checkpointed?
    groups.exclude defaults exclude console-consumer-.*, connect-.*, and __.* groups; a group matching those (e.g. a console-consumer test group) is never checkpointed.
  • Why might an old committed offset translate to nothing even though newer offsets translate fine?
    If no offset sync was retained covering that far-back source offset, the OffsetSyncStore can't map it; only positions covered by retained syncs translate.

saying these in an interview costs you the question

  • Looking for the checkpoints topic on the source cluster instead of the target.
  • Forgetting groups.exclude silently drops console-consumer/connect/internal groups.
  • Assuming a consumer that only assigns partitions (never commits) can be checkpointed.

context