skip to content

Remote Topic Naming and Replication Policies

Replication policies that decide whether mirrored topics get a source-cluster prefix, and how that prevents loops. Interviewers ask because the unprefixed identity policy is convenient and dangerous in bidirectional setups.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

When MirrorMaker 2 replicates a topic with the default settings, what does the replicated topic get named on the destination cluster, and why?

level: juniorimportance: must knowfreq 70%

answer

  1. DefaultReplicationPolicy
  2. source alias + separator
  3. us-west.orders
  4. provenance + no collisions
  5. enables cycle detection

basics

~20 s

By default MirrorMaker 2 prefixes the replicated topic with the source cluster's alias and a dot. A topic named 'orders' from cluster 'us-west' becomes 'us-west.orders' on the destination. This makes the topic's origin visible and prevents name clashes.

solid answer

~40 s

MirrorMaker 2 (MM2) uses DefaultReplicationPolicy, which renames replicated topics by prepending the source cluster alias plus a separator (a dot by default). So topic 'orders' on source cluster 'us-west' becomes 'us-west.orders' on the destination. This 'source-prefixed' naming serves two purposes: it makes the topic's provenance explicit, and it lets the same destination cluster hold replicas from multiple sources without collisions (e.g. 'us-west.orders' and 'eu-central.orders' coexist). Crucially, it also enables cycle detection in active/active setups: because the prefix encodes the path a record travelled, MM2 can refuse to replicate a topic back to a cluster it already came from. The prefix and separator are configurable; the separator is controlled by 'replication.policy.separator'.

go deeper

for a junior

Know that the default adds 'sourceAlias.' in front: orders becomes us-west.orders.

for a middle

Explain the two motivations (provenance, collision avoidance) and name DefaultReplicationPolicy and the separator config.

for a senior

Tie prefixing to cycle detection and multi-hop nested prefixes, and contrast with IdentityReplicationPolicy.

for a principal

Reason about when prefixing helps vs. hurts (DR failover ergonomics, downstream tooling, regex subscriptions) and design the cluster-aliasing scheme accordingly.

## What MirrorMaker 2 is MirrorMaker 2 (MM2) is Kafka's built-in tool for copying (replicating) records from one Kafka cluster to another — used for disaster recovery, geo-distribution, and migrations. It runs on Kafka Connect and copies topic data, consumer offsets, and ACLs. ## The naming problem If MM2 copied topic 'orders' from cluster A to cluster B and kept the name 'orders', two problems arise: (1) you couldn't tell, on B, whether 'orders' is B's own topic or a copy from A; (2) if both A and C replicate their 'orders' into B, the two would collide. ## DefaultReplicationPolicy and source-prefixing MM2's default behavior is governed by a class called `DefaultReplicationPolicy`. It renames each replicated topic by prepending the **source cluster alias** followed by a **separator**. A cluster alias is just the short name you assign each cluster in MM2 config (e.g. `us-west`, `eu-central`). The separator defaults to a dot (`.`). So: - Source cluster alias: `us-west` - Original topic: `orders` - Replicated topic on destination: `us-west.orders` This is called a **source-prefixed** or **remote** topic name. ## Why this matters 1. **Provenance**: anyone looking at `us-west.orders` immediately knows it's a replica originating from `us-west`. 2. **No collisions**: `us-west.orders` and `eu-central.orders` can both live on the same destination. 3. **Cycle detection**: the prefix records the replication path. In active/active topologies (A→B and B→A), MM2 inspects the prefix to detect that a topic already originated from a cluster and refuses to loop it back, preventing infinite replication. 4. **Consumer failover**: clients can subscribe with a regex/pattern (e.g. via `RemoteClusterUtils` / the offset-sync machinery) and follow a topic across a failover because the policy can map remote names back to the original. ## Configurability The separator is set with `replication.policy.separator` (default `.`). The whole renaming scheme is pluggable via `replication.policy.class`; the alternative built-in is `IdentityReplicationPolicy`, which keeps names unchanged. ## Edge cases - Multi-hop replication produces nested prefixes: A→B→C makes `B.us-west.orders` on C (the chain of hops is encoded). - Internal/system topics (offsets, config, MM2's own heartbeats/checkpoints) are filtered or handled specially so they aren't blindly re-prefixed in a way that breaks things.

  • What config controls the character between the prefix and the topic name?
    replication.policy.separator — it defaults to a dot ('.').
  • What happens to the name after two hops (A to B to C)?
    Prefixes nest: C ends up with B.us-west.orders, encoding the full replication path.

saying these in an interview costs you the question

  • Saying the replicated topic keeps the same name by default (that's IdentityReplicationPolicy, not the default)
  • Claiming the prefix is the destination alias (it's the SOURCE alias)
  • Thinking prefixing is purely cosmetic and serves no functional purpose (it enables cycle detection)

context

open as a page

What is IdentityReplicationPolicy, and what trade-off do you accept by using it instead of DefaultReplicationPolicy?

level: middleimportance: must knowfreq 60%

basics

~20 s

IdentityReplicationPolicy replicates topics keeping their original name (no source-cluster prefix), so 'orders' stays 'orders' on the destination. The trade-off: you lose automatic cycle detection and risk name collisions, so it's only safe for unidirectional (active/passive) replication.

open as a page

How does MM2 prevent infinite replication loops in an active/active topology, and how is this tied to topic naming?

level: seniorimportance: must knowfreq 50%

basics

~20 s

MM2 prevents loops by encoding a topic's origin in its name via the source prefix. Before replicating, it checks whether a topic already originated from the destination cluster (or matches the configured source-cluster prefix) and skips it, so records don't bounce back forever.

open as a page

How does replication.policy.separator work, what is its default, and why must it be consistent across an MM2 deployment?

level: middleimportance: should knowfreq 45%

basics

~20 s

replication.policy.separator is the character DefaultReplicationPolicy puts between the source cluster alias and the topic name. It defaults to a dot ('.'), so 'orders' becomes 'us-west.orders'. Changing it changes remote topic names; it must match everywhere or MM2 can't parse origins.

open as a page

Which topics does MM2 deliberately NOT replicate as ordinary data, and how does it identify them?

level: seniorimportance: should knowfreq 35%

basics

~20 s

MM2 skips internal and system topics: Kafka's own __consumer_offsets and transaction-state topics, and MM2's own heartbeats, checkpoints, and offset-sync topics. It identifies them via the ReplicationPolicy's isInternalTopic check and the topic filter, so they aren't replicated as user data.

open as a page

When would you implement a custom ReplicationPolicy, and what methods must it correctly implement to remain loop-safe and failover-correct?

level: principalimportance: should knowfreq 30%

basics

~20 s

You write a custom ReplicationPolicy when neither default prefixing nor identity naming fits — e.g. a bespoke naming scheme or org convention. It must correctly implement formatRemoteTopic plus the inverse parsing (topicSource, upstreamTopic) and isInternalTopic, or cycle detection and offset translation break.

open as a page