skip to content

A team runs IdentityReplicationPolicy so topic names stay identical across two clusters. What naming and double-replication hazards does this introduce, and how do you mitigate them?

level: seniorimportance: should knowfreq 45%

answer

  1. Identity = same name, no provenance
  2. loop + collision + divergent writes
  3. one-way only; topics.exclude allow/deny
  4. great for migration/DR, bad for active/active
  5. reserve Default for steady-state mesh

basics

~20 s

Without the alias prefix, replicas share the original name, so producers can write to the 'same' topic on both clusters and replication has no way to tell origin from replica. You must use one-way flows or strict topics.exclude filters to avoid loops and collisions.

solid answer

~40 s

IdentityReplicationPolicy keeps the source name on the target ('orders' -> 'orders'), which is desirable for migrations and consumer transparency but removes the provenance prefix that DefaultReplicationPolicy relies on for loop prevention. Hazards: (1) double-replication / loops in bidirectional setups, since neither side can distinguish an original from a replica by name; (2) naming collisions when both clusters legitimately have a local topic of the same name, so the replica silently merges into local production traffic; (3) producers writing to the same logical name on both clusters create divergent histories that replication then tries to reconcile. Mitigations: run replication strictly one-directional; use precise `topics` allow-lists and `topics.exclude` deny-lists so a topic is only ever sourced from one side; segregate by topic namespace conventions; and reserve IdentityReplicationPolicy for migration/DR cutovers rather than steady-state active/active meshes.

go deeper

for a junior

Know Identity keeps the same topic name and that this can cause loops if used both ways.

for a middle

Explain the loss of provenance and why one-way flows + topic filters are needed.

for a senior

Enumerate loop, collision, and divergent-write hazards and the allow/deny-list and namespace mitigations.

for a principal

Decide policy per topology, define namespace conventions, and architect a safe migration/cutover that retires cross-replication.

**Background.** MirrorMaker 2 (MM2) copies Kafka topics between clusters. Its `ReplicationPolicy` decides what a replicated topic is named on the target. `DefaultReplicationPolicy` prefixes the source alias (`A.orders`); `IdentityReplicationPolicy` leaves the name unchanged (`orders` stays `orders`). **Why teams choose IdentityReplicationPolicy.** Consumers and producers don't have to learn a prefixed name; this makes a *migration* (move all clients from cluster A to cluster B) or a *disaster-recovery cutover* transparent — clients keep using `orders`. That transparency is the whole appeal. **Hazard 1 — loops / double replication.** DefaultReplicationPolicy prevents cycles because the prefix records provenance. Identity erases that. In a bidirectional A<->B flow, MM2 on B sees a topic named `orders` (which is actually A's replica) and, having no prefix to mark it as remote, can replicate it back to A. Records bounce, partition leaders churn, and you get a replication storm that saturates inter-cluster bandwidth. **Hazard 2 — naming collisions.** Suppose B *already* has its own local `orders` topic. With Identity, A's `orders` replica lands on the same name and merges into B's local stream. There is no namespace separation, so you can't tell replicated records from locally-produced ones, and consumer semantics become ambiguous. With DefaultReplicationPolicy this can't happen because the replica is `A.orders`, distinct from B's local `orders`. **Hazard 3 — divergent writes.** If applications produce to `orders` on *both* clusters (true active/active), the two logs have independent partition offsets and orderings. Replication cannot merge two authoritative logs into one consistent ordering; you get duplicated or interleaved records and broken offset semantics. **Mitigations.** - Use Identity only for *single-direction* flows (migration, DR replica that is read-mostly until cutover). - Lock down the topic set with `topics` (regex allow-list) and `topics.exclude` (deny-list) so any given topic is replicated from exactly one source. Explicitly exclude already-replicated and internal topics (`.*\.internal`, `heartbeats`, `mm2-offset-syncs`, `checkpoints`). - Adopt topic *namespace conventions* (team or region prefixes baked into the original name) so collisions are impossible by construction. - Prefer DefaultReplicationPolicy for steady-state active/active meshes; reserve Identity for the cutover window, then retire the cross-replication. **Edge cases.** During a migration you often run Identity one-way A->B while clients drain; if you forget to disable the reverse connector you reintroduce the loop. Mixed policies across connectors in the same cluster make remote-topic parsing inconsistent and should be avoided.

  • When is IdentityReplicationPolicy actually the right choice?
    For one-directional flows where client transparency matters: cluster migrations and DR replicas that stay read-mostly until a cutover. The replica keeps the original topic name so consumers/producers don't need reconfiguration, and the absence of loop protection is fine because there is no reverse flow.
  • How do you stop a topic from being replicated from both sides at once under Identity?
    Use precise topics allow-lists and topics.exclude deny-lists so each topic has exactly one source connector, and exclude already-replicated and internal topics by regex. Operationally, run only one direction and disable the reverse connector during migration.

saying these in an interview costs you the question

  • Recommending IdentityReplicationPolicy for steady-state active/active without filters.
  • Believing the broker will reject the colliding name and protect you.
  • Assuming replication can merge two independently-written logs into one consistent ordering.

context