skip to content

How would you design an active-active bidirectional MM2 topology, and how do replication flows avoid infinite loops?

level: principalimportance: should knowfreq 35%

answer

  1. Two flows: A->B and B->A, full connector set each
  2. Loop prevention = remote prefixing + exclude already-remote
  3. Produce local, consume local + *.topic (fan-in on read)
  4. Checkpoints both ways for failover; heartbeats both ways
  5. MM2 = at-least-once, no global dedup/conflict resolution

basics

~20 s

You define two flows (A->B and B->A), each running its own set of MM2 connectors. Loops are avoided because mirrored topics are prefixed with the source cluster alias, and MM2's default filters exclude already-remote topics, so B never re-mirrors A's data back to A.

solid answer

~40 s

An active-active setup enables both `A->B.enabled=true` and `B->A.enabled=true`, each spawning a full connector set. The loop-prevention mechanism is the combination of (1) remote-topic prefixing — A's `orders` becomes `A.orders` on B — and (2) MM2 detecting and excluding already-remote topics so B->A doesn't re-replicate `A.orders` back to A. Producers write to local `orders`; consumers in each DC read both local `orders` and the remote `<other>.orders` to get the full global view. Configs sync per flow, checkpoints/heartbeats run both directions for failover and monitoring. Operationally you run one (or two) Connect cluster(s) hosting all four-plus connectors, size tasks.max per flow, and ensure ACL/config sync don't conflict. The naming policy is pluggable (ReplicationPolicy); the default prefixing is what makes cycle detection and consumer aggregation work.

go deeper

for a junior

Know that active-active needs replication in both directions and that prefixing keeps copies distinct.

for a middle

Explain the two-flow setup and that loop prevention relies on prefixing plus excluding remote topics.

for a senior

Detail the produce-local/consume-local+remote pattern, both-direction checkpoints/heartbeats, and at-least-once semantics.

for a principal

Architect the full mesh: per-flow filters, config/ACL sync ownership, RF per cluster, conflict-resolution responsibility, capacity, and failure modes.

**Goal of active-active.** Two datacenters/regions both accept writes for the same logical topics, and each region should be able to see the global stream (local + remote events). This supports geo-local latency, regional failover, and aggregated processing. **Topology.** A 'flow' is one directional replication (a set of MM2 connectors). For active-active you enable both directions: `A->B.enabled = true` and `B->A.enabled = true`. Each flow gets its own MirrorSourceConnector (data), MirrorCheckpointConnector (offsets, for failover), and MirrorHeartbeatConnector (liveness). You can host all of these on a single dedicated Connect/MM2 cluster, or split per region. **The loop problem.** Naively, A->B copies `orders` to B, then B->A would copy that same data back to A, and so on forever — an amplifying replication storm. MM2 prevents this with two cooperating mechanisms: 1. **Remote-topic prefixing** (the *naming* mechanism, owned by a sibling leaf, so just the dependency here): when A->B mirrors A's `orders`, the target topic is named with A's alias, e.g. `A.orders`, distinguishing source-local from remote topics by name. 2. **Excluding already-remote topics from re-replication.** The B->A flow's MirrorSourceConnector recognizes that `A.orders` is a *remote* topic (it originated from A) and does not mirror it back to A. This is enforced by the default topic-exclude patterns plus the connector's knowledge of remote prefixes / the configured ReplicationPolicy. The net effect: data flows A->B once and B->A once, but never bounces. **Application pattern.** Producers always write to the **local** non-prefixed topic (`orders`). A consumer that needs the *global* view subscribes to a pattern matching both local and remote variants — e.g. `orders` plus `*.orders` (or `A.orders`, `B.orders`) — often via a regex subscription. This 'fan-in on read' is the canonical MM2 active-active consumption model. (Some teams instead aggregate into a third 'aggregate' cluster — a different topology.) **Failover & offsets.** Checkpoint connectors run in both directions so that if region A dies, consumers can move to B and resume from translated offsets (mechanism owned by the offset-translation sibling leaf). Heartbeats both ways let you measure each direction's lag independently. **Operational design considerations (principal-level):** - **Config & ACL sync conflicts:** with `sync.topic.configs.enabled`, ensure both directions don't fight over the same topic's config — generally only the source-of-truth direction syncs a given topic's settings since the remote copy has a different name. - **Replication factor & sizing:** set `replication.factor` per cluster's broker count; A and B may differ. - **Filters define the mesh:** the `topics`/`topics.exclude` and `groups` filters per flow must be set so each flow replicates only the locally-originated topics, reinforcing loop prevention. - **Ordering / dedup:** active-active does not magically dedupe; if the same logical key can be produced in both regions, the application must handle conflict resolution — MM2 gives at-least-once delivery, not global ordering or conflict resolution. - **Capacity:** size tasks.max and Connect workers for the sum of both flows' partitions; WAN producer tuning matters in both directions. **Common failure modes:** clearing the default `topics.exclude` (re-enabling loops), mis-prefixing via a custom ReplicationPolicy that breaks remote-topic detection, or having consumers subscribe only to the local topic and silently missing remote events.

  • How does a consumer in region B get the full global view of an active-active topic?
    It subscribes to both the local topic (orders) and the remote-prefixed copies (e.g. A.orders), typically via a regex subscription like 'orders|.*\.orders'. Producers always write local; consumers fan-in on read across local + remote topics.
  • What prevents the B->A flow from re-replicating data that A originally sent to B?
    The data on B is named A.orders (remote-topic prefixing). The B->A MirrorSourceConnector recognizes A.orders as already-remote and excludes it from replication back to A, so the loop is broken.
  • Does active-active MM2 give you global ordering or automatic conflict resolution?
    No. MM2 provides at-least-once cross-cluster delivery with per-partition ordering only. If the same key is written in both regions, the application must resolve conflicts; MM2 does not dedupe or impose global order.

saying these in an interview costs you the question

  • Claiming MM2 provides global ordering or automatic dedup/conflict resolution in active-active — it does not.
  • Saying you only need one flow for bidirectional replication — you need both A->B and B->A enabled.
  • Forgetting consumers must read both local and remote-prefixed topics for the global view.
  • Believing loops are prevented by luck rather than remote-prefixing + already-remote exclusion (don't clear topics.exclude).

context