skip to content

In active-active, both regions accept writes to the same logical entity. What write-conflict and ordering hazards arise across regions, and how do idempotency keys and dedup help?

level: seniorimportance: must knowfreq 60%

answer

  1. MM2 = per-partition order only, at-least-once
  2. enable.idempotence is per-cluster, doesn't cross MM2
  3. idempotency key + dedup store / compacted topic
  4. conflict: LWW vs CRDT vs home-region affinity
  5. no dual-write; one producer region per record

basics

~20 s

Two regions can write conflicting updates to the same entity concurrently, and MM2 gives no cross-region ordering or conflict resolution — it just copies records. You handle conflicts at the application layer: attach idempotency keys so re-delivered or duplicate records are deduped, and use last-writer-wins, CRDTs, or entity-region affinity to resolve concurrent updates.

solid answer

~60 s

Active-active means the same logical key can be updated in both regions concurrently. MM2 replicates records as-is and preserves only **per-partition order within a single source topic** — it provides no global ordering and no conflict resolution. So if region A sets balance=100 and region B sets balance=120 at nearly the same time, each region first applies its own then receives the other's mirrored record, and they can converge to different states. Kafka's idempotent producer (`enable.idempotence=true`) only dedupes retries from one producer to one partition on one cluster — it does **not** survive replication, so the mirrored copy can be re-consumed as a duplicate. The fixes are application-level: (1) embed an **idempotency key / event ID** in each record and have consumers dedup against a store (or use a key so re-processing is naturally idempotent); (2) resolve concurrent conflicts with last-writer-wins using a logical clock/timestamp, CRDTs, or **entity-to-region affinity** (route all writes for a given key to one home region) to avoid the conflict entirely. Don't dual-write the same record into both regions.

go deeper

for a junior

Understand that two regions can write the same thing at once and Kafka won't auto-merge it; duplicates and conflicts are possible.

for a middle

Explain that MM2 is at-least-once with per-partition ordering only, and that idempotency keys plus a dedup store handle duplicate mirrored records.

for a senior

Reason about why enable.idempotence doesn't cross clusters, and select among LWW, CRDT, and home-region affinity per data shape; forbid dual-writes.

for a principal

Define per-domain conflict policies, clock/logical-clock strategy, dedup infrastructure (compacted topics/KV), and the consistency contract the org commits to under active-active.

## The core hazard In active-active, **both** clusters accept writes for the **same** logical entity (say, account `acct-42`). Region A and region B can each receive an update for `acct-42` at the same wall-clock moment, with neither having seen the other's. MM2 then mirrors each write to the other side. There is no coordination, no lock, no global sequencer. ## What MM2 guarantees — and what it does NOT MM2 (`MirrorSourceConnector`) **preserves per-partition ordering within a single replicated topic**: records from partition 3 of A's `orders` arrive in the same order on B's `A.orders` partition 3. It does **not** guarantee: - **Global / cross-topic / cross-region ordering** — `orders` (local on B) and `A.orders` (remote on B) are independent topics; their records interleave arbitrarily. - **Exactly-once across the replication boundary** — MM2 is effectively at-least-once; a mirrored record can be re-delivered, so **duplicates are expected**. - **Conflict resolution** — MM2 copies bytes; it has no notion that two records target the same entity. ## Why Kafka's built-in idempotence doesn't save you The **idempotent producer** (`enable.idempotence=true`, default in modern clients) attaches a producer ID + sequence number so the **broker** dedupes retried produces to **one partition on one cluster**. This scope is critical: it does not span clusters. When MM2 reads a record and produces it to the remote cluster, that's a *new* produce with a *new* producer ID — the original idempotence metadata is gone. So the remote copy is just a normal record that downstream consumers can see again (e.g., after a consumer restart and offset replay). Cross-region dedup is therefore an application responsibility. ## Hazard 1 — duplicates / dual-write A **dual-write hazard** is when the same logical record ends up written to *both* regions independently (e.g., a client or gateway that writes to A and B for 'safety'). Combined with MM2 mirroring, you now have multiple copies and no way to tell them apart unless they carry a stable ID. **Rule: produce each record to exactly one region and let MM2 fan it out.** Don't dual-write. ### Idempotency keys Give every business event a stable **idempotency key / event ID** (UUID, or a deterministic hash of the natural key + operation). Consumers maintain a dedup store (a compacted Kafka topic keyed by the ID, a fast KV store, or a processed-IDs set) and skip an ID they've already applied. This makes at-least-once delivery safe: re-delivered mirror copies are dropped. For idempotent *state* updates, designing the operation so re-applying it is a no-op (e.g., `SET balance=100` keyed by entity, not `ADD 10`) achieves the same effect. ## Hazard 2 — concurrent conflicting updates Two regions change the *same* key to *different* values concurrently. Resolution strategies: - **Last-writer-wins (LWW)**: tag each write with a timestamp or logical clock; on convergence keep the highest. Simple, but clock skew can drop a valid update — prefer logical clocks/version vectors over raw wall-clock. - **CRDTs**: model the value as a conflict-free replicated data type (counters, OR-sets) so any merge order converges deterministically without losing updates. - **Entity-region affinity (home region / single-writer per key)**: route all writes for a given key to one designated region; the other region only reads the mirrored copy. This *eliminates* the conflict rather than resolving it, at the cost of cross-region write latency for non-home traffic. This is the most robust pattern when strong consistency per entity matters. ## Ordering & dedup across regions — practical recipe 1. One producer region per record (no dual-write). 2. Stable idempotency key on every event. 3. Consumers dedup by key against a store / use compacted topics. 4. Conflict policy chosen per domain: LWW for tolerant data, CRDT for counters/sets, home-region affinity for strong per-entity consistency. 5. Accept that only per-partition order is preserved; never assume global order across the aggregate. ## Edge cases - **Offset translation, not value dedup**: MM2's `MirrorCheckpointConnector` + `RemoteClusterUtils` translate *consumer-group offsets* for failover; that is unrelated to deduping record *values* across regions. - **Compaction races**: if you dedup via a compacted topic, remember compaction is eventual; a consumer may still see two copies before cleanup. - **Clock skew with LWW**: synchronize clocks (NTP/PTP) or use hybrid logical clocks to limit lost updates.

  • Why doesn't enable.idempotence=true on the producer prevent duplicate processing of mirrored records in the other region?
    The idempotent producer dedupes retries via producer-ID + sequence number scoped to one partition on one cluster. When MM2 re-produces the record to the remote cluster it's a brand-new produce with a new producer ID, so that metadata doesn't carry over. Cross-region dedup must be done at the application layer with a stable idempotency key.
  • Give a concrete conflict-resolution strategy for a balance/counter that's updated in both regions, and why LWW is a poor fit.
    Use a CRDT counter (e.g., a PN-counter): each region increments its own sub-counter and merges by summing, so concurrent increments all survive regardless of merge order. LWW is poor here because picking the 'latest' write discards the other region's increment, losing money. LWW suits overwrite-style fields, not additive ones.
  • What is the dual-write hazard and how do you avoid it?
    Dual-write is when a client writes the same logical record directly into both regions to 'be safe', producing independent copies that MM2 then also mirrors, causing duplicates with no shared identity. Avoid it by producing each record to exactly one region and letting MM2 replicate it; if you need it in both for reads, that's MM2's job, not the producer's.

saying these in an interview costs you the question

  • Claiming enable.idempotence or transactions give exactly-once across clusters via MM2
  • Saying MM2 resolves write conflicts or guarantees global ordering
  • Recommending dual-writing the same record to both regions
  • Using wall-clock last-writer-wins for additive values (counters/balances), silently losing updates
  • Confusing MM2 offset translation (failover) with value-level dedup

context