What is partition-count and config drift between a source and a mirrored Kafka cluster, why is it dangerous, and how does MirrorMaker 2 try to keep them in sync?
answer
- drift = partition count + topic configs diverge
- hash(key) % numPartitions => key->partition breaks
- MM2 propagates partition increases, never shrinks
- sync.topic.configs.enabled + exclude list
- never hand-edit mirrored target topics
basics
~20 sDrift is when the target topic ends up with a different partition count or different topic configs than the source. It's dangerous because key-based ordering and partitioning break after failover. MM2 periodically syncs partition counts and topic configs to keep them matched.
solid answer
~50 sConfig/partition drift means the mirrored topic on the target diverges from the source: a different number of partitions, or different topic-level configs (retention.ms, cleanup.policy, max.message.bytes, etc.). Partition-count drift is especially dangerous: Kafka routes a keyed record by hash(key) % numPartitions, so if the target has a different partition count, the same key lands on a different partition number — breaking per-key ordering and any partition-affinity assumptions after a DR failover. Config drift can silently change retention or compaction semantics, causing data to expire differently on each side. MM2's MirrorSourceConnector periodically reconciles: it auto-creates target topics, propagates source partition counts (adding partitions when the source grows), and syncs allowed topic configs on an interval controlled by settings like sync.topic.configs.enabled and the refresh intervals. Caveats: MM2 doesn't shrink partitions, certain configs are excluded from sync (config.properties.exclude / a default blacklist), and manual changes on the target can re-introduce drift between sync cycles.
go deeper
Know that the copy should have the same number of partitions and similar settings, and that mismatches cause problems.
Explain hash(key)%partitions and why differing counts break ordering after failover.
Describe MM2's MirrorSourceConnector reconciliation, sync.topic.configs, exclusion lists, and the no-shrink caveat.
Govern drift across a fleet: change-management for partition counts, what's intentionally not synced (ACLs/quotas), and failover correctness guarantees.
**What 'drift' means.** When MirrorMaker 2 (MM2) mirrors a topic, the target copy should structurally match the source. *Drift* is any divergence between them: most importantly the **partition count**, but also **topic-level configuration** such as `retention.ms`, `cleanup.policy` (delete vs compact), `max.message.bytes`, `min.insync.replicas`, and `segment.bytes`. **Why partition-count drift is dangerous.** A Kafka *partition* is the unit of ordering and parallelism; a topic with N partitions is N ordered logs. When a producer sends a record with a key, the default partitioner computes `partition = hash(key) % numPartitions`. So which partition a key lands in depends on the *partition count*. If the source has 12 partitions and the target has 6, then after a DR *failover* (clients now produce to the target) the same key `user-42` hashes to a different partition number than it did on the source. Consequences: (1) **per-key ordering breaks** — a consumer that assumed all records for a key are in one partition now sees them split or interleaved across the source-era and target-era histories; (2) **stateful/co-partitioned joins** (Kafka Streams) that rely on matching partition counts across topics fail or rebalance incorrectly; (3) **offset/position assumptions** about specific partitions become invalid. This is a *silent correctness* bug, not a crash, which makes it worse. **Why config drift is dangerous.** If `retention.ms` differs, data expires at different times on each cluster — the DR copy may have already deleted records you assumed were safe, blowing your RPO assumptions. If `cleanup.policy` differs (source compacted, target delete), the log *content* diverges. If `max.message.bytes` is smaller on the target, large records that the source accepted get rejected on the replica, silently dropping data. **How MM2 reconciles.** The `MirrorSourceConnector` is responsible for keeping topics aligned: - **Auto-create + partition propagation:** it creates missing target topics and, on a refresh interval, detects when the *source* partition count grew and adds partitions to the target so counts match. (It does NOT shrink partitions — Kafka can't reduce partition count anyway.) - **Config sync:** when `sync.topic.configs.enabled=true` (default), it periodically copies allowed source topic configs to the target. A default exclusion list / `config.properties.exclude` keeps replication-specific or broker-managed configs from being overwritten (e.g. you don't blindly copy `min.insync.replicas` if topologies differ). - **Refresh intervals:** `refresh.topics.interval.seconds` and `sync.topic.configs.interval.seconds` (names vary slightly by version) control how often reconciliation runs — so there is always a window where manual target changes cause temporary drift until the next cycle. **Edge cases and pitfalls.** - If someone manually adds partitions on the *target*, MM2 won't remove them, and now hashing differs permanently — never edit mirrored topics by hand. - Excluded configs (by design) can drift legitimately; know which ones MM2 ignores. - Increasing partitions on the source *also* breaks key->partition mapping going forward (a general Kafka caveat), and MM2 faithfully propagates that — so partition increases should be rare and deliberate on both sides. - ACLs and quotas are generally NOT synced by MM2 (separate concern), another form of operational drift.
- Exactly why does a different partition count on the target break things after failover?Default keyed routing is partition = hash(key) % numPartitions. A different count changes the modulus, so the same key maps to a different partition. Per-key ordering, co-partitioned Streams joins, and partition-affinity logic all break, because records for a key are no longer guaranteed to live in one consistent partition.
- Which kinds of config does MM2 deliberately NOT sync, and why?Replication- and broker-managed properties on an exclusion list (config.properties.exclude / a default blacklist) — e.g. things tied to local cluster topology — plus generally ACLs and quotas. They're excluded so MM2 doesn't clobber intentionally different local settings or impose source-only assumptions on the target.
saying these in an interview costs you the question
- Saying MM2 keeps partition counts identical by shrinking the target (Kafka can't shrink partitions).
- Assuming all topic configs are blindly copied (there's an exclusion list; ACLs/quotas aren't synced).
- Hand-editing partitions on the mirrored target and expecting MM2 to fix it.