skip to content

MM2 mirrors topics and topic configs across clusters, but it does NOT mirror ACLs or SASL credentials. As a principal engineer, how do you keep authorization and credentials consistent across clusters for failover?

level: principalimportance: should knowfreq 30%

answer

  1. MM2 mirrors data/offsets/configs, NOT ACLs/creds
  2. Identity-as-code: Terraform/GitOps/Strimzi to all clusters
  3. Stable identical principals across clusters
  4. Secrets manager for SCRAM/certs, not ZK/KRaft copy
  5. IdentityReplicationPolicy keeps topic names portable

basics

~20 s

MM2 copies topic data and (optionally) topic configs, but it does not replicate ACLs or SCRAM credentials between clusters. You propagate those out-of-band with infrastructure-as-code (e.g. Terraform/GitOps) or a control plane, so the failover cluster already grants every consumer/producer the same access under the same identity.

solid answer

~50 s

MM2's `MirrorSourceConnector` replicates records and, with `sync.topic.configs.enabled`, topic configurations — but ACLs and authentication credentials are deliberately out of scope. If you fail consumers over to the DR cluster, their offsets are translated by the checkpoint connector, yet they'll hit `TopicAuthorizationException` unless the DR cluster already grants identical ACLs, and `SaslAuthenticationException` unless their SASL/SCRAM users exist there. There was an `acl.sync.enabled` option in MM2 with narrow semantics, but the robust approach is to manage authorization and identity as **infrastructure-as-code**: a single source of truth (Terraform Kafka provider, GitOps, or a control plane like Confluent/Strimzi) that provisions the same principals, SCRAM credentials, RBAC role bindings, and ACLs to every cluster. Principal naming must be **stable and identical across clusters** (consistent `ssl.principal.mapping.rules` or SCRAM usernames) so one ACL definition applies everywhere. Treat replication topology, ACLs, and credentials as one deployable unit so DR cutover is authorization-transparent.

go deeper

for a junior

Know MM2 copies data but not ACLs or credentials — those are set up separately per cluster.

for a middle

Recognize failover needs matching ACLs + SCRAM users on the DR cluster, and that topic prefixes affect ACL names.

for a senior

Manage ACLs/credentials via IaC with stable principals and align offset translation with Group ACLs.

for a principal

Architect identity-as-code across the whole failover group: single source of truth, secrets-manager-backed credential fan-out, drift detection, and replication-policy choice for ACL portability.

## The core gap MM2 replicates the **data plane**: records, consumer-group offsets (translated), and optionally topic configs (`sync.topic.configs.enabled`, default true) and topic ACLs in some versions. It does **not** durably replicate the **identity/authorization plane**: - **SASL/SCRAM credentials** live in cluster metadata (historically a ZooKeeper/KRaft store) and are per-cluster. - **ACLs** are per-cluster authorizer state. - **RBAC role bindings** (in commercial distros) are per-cluster. So a perfectly replicated topic on the DR cluster is useless to a failed-over client if (a) the client's identity doesn't exist there, or (b) no ACL grants it access. ## Why MM2 doesn't (reliably) do it Authorization is sensitive and cluster-specific; blindly copying ACLs can leak access if principal mappings differ, and copying SCRAM secrets across a trust boundary is a credential-distribution problem you generally don't want a data-replication tool owning. MM2 has had an `acl.sync` capability but it is narrow (only topic Read/Write derived from MM2's own config) and is not a substitute for full identity management. Credentials it never syncs. ## The principal-engineering answer: identity as code 1. **Single source of truth.** Define principals, SCRAM users, ACLs, and RBAC bindings in a declarative system — Terraform (Kafka/Confluent providers), Strimzi `KafkaUser`/`KafkaTopic` CRDs, or a GitOps pipeline. Apply the same definitions to **every** cluster in the failover group. 2. **Stable, cluster-independent principals.** The same client must authenticate as the *same* `KafkaPrincipal` on every cluster. For mTLS that means consistent `ssl.principal.mapping.rules` so different per-cluster certs still map to `User:orders-svc`. For SCRAM it means the same usernames provisioned everywhere. For OAuth it means the same token claim mapping. If principals differ per cluster, one ACL set can't cover all clusters and DR becomes bespoke. 3. **Credential distribution.** SCRAM passwords / certificates must be provisioned to each cluster and to clients via a secrets manager (Vault, AWS Secrets Manager, K8s Secrets) — not by copying ZooKeeper/KRaft state. Rotation must fan out to all clusters. 4. **Topic-name awareness.** If you use the default `DefaultReplicationPolicy`, DR topics are prefixed (`primary.orders`); ACLs and client configs on the DR side must match the prefixed names — or use `IdentityReplicationPolicy` so names are identical across clusters, which makes ACL definitions portable (at the cost of losing loop-prevention in active/active). 5. **Validation/drift detection.** Continuously assert that every cluster's ACL/credential state matches the source of truth (e.g. `kafka-acls --list` diffed against the declared set), so silent drift doesn't surface only during a real disaster. ## Active/active vs active/passive nuance - **Active/passive (DR):** the passive cluster needs the full ACL+credential set staged so cutover is instant. - **Active/active:** both clusters serve clients, so both must always hold the complete identity set, and `RemoteClusterUtils`/checkpoint offset translation plus consistent ACLs let a consumer migrate either direction. ## Edge cases - **Principal mapping drift** silently breaks failover: ACLs reference `User:svc` but the DR cluster maps the cert to `User:CN=svc,...` — the cutover authorizes nothing. - **Group ACLs**: failed-over consumers may use a different consumer group name on DR; ensure Group ACLs and offset translation align. - **Credential rotation race**: rotating a SCRAM password on primary but not DR leaves clients unable to authenticate post-failover. - **Auditability**: managing identity as code gives you a reviewable, versioned trail — important for compliance that asks 'who can read PII topic X on every cluster'.

  • A consumer fails over to DR; offsets are correct but it gets TopicAuthorizationException. Root cause?
    The DR cluster lacks an ACL granting that principal Read on the (possibly prefixed) topic, or the principal mapping resolves to a different name than the ACLs reference. MM2 translated offsets but never propagated authorization.
  • Why prefer IdentityReplicationPolicy when portability of ACLs matters?
    It keeps topic names identical across clusters (no source prefix), so one ACL/client-config definition works on every cluster. The tradeoff is losing the prefix-based loop prevention, so it's mainly for active/passive or carefully managed flows.
  • How do you keep SCRAM credentials consistent without copying KRaft/ZooKeeper state?
    Provision users declaratively (Terraform/Strimzi KafkaUser) sourced from a secrets manager, applied to every cluster, with rotation fanned out atomically to all clusters and their clients.

saying these in an interview costs you the question

  • Believing MM2 automatically keeps ACLs and credentials in sync across clusters
  • Copying ZooKeeper/KRaft SCRAM state between clusters as a credential strategy
  • Letting principal names differ per cluster, so one ACL set can't cover failover
  • Forgetting prefixed topic names break DR-side ACLs under DefaultReplicationPolicy
  • Treating data replication as sufficient for DR readiness without identity readiness

context