How do default.replication.factor and broker-level defaults relate to per-topic overrides, and how would you use them to manage durability across a mixed-workload or multi-DC cluster?
answer
- default.replication.factor = auto-created topics only
- topic override > broker default
- RF needs reassignment; MISR/unclean are dynamic
- tier topics: critical vs disposable
- multi-DC: rack awareness + cross-DC MISR + unclean=false
basics
~20 sBroker configs like default.replication.factor and min.insync.replicas set what new auto-created topics get. Each topic can override these at creation or via config alteration. Use safe broker defaults, then override per topic when a workload needs different durability or spans data centers.
solid answer
~40 sBroker-level settings define cluster-wide defaults: default.replication.factor governs the RF assigned to auto-created topics, and the broker's min.insync.replicas / unclean.leader.election.enable become the baseline that topics inherit. Per-topic overrides (kafka-topics --create --config, or kafka-configs --alter --entity-type topics) take precedence over the broker defaults. The pattern is: set conservative, durable broker defaults (e.g. default.replication.factor=3, min.insync.replicas=2, unclean.leader.election.enable=false) so any accidentally auto-created topic is safe, then raise or relax per topic. For mixed workloads, financial/audit topics might use RF=3/MISR=2 while a high-volume metrics topic uses RF=2/MISR=1 for cost. For multi-DC stretch clusters you bump RF (e.g. 4 or 6) with rack/DC awareness so copies span DCs, set MISR to require a cross-DC ack, and keep unclean=false so a DC partition never silently truncates.
go deeper
Know broker defaults seed new topics and each topic can override them.
Distinguish auto-create defaults from explicit creation and which configs are dynamic vs require reassignment.
Design per-tier overrides and reason about multi-DC RF/MISR/rack-awareness trade-offs.
Own cluster-wide governance: safe defaults, auto-create policy, tiering standards, and cross-DC durability architecture.
**Two configuration scopes.** Kafka settings exist at the *broker* level (in `server.properties` / KRaft controller config, cluster-wide) and the *topic* level (per topic). For replication, the relevant broker defaults are: - `default.replication.factor` — RF used when a topic is **auto-created** (i.e. `auto.create.topics.enable=true` and a producer/consumer touches a nonexistent topic). It does NOT retroactively change existing topics and is ignored if you specify RF explicitly at creation. - `min.insync.replicas` — the cluster default MISR inherited by topics that don't override it. - `unclean.leader.election.enable` — the cluster default for unclean elections. - (`num.partitions` similarly defaults partition count for auto-created topics.) **Override precedence.** A per-topic config always wins over the broker default. You set it at creation: `kafka-topics.sh --create --topic payments --partitions 12 --replication-factor 3 --config min.insync.replicas=2` or change it later: `kafka-configs.sh --alter --entity-type topics --entity-name payments --add-config min.insync.replicas=2` Note: RF itself cannot be changed by `--alter`; increasing RF requires a reassignment (`kafka-reassign-partitions.sh`). MISR and unclean.leader.election.enable, however, are dynamic and can be altered live. **Why conservative broker defaults matter.** If `auto.create.topics.enable` is on, a typo or a new service can spawn a topic with whatever the broker defaults are. If those defaults are RF=1/MISR=1, you silently create undurable topics. Best practice: default.replication.factor=3, min.insync.replicas=2, unclean.leader.election.enable=false — so even accidental topics are durable — and ideally disable auto-creation in production entirely. **Mixed-workload tiers.** Different topics warrant different points on the durability/availability/cost curve: - Critical (payments, audit): RF=3, MISR=2, acks=all, unclean=false. - Bulk/observability (metrics, logs): RF=2, MISR=1 (or even RF=1 for truly disposable data) to cut storage and network cost, accepting weaker guarantees. Overrides let one cluster serve both without separate clusters. **Multi-DC considerations.** For a stretch cluster spanning DCs: - Increase RF so copies land in multiple DCs (e.g. RF=4 split 2+2, or RF=6). - Use **rack awareness** (`broker.rack` + the rack-aware replica assignment) so the RF copies are distributed across DCs/racks rather than piling into one. - Set MISR so a committed write must be acknowledged across DCs (e.g. MISR=3 with a 2+2 layout forces at least one replica in the remote DC), giving cross-DC durability. - Keep `unclean.leader.election.enable=false` so that during a DC-to-DC network partition the minority side cannot elect a stale leader and truncate data. - Be aware of the latency cost: cross-DC acks add round-trip latency to every acks=all produce. **Edge cases.** default.replication.factor only affects auto-creation; explicit creation must pass `--replication-factor`. Lowering MISR on a live topic instantly relaxes the durability contract for new writes; raising it can immediately make under-replicated partitions reject writes. Always reconcile RF and MISR (MISR must be <= RF) when scripting topic provisioning.
- Can you increase a topic's replication factor with kafka-configs --alter?No. MISR and unclean.leader.election.enable are dynamic topic configs you can --alter, but RF is changed by a partition reassignment (kafka-reassign-partitions.sh) that adds the new replicas and lets them catch up.
- Why disable auto.create.topics.enable in production even if your broker defaults are safe?To prevent typos or rogue clients from silently creating topics with default partition counts and configs, which can fragment data, bypass naming/governance conventions, and surprise capacity planning — explicit, reviewed topic provisioning is safer.
saying these in an interview costs you the question
- Claiming default.replication.factor changes existing topics' RF (it only affects auto-created ones).
- Saying you can raise RF with a simple config --alter (it requires reassignment).
- Ignoring rack awareness when spreading replicas across DCs, so all 'copies' land in one DC.