Replication and Durability
How partitions survive broker loss: leader/follower fetching, the ISR, high watermark and leader epoch, and election choices. Interviewers push here to see whether you can reason about the durability-versus-availability dial.
part ofApache Kafkaoverview, primer and where to startread it →on this pageshowhide
explore
- In-Sync Replicas and Replica Lag5 questions
- Unclean Leader Election5 questions
- Rack Awareness and Replica Placement5 questions
- End-to-End Durability vs Availability Tuning5 questions
- Leadership Failover and Replica State5 questions
questions
page 2 of 2How does Kafka's rack-aware replica assignment algorithm decide where to place replicas, and what is the relationship between replication factor and the number of racks?
basics
~20 sKafka assigns each partition's replicas by cycling through racks so no two replicas share a rack until it runs out of racks. With replication factor 3 and 3 racks, each replica lands in a different rack; with fewer racks than RF, some replicas double up.
Give an example of a topic where you'd deliberately enable unclean leader election, and one where you'd never, and justify each.
basics
~20 sEnable it for high-volume, loss-tolerant streams like metrics or clickstream logs where staying available matters more than a few lost records. Never enable it for financial transactions, payments, or audit logs where losing a committed record is unacceptable — keep those fail-closed.
How do default.replication.factor and broker-level defaults relate to per-topic overrides, and how would you use them to manage durability across a mixed-workload or multi-DC cluster?
basics
~20 sBroker configs like default.replication.factor and min.insync.replicas set what new auto-created topics get. Each topic can override these at creation or via config alteration. Use safe broker defaults, then override per topic when a workload needs different durability or spans data centers.
What is the leader epoch, and why is it bumped on every leadership change?
basics
~20 sThe leader epoch is a number that increases by one each time a partition gets a new leader. It lets brokers and clients tell stale leadership apart from current, so followers truncate correctly and old leaders can't corrupt the log.
During failover, what is unclean leader election, what does it trade off, and when (if ever) would you enable it?
basics
~20 sUnclean leader election lets an out-of-sync replica become leader when no in-sync replica survives. It restores availability but can lose acknowledged data. It's off by default; enable it only when uptime matters more than durability.
What do the flush.messages and flush.ms settings control, and why are they usually left at their defaults?
basics
~20 sflush.messages and flush.ms tell a broker to fsync its log to disk after a certain number of messages or after a time interval. They're usually left at defaults because Kafka relies on replication for durability, so forcing extra fsyncs just hurts throughput.
Your monitoring shows UnderReplicatedPartitions spiking and IsrShrinksPerSec elevated, but brokers are all up. How do you diagnose and act?
basics
~20 sUnder-replicated partitions with all brokers up means followers can't keep up, not that brokers are down. Check follower I/O (disk, network, GC), look at IsrShrinks/ExpandsPerSec for flapping, inspect which brokers are dropping out, and either fix the bottleneck or tune replica.lag.time.max.ms / replica fetcher threads.
Walk through what happens, step by step, when a follower falls out of the ISR and later rejoins. Who decides, and how is the change propagated?
basics
~20 sThe leader detects a follower exceeding replica.lag.time.max.ms, shrinks the ISR, and persists the new ISR via the controller (AlterPartition in KRaft, ZooKeeper in older versions). When the follower catches back up to the leader's LEO, the leader expands the ISR and re-propagates it.
Walk through replica.fetch.max.bytes, replica.fetch.wait.max.ms, and num.replica.fetchers — what each controls and how you'd tune them.
basics
~10 sreplica.fetch.max.bytes caps bytes per partition per fetch; replica.fetch.wait.max.ms is the max long-poll wait when no data is ready; num.replica.fetchers sets how many fetcher threads a follower runs per source broker to parallelize replication.
A team reports intermittent producer failures during routine broker maintenance. Their topic is RF=3, min.insync.replicas=3, acks=all. What is wrong and how would you reason about the fix?
basics
~20 sWith min.insync.replicas equal to RF, taking any one broker down for maintenance drops the ISR from 3 to 2, below the floor of 3, so all acks=all writes fail. Lowering min.insync.replicas to 2 fixes it while keeping dual-copy durability.
When is kafka-reassign-partitions.sh needed instead of preferred leader election, and how does reassignment differ from a leader election?
basics
~20 sPreferred leader election only changes which existing replica leads. kafka-reassign-partitions.sh actually moves or changes replicas across brokers — needed for adding/removing brokers, fixing skewed data placement, or changing replication factor. It copies data, so it is heavyweight.
How does unclean.leader.election.enable interact with min.insync.replicas and acks to position a topic on the consistency-vs-availability spectrum?
basics
~20 sacks=all + min.insync.replicas controls write-time durability (rejecting writes when too few replicas are in sync); unclean.leader.election.enable controls failover behavior when all ISR are gone. Together false + acks=all + min.insync.replicas>=2 gives a CP topic that fails closed; unclean=true makes it fail open with possible loss.
You must guarantee no acknowledged message is ever lost or silently duplicated end-to-end. Which producer, topic, and broker settings do you combine, and why is acks=all/MISR=2 alone insufficient?
basics
~20 sUse RF=3, min.insync.replicas=2, acks=all, unclean.leader.election.enable=false on the broker/topic, plus enable.idempotence=true on the producer (with retries and proper delivery.timeout.ms). acks=all/MISR=2 stops loss on the broker side but doesn't prevent producer retries from creating duplicates or unclean election from truncating data.
As a principal engineer, how do acks, min.insync.replicas, the high watermark, and leader epochs together determine whether an acknowledged write can ever be lost? What configuration gives the strongest durability and what are the trade-offs?
basics
~20 sStrongest durability: acks=all, min.insync.replicas=2 (RF=3), and unclean.leader.election.enable=false. Then a write is acked only after the HW advances past it (committed on >=2 in-sync replicas), and leader-epoch truncation prevents divergence on failover. The trade-off is higher latency and reduced availability when replicas fall behind.
How do ISR shrink dynamics interact with min.insync.replicas, acks=all, and unclean.leader.election.enable to shape the durability vs. availability tradeoff?
basics
~20 sWhen the ISR shrinks, fewer replicas hold committed data. min.insync.replicas sets the floor: below it, acks=all writes are rejected (availability lost to protect durability). unclean.leader.election lets an out-of-sync replica become leader if ISR is empty — restoring availability but risking data loss.
As a platform architect, how would you design rack/AZ topology, replication, and fetch strategy for a multi-AZ Kafka cluster to balance durability, availability, and cross-AZ cost? What are the principal tradeoffs?
basics
~20 sMap broker.rack to AZs, use RF=3 across 3 AZs with min.insync.replicas=2 and unclean.leader.election=false for durability. Enable KIP-392 fetch-from-follower (RackAwareReplicaSelector + client.rack) to cut cross-AZ read costs. The main tradeoffs are cost vs freshness and balanced AZ sizing.
All ISR replicas for a partition are dead and the partition is offline because unclean election is disabled. As the on-call operator, how do you decide and how do you force an unclean election to restore availability?
basics
~20 sFirst weigh data loss vs downtime: if waiting for an ISR replica to return is acceptable, wait. If availability must be restored now and loss is tolerable, force it — temporarily set unclean.leader.election.enable=true (per-topic), or run kafka-leader-election.sh with --election-type UNCLEAN, then revert. Document the loss.
After a failover, how does a former-leader replica that rejoins reconcile its log with the new leader, and what role does the high watermark play?
basics
~20 sWhen the old leader comes back as a follower, it may hold uncommitted records past the high watermark that the new leader never had. Using leader-epoch lookups it finds the divergence point, truncates the extra records, then re-fetches from the new leader to catch up and rejoin the ISR.
Design a leadership-balancing strategy for a large, latency-sensitive cluster during rolling restarts. What do you tune, and what are the failure modes?
basics
~20 sOften disable auto.leader.rebalance.enable and trigger preferred election deliberately after each restart batch, scoped via JSON and staggered, so leadership churn is controlled. Watch for client metadata storms, imbalance from out-of-sync preferred replicas, and unclean-election durability risks.
showing 31–50 of 50