skip to content

Replication and Durability

How partitions survive broker loss: leader/follower fetching, the ISR, high watermark and leader epoch, and election choices. Interviewers push here to see whether you can reason about the durability-versus-availability dial.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

How does Kafka's rack-aware replica assignment algorithm decide where to place replicas, and what is the relationship between replication factor and the number of racks?

level: middleimportance: should knowfreq 45%

basics

~20 s

Kafka assigns each partition's replicas by cycling through racks so no two replicas share a rack until it runs out of racks. With replication factor 3 and 3 racks, each replica lands in a different rack; with fewer racks than RF, some replicas double up.

open as a page

Give an example of a topic where you'd deliberately enable unclean leader election, and one where you'd never, and justify each.

level: middleimportance: should knowfreq 40%

basics

~20 s

Enable it for high-volume, loss-tolerant streams like metrics or clickstream logs where staying available matters more than a few lost records. Never enable it for financial transactions, payments, or audit logs where losing a committed record is unacceptable — keep those fail-closed.

open as a page

How do default.replication.factor and broker-level defaults relate to per-topic overrides, and how would you use them to manage durability across a mixed-workload or multi-DC cluster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Broker configs like default.replication.factor and min.insync.replicas set what new auto-created topics get. Each topic can override these at creation or via config alteration. Use safe broker defaults, then override per topic when a workload needs different durability or spans data centers.

open as a page

What is the leader epoch, and why is it bumped on every leadership change?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The leader epoch is a number that increases by one each time a partition gets a new leader. It lets brokers and clients tell stale leadership apart from current, so followers truncate correctly and old leaders can't corrupt the log.

open as a page

During failover, what is unclean leader election, what does it trade off, and when (if ever) would you enable it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Unclean leader election lets an out-of-sync replica become leader when no in-sync replica survives. It restores availability but can lose acknowledged data. It's off by default; enable it only when uptime matters more than durability.

open as a page

Explain the data-loss window that exists with Kafka's default flush settings, and what conditions are required to actually lose acknowledged data.

level: seniorimportance: should knowfreq 40%

basics

~20 s

Because Kafka acknowledges writes when replicas have the data in memory (not on disk), acknowledged records can be lost only if every in-sync replica loses power at nearly the same time before any of them flushes to disk. That's the correlated power-loss window.

open as a page

What do the flush.messages and flush.ms settings control, and why are they usually left at their defaults?

level: seniorimportance: should knowfreq 45%

basics

~20 s

flush.messages and flush.ms tell a broker to fsync its log to disk after a certain number of messages or after a time interval. They're usually left at defaults because Kafka relies on replication for durability, so forcing extra fsyncs just hurts throughput.

open as a page

Your monitoring shows UnderReplicatedPartitions spiking and IsrShrinksPerSec elevated, but brokers are all up. How do you diagnose and act?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Under-replicated partitions with all brokers up means followers can't keep up, not that brokers are down. Check follower I/O (disk, network, GC), look at IsrShrinks/ExpandsPerSec for flapping, inspect which brokers are dropping out, and either fix the bottleneck or tune replica.lag.time.max.ms / replica fetcher threads.

open as a page

Walk through what happens, step by step, when a follower falls out of the ISR and later rejoins. Who decides, and how is the change propagated?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The leader detects a follower exceeding replica.lag.time.max.ms, shrinks the ISR, and persists the new ISR via the controller (AlterPartition in KRaft, ZooKeeper in older versions). When the follower catches back up to the leader's LEO, the leader expands the ISR and re-propagates it.

open as a page

Walk through replica.fetch.max.bytes, replica.fetch.wait.max.ms, and num.replica.fetchers — what each controls and how you'd tune them.

level: seniorimportance: should knowfreq 45%

basics

~10 s

replica.fetch.max.bytes caps bytes per partition per fetch; replica.fetch.wait.max.ms is the max long-poll wait when no data is ready; num.replica.fetchers sets how many fetcher threads a follower runs per source broker to parallelize replication.

open as a page

A team reports intermittent producer failures during routine broker maintenance. Their topic is RF=3, min.insync.replicas=3, acks=all. What is wrong and how would you reason about the fix?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With min.insync.replicas equal to RF, taking any one broker down for maintenance drops the ISR from 3 to 2, below the floor of 3, so all acks=all writes fail. Lowering min.insync.replicas to 2 fixes it while keeping dual-copy durability.

open as a page

When is kafka-reassign-partitions.sh needed instead of preferred leader election, and how does reassignment differ from a leader election?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Preferred leader election only changes which existing replica leads. kafka-reassign-partitions.sh actually moves or changes replicas across brokers — needed for adding/removing brokers, fixing skewed data placement, or changing replication factor. It copies data, so it is heavyweight.

open as a page

How does unclean.leader.election.enable interact with min.insync.replicas and acks to position a topic on the consistency-vs-availability spectrum?

level: seniorimportance: should knowfreq 45%

basics

~20 s

acks=all + min.insync.replicas controls write-time durability (rejecting writes when too few replicas are in sync); unclean.leader.election.enable controls failover behavior when all ISR are gone. Together false + acks=all + min.insync.replicas>=2 gives a CP topic that fails closed; unclean=true makes it fail open with possible loss.

open as a page

You must guarantee no acknowledged message is ever lost or silently duplicated end-to-end. Which producer, topic, and broker settings do you combine, and why is acks=all/MISR=2 alone insufficient?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use RF=3, min.insync.replicas=2, acks=all, unclean.leader.election.enable=false on the broker/topic, plus enable.idempotence=true on the producer (with retries and proper delivery.timeout.ms). acks=all/MISR=2 stops loss on the broker side but doesn't prevent producer retries from creating duplicates or unclean election from truncating data.

open as a page

As a principal engineer, how do acks, min.insync.replicas, the high watermark, and leader epochs together determine whether an acknowledged write can ever be lost? What configuration gives the strongest durability and what are the trade-offs?

level: principalimportance: should knowfreq 35%

basics

~20 s

Strongest durability: acks=all, min.insync.replicas=2 (RF=3), and unclean.leader.election.enable=false. Then a write is acked only after the HW advances past it (committed on >=2 in-sync replicas), and leader-epoch truncation prevents divergence on failover. The trade-off is higher latency and reduced availability when replicas fall behind.

open as a page

How do ISR shrink dynamics interact with min.insync.replicas, acks=all, and unclean.leader.election.enable to shape the durability vs. availability tradeoff?

level: principalimportance: should knowfreq 48%

basics

~20 s

When the ISR shrinks, fewer replicas hold committed data. min.insync.replicas sets the floor: below it, acks=all writes are rejected (availability lost to protect durability). unclean.leader.election lets an out-of-sync replica become leader if ISR is empty — restoring availability but risking data loss.

open as a page

As a platform architect, how would you design rack/AZ topology, replication, and fetch strategy for a multi-AZ Kafka cluster to balance durability, availability, and cross-AZ cost? What are the principal tradeoffs?

level: principalimportance: should knowfreq 30%

basics

~20 s

Map broker.rack to AZs, use RF=3 across 3 AZs with min.insync.replicas=2 and unclean.leader.election=false for durability. Enable KIP-392 fetch-from-follower (RackAwareReplicaSelector + client.rack) to cut cross-AZ read costs. The main tradeoffs are cost vs freshness and balanced AZ sizing.

open as a page

All ISR replicas for a partition are dead and the partition is offline because unclean election is disabled. As the on-call operator, how do you decide and how do you force an unclean election to restore availability?

level: principalimportance: should knowfreq 35%

basics

~20 s

First weigh data loss vs downtime: if waiting for an ISR replica to return is acceptable, wait. If availability must be restored now and loss is tolerable, force it — temporarily set unclean.leader.election.enable=true (per-topic), or run kafka-leader-election.sh with --election-type UNCLEAN, then revert. Document the loss.

open as a page

After a failover, how does a former-leader replica that rejoins reconcile its log with the new leader, and what role does the high watermark play?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

When the old leader comes back as a follower, it may hold uncommitted records past the high watermark that the new leader never had. Using leader-epoch lookups it finds the divergence point, truncates the extra records, then re-fetches from the new leader to catch up and rejoin the ISR.

open as a page

Design a leadership-balancing strategy for a large, latency-sensitive cluster during rolling restarts. What do you tune, and what are the failure modes?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Often disable auto.leader.rebalance.enable and trigger preferred election deliberately after each restart batch, scoped via JSON and staggered, so leadership churn is controlled. Watch for client metadata storms, imbalance from out-of-sync preferred replicas, and unclean-election durability risks.

open as a page

showing 31–50 of 50