skip to content

End-to-End Durability vs Availability Tuning

Tuning replication factor, min.insync.replicas, acks, and unclean election together to hit a stated durability target. Interviewers use it as the synthesis question after the individual knobs.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What are the three main Kafka settings that work together to control write durability, and what does each do at a high level?

level: juniorimportance: must knowfreq 78%

answer

  1. RF = how many copies
  2. min.insync.replicas = how many must be caught up
  3. acks = how many producer waits for
  4. RF=3, MISR=2, acks=all baseline
  5. MISR only bites with acks=all

basics

~20 s

Replication factor (how many copies of each partition), min.insync.replicas (how many copies must be caught up to accept a write), and acks (how many copies the producer waits for: 0, 1, or all). Together they decide how many failures you can survive without losing data.

solid answer

~40 s

Three knobs interact. Replication factor (RF), a topic property, sets how many brokers hold a copy of each partition; RF=3 means a leader plus two followers. min.insync.replicas (a topic/broker config) sets the minimum number of in-sync replicas (ISR) that must be present for an acks=all write to be accepted; if fewer are in sync, the broker rejects the write with NotEnoughReplicas. acks (a producer config) controls how many replicas the producer waits to acknowledge: acks=0 (fire and forget), acks=1 (leader only), acks=all (every member of the ISR). The classic durable baseline is RF=3, min.insync.replicas=2, acks=all: it tolerates one broker loss and still guarantees no acknowledged write is lost.

go deeper

for a junior

Memorize the three knobs and the RF=3/MISR=2/acks=all baseline and what each word means.

for a middle

Explain ISR dynamics and that min.insync.replicas only applies to acks=all.

for a senior

Articulate the exact guarantee (acknowledged write survives N broker losses) and the durability-vs-availability tension when MISR approaches RF.

for a principal

Reason about org-wide defaults, cost of RF, and how these defaults compose with idempotence/transactions for end-to-end guarantees.

## What governs durability Kafka stores each topic partition as an ordered log replicated across brokers for fault tolerance. Three settings govern the durability of a write: - **Replication factor (RF)** — a per-topic setting (`--replication-factor` at creation, or the broker default `default.replication.factor`). RF=N means each partition has N copies on N different brokers: one **leader** that handles all reads and writes, and N-1 **followers** that continuously fetch from the leader to stay current. RF determines the maximum number of broker failures the data can physically survive (you can lose up to N-1 copies and still have one). - **In-Sync Replicas (ISR)** — the subset of replicas (including the leader) that are fully caught up to the leader's log. A follower drops out of the ISR if it falls behind by more than `replica.lag.time.max.ms` (default 30s). The ISR shrinks and grows dynamically. - **min.insync.replicas** — a topic/broker config naming the minimum ISR size required to accept a write *when the producer uses acks=all*. If the current ISR is smaller than this, the leader rejects the produce request with `NotEnoughReplicasException` / `NotEnoughReplicasAfterAppendException`. This protects you from the dangerous case where the ISR has collapsed to just the leader: without it, an acks=all write to a lone leader could be lost if that leader then dies. - **acks** — a producer config: `acks=0` means the producer never waits (highest throughput, can silently lose data); `acks=1` waits only for the leader to write to its log (lost if the leader dies before a follower replicates); `acks=all` (a.k.a. `acks=-1`) waits for every member of the current ISR to acknowledge. ## How they combine Durability comes from `acks=all` AND `min.insync.replicas >= 2`. With acks=all, the producer is told a write succeeded only after all ISR members have it; with min.insync.replicas=2, the write is only accepted while at least two replicas are in sync, so an acknowledged record always lives on at least two brokers. The canonical durable config is **RF=3, min.insync.replicas=2, acks=all**: it survives one broker failure with zero acknowledged-data loss and stays available for writes. Setting min.insync.replicas equal to RF (e.g. 3) maximizes durability but means *any* single replica falling out of the ISR halts writes — a **durability/availability trade-off**. ## Edge cases - `min.insync.replicas` only matters with acks=all; with acks=1 it is ignored. - RF without acks=all gives you read availability but not write durability guarantees.

  • Does min.insync.replicas have any effect when the producer uses acks=1?
    No. min.insync.replicas is only enforced for acks=all (acks=-1) writes. With acks=1 the leader acknowledges on its own write regardless of ISR size, so the setting is ignored.
  • Why is RF=3 with min.insync.replicas=2 preferred over RF=2 with min.insync.replicas=2?
    With RF=2/MISR=2 you have no headroom: losing one broker drops the ISR to 1, below min.insync.replicas, so all writes stop immediately. RF=3/MISR=2 tolerates one failure and stays writable.

saying these in an interview costs you the question

  • Saying min.insync.replicas guarantees durability even with acks=1 (it does nothing there).
  • Confusing RF with min.insync.replicas — RF is total copies, MISR is the required caught-up subset.
  • Claiming acks=all waits for ALL replicas (it waits for all members of the current ISR, which may be smaller than RF).

context

open as a page

Explain why setting min.insync.replicas equal to the replication factor hurts availability, and how to pick the value to hit a specific durability/availability target.

level: middleimportance: must knowfreq 70%

basics

~20 s

If min.insync.replicas equals RF, every replica must be in sync to accept writes, so losing or lagging just one broker stops all writes. Setting it to RF-1 (e.g. 2 with RF=3) keeps you durable while tolerating one failure.

open as a page

What does unclean.leader.election.enable do, and how does it change the durability/availability trade-off when combined with the other replication settings?

level: seniorimportance: must knowfreq 62%

basics

~20 s

It decides whether a partition can elect a leader from a replica that was NOT in the in-sync set when all in-sync replicas are down. Enabled = stay available but possibly lose data; disabled (default) = stay consistent but the partition goes offline until an in-sync replica returns.

open as a page

How do default.replication.factor and broker-level defaults relate to per-topic overrides, and how would you use them to manage durability across a mixed-workload or multi-DC cluster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Broker configs like default.replication.factor and min.insync.replicas set what new auto-created topics get. Each topic can override these at creation or via config alteration. Use safe broker defaults, then override per topic when a workload needs different durability or spans data centers.

open as a page

You must guarantee no acknowledged message is ever lost or silently duplicated end-to-end. Which producer, topic, and broker settings do you combine, and why is acks=all/MISR=2 alone insufficient?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use RF=3, min.insync.replicas=2, acks=all, unclean.leader.election.enable=false on the broker/topic, plus enable.idempotence=true on the producer (with retries and proper delivery.timeout.ms). acks=all/MISR=2 stops loss on the broker side but doesn't prevent producer retries from creating duplicates or unclean election from truncating data.

open as a page