During broker turnover (adding/removing nodes), explain unclean.leader.election.enable and the durability tradeoff it represents.
answer
- no ISR replica -> elect out-of-sync replica
- default false since 0.11 (durability)
- true = availability but committed data truncated/lost
- per-topic override possible
- RF=3 + min.insync=2 + acks=all keeps you away from it
- controlled shutdown avoids, unclean is the fallback
basics
~10 sUnclean leader election lets Kafka elect an out-of-sync (lagging) replica as leader when no in-sync replica is available. It restores availability but can permanently lose committed messages. It defaults to false.
solid answer
~50 sA partition's leader must normally come from the ISR (in-sync replicas) — replicas fully caught up with the previous leader. If every ISR member is offline (e.g., you removed/restarted the wrong brokers), the partition is unavailable. With unclean.leader.election.enable=true, Kafka will then elect a replica that was NOT in sync, making it the new leader; its log is shorter, so all messages the old leader had committed but this replica never received are lost (followers truncate to match). Default is false (since 0.11) — prioritizing durability over availability. During broker turnover this matters because draining/removing brokers can shrink an ISR; if you go too far and lose the last in-sync replica, the partition either stays offline (default) or recovers lossily (if unclean is on). Best practice: keep RF>=3 and min.insync.replicas=2, never decommission a broker while a partition's ISR is at minimum, and keep unclean election disabled except for explicitly loss-tolerant data, re-enabling per-topic only as a deliberate emergency recovery choice.
go deeper
Know it trades availability for durability: electing a lagging replica can lose data; default is off.
Explain ISR, committed messages, and the truncation that loses data; know the per-topic override.
Connect it to broker drain/restart ordering and the RF/min.insync/acks invariants that keep you out of it.
Set fleet durability policy, define emergency-recovery runbooks, and reason about CP vs AP tradeoffs per data class.
**Background — leaders, ISR, and 'committed':** Each partition has one **leader** replica and several **followers**. Followers continuously fetch from the leader; those caught up within `replica.lag.time.max.ms` form the **ISR (in-sync replicas)**. A message is **committed** (and acknowledged to an `acks=all` producer) only once it's replicated to all current ISR members. Normal leader election picks the new leader **only from the ISR**, guaranteeing the new leader already has every committed message — this is a **clean** election. **The failure scenario:** Suppose a partition's ISR shrinks (followers lag, brokers go down, or you drain/remove brokers during maintenance) until **no ISR replica is available** — every in-sync copy is offline. The only surviving replicas are **out-of-sync** (they lagged behind and were kicked from the ISR). Now Kafka faces a choice. **The two policies (`unclean.leader.election.enable`):** - **`false` (default since 0.11):** Kafka refuses to elect an out-of-sync replica. The partition stays **offline/unavailable** until one of its former in-sync replicas comes back. **No data is lost**, but the partition can't serve reads/writes meanwhile. This favors **durability/consistency** (a CP-leaning choice). - **`true`:** Kafka elects one of the surviving **out-of-sync** replicas as the new leader to **restore availability immediately**. But that replica's log is *shorter* — it's missing some messages the old leader had committed. Those messages are **permanently lost**; when the old in-sync replicas return, they **truncate** their logs to match the new (shorter) leader, discarding the extra committed records. This favors **availability** (an AP-leaning choice). **Why it's central to broker lifecycle ops:** Adding and especially **removing/draining** brokers changes ISR membership. If you decommission a broker that holds the last in-sync replica of some partition — or do a rolling restart faster than followers can re-sync — you can drive a partition's available ISR to zero. With the safe default, that partition just goes offline (operational pain, but recoverable). With unclean election on, the cluster silently picks a stale replica and you **lose committed data without an obvious error**. That's why durability invariants (don't remove a broker while a partition is at minimum ISR; verify URP=0; controlled shutdown to drain leaders) exist. **Configuration scope & related settings:** - `unclean.leader.election.enable` is a **cluster default** that can be **overridden per topic**. The safe pattern is global `false`, opting individual loss-tolerant topics into `true` only if truly appropriate. - Pair `replication.factor=3` with `min.insync.replicas=2` and `acks=all`: producers only get acks when >=2 copies hold the data, so a single broker loss never strands the only committed copy — keeping you far from ever *needing* an unclean election. - Re-enabling unclean election on a stuck topic can be used as a **deliberate emergency** to bring an offline partition back, accepting the loss; you should turn it back off afterward. **Common misconceptions / pitfalls:** - 'Unclean election just picks the next replica' — no, it specifically allows an **out-of-sync** one, which is the whole risk. - 'It's on by default' — it has defaulted to **false** since Kafka 0.11. - 'acks=all prevents all loss' — only if `min.insync.replicas` is set appropriately *and* unclean election stays off; with unclean=true, even acks=all data can be truncated away. - Confusing it with **controlled shutdown** — controlled shutdown *prevents* unclean situations by migrating leadership to in-sync replicas before a broker exits; unclean election is the *fallback* when no in-sync replica remains.
- With acks=all set, can unclean leader election still lose committed messages?Yes. acks=all guarantees a write reached all current ISR members, but if every ISR replica later becomes unavailable and unclean election elects an out-of-sync replica, that replica never received the latest committed messages — they're truncated away and lost despite acks=all. min.insync.replicas plus keeping unclean=false is what actually protects you.
- How does controlled shutdown relate to avoiding unclean leader elections?Controlled shutdown proactively migrates a stopping broker's leaderships to in-sync replicas, so partitions retain an in-sync leader and never fall into the no-ISR state. Unclean leader election is only the last-resort fallback that triggers when no in-sync replica remains.
- Is unclean.leader.election.enable a cluster-only setting?No — it's a cluster default that can be overridden per topic. The recommended pattern is global false with selective per-topic true only for explicitly loss-tolerant topics or as a temporary emergency-recovery override.
saying these in an interview costs you the question
- Saying unclean leader election is enabled by default (it's false since 0.11).
- Claiming acks=all alone prevents all data loss regardless of unclean election.
- Describing it as electing 'the next available replica' without noting it's specifically an out-of-sync one.
- Confusing it with controlled shutdown (which prevents the situation rather than causing loss).