skip to content

Unclean Leader Election

Whether to promote an out-of-sync replica when every in-sync one is gone, and the data loss that follows. The canonical availability-versus-durability trade-off question.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is unclean leader election in Kafka, and what does the unclean.leader.election.enable setting control?

level: juniorimportance: must knowfreq 70%

answer

  1. ISR = caught-up replicas
  2. all ISR down -> wait vs promote lagging
  3. false = durability, true = availability
  4. per-topic override
  5. committed records can be lost

basics

~20 s

Unclean leader election lets Kafka promote a replica that is NOT fully up to date (out of the in-sync set) to leader when no in-sync replica is available. The unclean.leader.election.enable flag turns this on or off. It trades possible data loss for availability.

solid answer

~40 s

Each Kafka partition has one leader and several follower replicas. Replicas that are caught up with the leader form the ISR (in-sync replica set). Normally Kafka only elects a new leader from the ISR, guaranteeing no committed data is lost. Unclean leader election allows promoting an out-of-sync replica (one not in the ISR) to leader when every ISR member is unavailable. The unclean.leader.election.enable config (broker-level default, overridable per topic) controls this. With it false (the modern default), the partition goes offline and stays unavailable until an ISR replica returns — protecting data. With it true, Kafka picks a lagging replica so the partition stays writable, but any records the old leader had that the new leader never received are permanently lost.

go deeper

for a junior

Know the one-line trade-off: lets Kafka promote a not-fully-caught-up replica to keep the partition available, at the risk of losing data.

for a middle

Tie it to ISR and committed records; know the config is false by default and overridable per topic.

for a senior

Frame it as the CP-vs-AP knob; explain exactly when it triggers (all ISR down) and what specifically is lost.

for a principal

Reason about which topic classes warrant true vs false and the operational/runbook implications of either choice.

## Background In Kafka, every topic-partition is replicated across multiple brokers for fault tolerance. One replica is the **leader** (handles all reads and writes); the others are **followers** that continuously fetch records from the leader to stay current. A follower that is sufficiently caught up is part of the **ISR — the In-Sync Replica set**. A record is considered **committed** only once every member of the ISR has it; consumers can only read committed records, and (with acks=all) producers only get acknowledgement after all ISR members store the record. This is what makes a committed record durable. ## What unclean leader election is When a leader fails, Kafka must elect a new leader. A **clean** election picks the new leader from the current ISR — those replicas have every committed record, so nothing is lost. But if **all** ISR replicas are down at once (e.g. the leader and every in-sync follower crash), there is no in-sync replica to promote. At that point Kafka has two choices: 1. **Wait** — keep the partition offline until one of the ISR replicas comes back. No data loss, but the partition is unavailable (producers and consumers for it are blocked). 2. **Unclean leader election** — promote a surviving **out-of-sync** replica (one that had fallen out of the ISR and is therefore behind). The partition becomes available again immediately, but any committed records that the old ISR had and this lagging replica never fetched are gone forever. ## The config `unclean.leader.election.enable` controls which path Kafka takes. - **false** (the default since Kafka 0.11): never elect an out-of-sync replica. Favors **durability/consistency** (CP-leaning). - **true**: allow it. Favors **availability** (AP-leaning). It is a **broker/cluster default** but can be **overridden per topic** (topic config `unclean.leader.election.enable`), so you can keep most topics safe while allowing a specific low-value, availability-critical topic to fail open. ## Why it matters This setting is the concrete knob where Kafka lets you choose your side of the CAP-style trade-off for a partition: when the cluster cannot give you both, do you want the data correct (stay offline) or the system writable (accept loss)?

  • What is the default value of unclean.leader.election.enable in modern Kafka?
    false. Since Kafka 0.11 the default favors durability — the partition stays offline rather than risk losing committed data.
  • Why can't a clean election just always be used?
    A clean election needs at least one ISR member alive. If every in-sync replica is down simultaneously, there is no caught-up replica to promote, so the only options are wait (offline) or elect an out-of-sync replica (unclean).

saying these in an interview costs you the question

  • Saying unclean election causes data loss in normal operation — it only triggers when ALL ISR replicas are unavailable.
  • Thinking the default is true (it has been false since 0.11).
  • Confusing it with electing any follower — clean election from the ISR is normal; unclean specifically means an OUT-of-sync replica.

context

open as a page

Walk through exactly what data loss and log divergence happens when an unclean leader election promotes an out-of-sync replica.

level: seniorimportance: must knowfreq 60%

basics

~20 s

The lagging replica becomes leader at its own (shorter) log end. Records the old leader had committed but this replica never fetched are gone. When the old replica returns, it must truncate its longer log down to the new leader's offset to re-join, throwing away those records.

open as a page

Give an example of a topic where you'd deliberately enable unclean leader election, and one where you'd never, and justify each.

level: middleimportance: should knowfreq 40%

basics

~20 s

Enable it for high-volume, loss-tolerant streams like metrics or clickstream logs where staying available matters more than a few lost records. Never enable it for financial transactions, payments, or audit logs where losing a committed record is unacceptable — keep those fail-closed.

open as a page

How does unclean.leader.election.enable interact with min.insync.replicas and acks to position a topic on the consistency-vs-availability spectrum?

level: seniorimportance: should knowfreq 45%

basics

~20 s

acks=all + min.insync.replicas controls write-time durability (rejecting writes when too few replicas are in sync); unclean.leader.election.enable controls failover behavior when all ISR are gone. Together false + acks=all + min.insync.replicas>=2 gives a CP topic that fails closed; unclean=true makes it fail open with possible loss.

open as a page

All ISR replicas for a partition are dead and the partition is offline because unclean election is disabled. As the on-call operator, how do you decide and how do you force an unclean election to restore availability?

level: principalimportance: should knowfreq 35%

basics

~20 s

First weigh data loss vs downtime: if waiting for an ISR replica to return is acceptable, wait. If availability must be restored now and loss is tolerable, force it — temporarily set unclean.leader.election.enable=true (per-topic), or run kafka-leader-election.sh with --election-type UNCLEAN, then revert. Document the loss.

open as a page