skip to content

All ISR replicas for a partition are dead and the partition is offline because unclean election is disabled. As the on-call operator, how do you decide and how do you force an unclean election to restore availability?

level: principalimportance: should knowfreq 35%

answer

  1. first choice: wait if ISR broker can return (lossless)
  2. force only if loss tolerable / brokers gone
  3. kafka-configs topic override OR kafka-leader-election --election-type UNCLEAN
  4. estimate loss from survivor LEO vs HW
  5. revert flag, notify consumers, document

basics

~20 s

First weigh data loss vs downtime: if waiting for an ISR replica to return is acceptable, wait. If availability must be restored now and loss is tolerable, force it — temporarily set unclean.leader.election.enable=true (per-topic), or run kafka-leader-election.sh with --election-type UNCLEAN, then revert. Document the loss.

solid answer

~50 s

When a partition is offline with an empty ISR, you have a business decision, not just a technical one: is the downtime worse than losing the committed records that only the dead ISR replicas hold? If an ISR broker can come back soon, wait — that's lossless. If not and availability is critical, force an unclean election. Two ways: (1) override the topic config `unclean.leader.election.enable=true` (e.g. kafka-configs.sh --alter --entity-type topics), which lets the controller promote a surviving out-of-sync replica, then set it back to false afterward; or (2) explicitly trigger it with kafka-leader-election.sh --election-type UNCLEAN for the specific topic-partition. Before doing so, confirm which replicas are alive and their LEO so you know roughly how much data you'll lose. After recovery, document the loss, alert downstream consumers (offsets may now be invalid / OffsetOutOfRange), and verify returning replicas truncate and re-sync. The decision should ideally be pre-agreed per topic class in a runbook.

go deeper

for a junior

Recognize that forcing unclean election trades data loss for getting the partition back online.

for a middle

Know at least one mechanism (topic config override or kafka-leader-election --election-type UNCLEAN) and to revert it.

for a senior

Run the full decision: assess blast radius, choose wait vs force, execute surgically, reconcile consumers.

for a principal

Own the per-topic policy/runbook, rack-aware replica placement to prevent it, and the data-integrity contract with downstream teams.

## The situation The partition is **offline**: every replica in its ISR is down, and because `unclean.leader.election.enable=false` Kafka refuses to promote a lagging survivor. Producers and consumers for that partition are blocked. Other partitions are fine. ## Step 1 — Decide: data loss vs downtime This is fundamentally a **business trade-off**, which is why Kafka leaves it to the operator: - **Wait (lossless):** if any ISR broker can be recovered quickly (reboot, disk reattach, network heal), bring it back. It still holds every committed record, so a normal clean election restores the partition with **zero loss**. Always the first choice when feasible. - **Force unclean (lossy):** if the ISR brokers are gone for good (lost disks) or the downtime cost outweighs the data value, promote a surviving out-of-sync replica and accept the loss. Key inputs to the decision: how stale is the surviving replica (its LEO vs the last known HW), how much downtime is accumulating, whether the topic carries reconstructable/idempotent data, and any compliance constraints on losing records. ## Step 2 — Assess the blast radius Before forcing it, inspect replica state — `kafka-topics.sh --describe` shows replicas and ISR; check each surviving broker's log end offset for the partition. The gap between the last committed HW and the survivor's LEO approximates the records you'll lose. ## Step 3 — Force the unclean election Two supported mechanisms: 1. **Config override (broad):** `kafka-configs.sh --bootstrap-server ... --alter --entity-type topics --entity-name <topic> --add-config unclean.leader.election.enable=true` The controller then promotes an out-of-sync replica for any offline partition of that topic. **Revert to false** once recovered so you don't silently stay in fail-open mode. 2. **Targeted election (surgical):** `kafka-leader-election.sh --bootstrap-server ... --election-type UNCLEAN --topic <topic> --partition <n>` Triggers an unclean election for just that partition without leaving the config flipped. Preferred when you want to scope the loss to one partition. ## Step 4 — Recover and reconcile - The promoted replica becomes leader at a new leader epoch; returning replicas perform leader-epoch truncation and re-sync to it. - **Notify downstream consumers**: stored offsets may now exceed the new LEO → `OffsetOutOfRange`, handled per `auto.offset.reset`. Some records they already processed no longer exist (phantom reads) — application teams may need to reconcile or replay from a source of truth. - **Document** the incident: which partition, estimated records lost, time window, decision rationale. ## Step 5 — Prevent recurrence The real fix is upstream: spread replicas across racks/AZs (`broker.rack` + rack-aware assignment) so all ISR members rarely die together, and pre-decide per-topic policy in a runbook so on-call isn't improvising under pressure. For topics where any loss is unacceptable, the answer is to **wait**, full stop, and invest in faster broker recovery.

  • What's the difference between flipping the topic config and using kafka-leader-election.sh --election-type UNCLEAN?
    The config override enables unclean election for the whole topic until you revert it, affecting all its offline partitions. kafka-leader-election.sh --election-type UNCLEAN triggers it for specific partition(s) without leaving a fail-open config set, so it's more surgical.
  • Why is forcing unclean election an operator decision rather than automatic?
    Because it's a business trade-off: the value of the lost committed records versus the cost of continued downtime. Only the owning team can weigh that, so Kafka defaults to false and leaves the override to humans.
  • How do you reduce the chance of ever facing this?
    Spread replicas across racks/AZs with broker.rack and rack-aware assignment so all ISR members don't fail together, keep RF=3/min.insync.replicas=2/acks=all, and invest in fast broker/disk recovery so waiting is viable.

saying these in an interview costs you the question

  • Always forcing unclean election to clear the alert without weighing data loss — it's a deliberate trade-off, not a reflex.
  • Forgetting to revert unclean.leader.election.enable to false after recovery, silently leaving the topic fail-open.
  • Ignoring downstream consumers — offsets can become invalid (OffsetOutOfRange) and already-read records may vanish.
  • Assuming the data 'comes back' when the old leader restarts — it truncates and discards the diverged suffix.

context