skip to content

Give an example of a topic where you'd deliberately enable unclean leader election, and one where you'd never, and justify each.

level: middleimportance: should knowfreq 40%

answer

  1. cost of lost record vs cost of downtime
  2. enable: metrics, clickstream, logs (loss-tolerant)
  3. never: payments, ledger, audit, exactly-once
  4. per-topic override, cluster default false
  5. pair safe topics with RF3/MISR2/acks=all

basics

~20 s

Enable it for high-volume, loss-tolerant streams like metrics or clickstream logs where staying available matters more than a few lost records. Never enable it for financial transactions, payments, or audit logs where losing a committed record is unacceptable — keep those fail-closed.

solid answer

~40 s

The choice follows the value of an individual record versus the cost of downtime. Enable unclean.leader.election.enable=true on topics carrying high-volume, individually-cheap, ideally reconstructable data — application metrics, clickstream/telemetry, verbose logs — where a brief partition outage hurts more than losing a handful of records, and downstream systems already tolerate gaps. Keep it false (the default) on topics where every committed record matters: payment/order events, financial ledgers, audit trails, anything used as a system of record or feeding exactly-once pipelines. For those, you'd rather the partition go offline and page an operator than silently drop committed data. Because the setting is overridable per topic, you typically keep the cluster default false and selectively flip specific low-value topics to true, rather than making a single cluster-wide choice.

go deeper

for a junior

Match the setting to data value: on for cheap/loss-tolerant data, off for important data.

for a middle

Give concrete topic examples both ways and justify by record value vs downtime cost.

for a senior

Recommend per-topic override pattern with default-false and pairing safe topics with RF3/MISR2/acks=all.

for a principal

Define org-wide topic-classification policy and the placement/runbook strategy that backs each class.

## The deciding question For each topic, ask: **what is the cost of losing one committed record, versus the cost of the partition being unavailable until an in-sync replica returns?** Unclean leader election trades the former to avoid the latter. So your answer depends entirely on the data's value and your downstream tolerance. ## Enable (unclean=true) — examples - **Application metrics / monitoring streams**: a momentary gap in CPU or request-rate samples is harmless; you'd much rather keep ingesting than block dashboards and alerting. - **Clickstream / telemetry / behavioral events** feeding analytics: aggregate trends survive a few lost events; availability of the pipeline is what matters. - **Verbose application logs**: high volume, low per-record value, often best-effort already. Common thread: **high throughput, low per-record value, loss-tolerant or reconstructable, and consumers already cope with gaps.** Availability beats perfect completeness. ## Never enable (keep unclean=false) — examples - **Payments / order / transaction events**: losing a committed 'payment captured' record can mean real money or correctness bugs. - **Financial ledgers and system-of-record topics**: these are the source of truth; silent loss corrupts downstream state irrecoverably. - **Audit / compliance logs**: regulatory requirements forbid dropping committed records. - **Topics feeding exactly-once / transactional pipelines**: unclean election breaks the delivery guarantee. For these you accept that a total-ISR-loss event makes the partition **offline** and pages a human, because correctness outranks availability. ## How to apply it Because `unclean.leader.election.enable` is **overridable per topic**, the recommended pattern is: keep the **cluster default false**, then explicitly set `true` only on the specific low-value topics that want it. This way the safe behavior is the default and enabling fail-open is a conscious, auditable per-topic decision. ## Reinforcing the safe topics For the never-enable topics, pair unclean=false with **RF=3, min.insync.replicas=2, acks=all**, and **rack/AZ-aware replica placement** so all ISR members rarely die together — minimizing how often you even hit the offline state.

  • Should this be a cluster-wide setting or per-topic?
    Per-topic. Keep the cluster default false and override specific low-value, availability-critical topics to true, so the safe behavior is the default and each exception is a deliberate, auditable choice.
  • For a topic you keep unclean=false, what else reduces the chance it ever goes offline?
    RF=3 with min.insync.replicas=2 and acks=all, plus rack/AZ-aware replica placement (broker.rack) so all ISR members don't fail simultaneously, and fast broker recovery so waiting is viable.

saying these in an interview costs you the question

  • Enabling unclean election cluster-wide 'for availability' without considering which topics hold critical data.
  • Enabling it on payment/ledger/audit topics — these must fail closed.
  • Assuming the choice is all-or-nothing rather than a per-topic override.
  • Justifying it purely by throughput rather than by per-record value and downstream loss tolerance.

context