skip to content

Explain the data-loss window that exists with Kafka's default flush settings, and what conditions are required to actually lose acknowledged data.

level: seniorimportance: should knowfreq 40%

answer

  1. ack'd but only in RAM across ISR
  2. lose data only on correlated abrupt power loss
  3. single-broker loss is safe
  4. window bounded by OS writeback / flush.ms
  5. mitigate with cross-AZ rack awareness

basics

~20 s

Because Kafka acknowledges writes when replicas have the data in memory (not on disk), acknowledged records can be lost only if every in-sync replica loses power at nearly the same time before any of them flushes to disk. That's the correlated power-loss window.

solid answer

~50 s

With default flush settings, an acks=all write is acknowledged once all in-sync replicas hold the record in their OS page cache (RAM) — before it's flushed to disk. So there's a window where acknowledged data exists only in volatile memory across the ISR. To actually lose that data, you need a **correlated failure**: every ISR member must suffer an abrupt power loss (or kernel crash) before any of them flushes those pages to disk. A normal broker restart or single-broker power loss doesn't lose data, because surviving ISR members still hold the record and one becomes the new leader. The window is bounded by how long pages stay dirty (OS writeback interval, plus flush.ms/flush.messages if set). You shrink the risk by spreading replicas across racks/AZs with independent power, raising min.insync.replicas, and optionally tightening flush settings at a throughput cost. This residual window is the deliberate trade-off Kafka makes for performance.

go deeper

for a junior

Know acknowledged data can be lost only in a rare case where all replicas lose power at once before flushing.

for a middle

Explain why single-broker loss is safe and what makes the correlated case different.

for a senior

Bound the window with OS writeback and flush settings, and prescribe cross-AZ placement plus min.insync.replicas to mitigate.

for a principal

Quantify the residual risk vs. throughput trade-off and set durability SLAs, replica placement, and storage choices for the fleet.

## Setting up the window Kafka's durability model acknowledges an `acks=all` write once **all in-sync replicas (ISR)** have the record — but "have the record" means "appended it to their log, which sits in the **OS page cache** (RAM)." The actual flush to physical disk happens **asynchronously** later (or when `flush.ms`/`flush.messages` force it). Therefore, for some interval after acknowledgement, the only copies of the record may be in volatile RAM on the ISR brokers. ## What does NOT lose data - **Single broker power loss / crash.** The other ISR members still have the record. On recovery, a surviving in-sync replica is elected leader; the record is intact. The crashed broker re-replicates on restart. - **Graceful shutdown/restart.** Kafka flushes its logs on a clean shutdown, so nothing is lost. - **Disk failure on one broker.** Replication covers it. ## What DOES lose acknowledged data All of these must coincide: 1. The record is acknowledged but still only in page cache (not yet flushed) on the relevant brokers. 2. **Every** member of the ISR that holds it suffers an **abrupt** failure — power loss or a hard kernel crash that doesn't flush dirty pages — at essentially the same time. 3. They fail before the OS (or a forced fsync) writes those dirty pages to disk. This is the **correlated power-loss window**. It typically requires a shared-fate event: a whole rack, power distribution unit, or availability zone going dark at once. ## What bounds the window - **OS writeback interval** — kernel tunables (`vm.dirty_expire_centisecs`, `vm.dirty_ratio`, etc.) determine how long pages stay dirty before being flushed. - **flush.ms / flush.messages** — if set, they force fsyncs sooner, shrinking the window at the cost of throughput. - **min.insync.replicas** — a higher value means more independent copies must all fail; combined with cross-AZ placement this makes the correlated event far less likely. ## Mitigations - Spread replicas across **racks and availability zones** with independent power (`broker.rack` + rack-aware assignment) so a single power event can't take out the whole ISR. - Use **replication.factor=3, min.insync.replicas=2, acks=all** as a baseline. - For extreme durability requirements, lower `flush.ms` (accepting throughput loss) or use storage with power-loss protection / battery-backed write cache. ## The philosophical point Kafka could fsync every message to close this window entirely, but the throughput cost is enormous and the residual risk — *all* independent replicas losing power simultaneously before any flush — is rare in a well-distributed cluster. Accepting this small window in exchange for high throughput is an intentional design decision, not a bug.

  • Why doesn't a single broker's sudden power loss cause data loss with acks=all?
    The other in-sync replicas still hold the record in their logs. One of them is elected the new leader, so the acknowledged record survives; the failed broker re-syncs on restart.
  • How does rack/AZ-aware replica placement reduce the correlated-loss risk?
    broker.rack plus rack-aware partition assignment spreads a partition's replicas across independent failure domains (racks/AZs) with separate power, so one power event can't drop every ISR member's page-cache copy at once.

saying these in an interview costs you the question

  • Claiming a single broker power loss loses acknowledged data under acks=all.
  • Saying Kafka can never lose acknowledged data regardless of configuration or failure mode.
  • Ignoring that the loss requires the failure to be abrupt (no flush) AND correlated across the whole ISR.

context