skip to content

Explain how Kafka's append-only log and sequential disk writes make it fast even on spinning disks.

level: middleimportance: should knowfreq 55%

answer

  1. append-only log, immutable segments
  2. sequential = bandwidth, random = seeks/IOPS
  3. writes land as dirty pages, writeback flushes
  4. durability via replication, not per-write fsync
  5. retention deletes whole segments

basics

~20 s

Kafka only appends to the end of partition log files instead of writing in random places. Sequential writes are far faster than random writes (no disk seeks), so even HDDs can sustain high throughput, and writes go through the page cache first.

solid answer

~40 s

Each partition is an immutable, append-only log split into segment files. Producers never update in place; the broker just appends new records to the tail of the active segment. This turns all writes into a single sequential stream per partition, which avoids the costly random seeks that dominate disk latency on HDDs and even matter on SSDs. The OS receives these appends into the page cache as dirty pages and flushes them to disk in large, contiguous, batched writes via background writeback. Combined with producer-side batching and compression, Kafka achieves high write throughput that is bounded by sequential disk bandwidth rather than IOPS. Reads of recent data are similarly sequential and usually served from page cache. The append-only design also simplifies replication and retention (delete/compact whole old segments).

go deeper

for a junior

Know Kafka appends to the end of a log and that sequential writes are faster than random.

for a middle

Explain segments, dirty pages, writeback, and that durability is replication-based by default.

for a senior

Quantify the random-vs-sequential gap, relate to producer batching/compression and partition-count effects.

for a principal

Reason about disk-layout, partition density, and the durability/throughput trade-off across the cluster.

## Append-only logs A Kafka topic partition is a **log**: an ordered, immutable sequence of records that you can only **append** to. Physically it is a directory of **segment files** (e.g. `00000000000000000000.log`), each holding a contiguous range of offsets. When the active segment reaches `segment.bytes` (default 1 GB) or `segment.ms`, it is rolled and a new active segment begins. ## Why sequential beats random Disk performance has two very different regimes: - **Random I/O**: many small reads/writes scattered across the disk. On a spinning HDD each one may require a head **seek** (milliseconds), so throughput is limited by IOPS and is very low for small ops. - **Sequential I/O**: reading/writing contiguous bytes. The head stays put (HDD) or the controller streams pages (SSD), so throughput approaches the raw bandwidth of the device — orders of magnitude higher. Because Kafka **only ever appends**, every write to a partition is sequential. There is no in-place update, no B-tree to rebalance, no random scatter. This is the central reason a Kafka broker can saturate disk bandwidth even on cheap commodity disks. ## How a write actually flows 1. Producer batches records and sends them to the partition leader. 2. The broker appends the batch to the active segment **file**. This write lands in the **page cache** as **dirty pages** (modified pages not yet on disk). 3. The kernel's **writeback** mechanism flushes dirty pages to disk asynchronously in large, contiguous chunks (tunable via `vm.dirty_ratio`, `vm.dirty_background_ratio`). 4. Durability across the cluster comes from **replication** (acks + in-sync replicas), not from forcing an fsync on every write — Kafka by default does **not** fsync per message and relies on replicas to survive a single-node crash. ## Reinforcing optimizations - **Producer batching** (`batch.size`, `linger.ms`) groups many records into one larger sequential write. - **Compression** (`compression.type`) shrinks the bytes written. - **Segment + index layout** keeps related data physically contiguous, so consumer reads are also sequential. - **Whole-segment retention**: old data is deleted by removing entire segment files, an O(1) operation, never row-by-row deletes. ## Edge cases - Many partitions on one disk means several interleaved sequential streams, which can look partly random at the device — hence per-broker partition-count guidance. - If consumers read very old (cold) data not in page cache, those reads still hit disk, though laid out sequentially per segment. - On SSDs, sequential vs random matters less for latency but write amplification and batching still favor the append-only design.

  • Does Kafka fsync each write to guarantee durability?
    No. By default Kafka does not fsync per record; it relies on replication across in-sync replicas to survive node failure. Flush is governed by flush.messages/flush.ms and OS writeback, with fsync being a tunable trade-off.
  • How does having very many partitions per broker affect the 'sequential' assumption?
    Each partition is sequential on its own, but many active partitions interleave writes/reads on the same disk, which the device sees as partly random, reducing the sequential benefit and motivating partition-count and disk-layout planning.

saying these in an interview costs you the question

  • Saying Kafka updates records in place.
  • Claiming durability comes from a per-message fsync by default.
  • Believing sequential I/O only matters on SSDs (it matters most on HDDs).
  • Thinking retention deletes records individually rather than whole segments.

context