skip to content

What does configuring multiple log.dirs (JBOD) give you on a Kafka broker, and how are partitions placed across the directories?

level: middleimportance: should knowfreq 55%

answer

  1. log.dirs = comma-separated, JBOD
  2. one partition = one dir (no striping)
  3. placement = fewest-partitions dir (count-based)
  4. disk fail -> dir offline, broker survives
  5. replication across brokers replaces RAID

basics

~20 s

Setting multiple paths in log.dirs (JBOD = Just a Bunch Of Disks) lets one broker spread partition data across several independent disks for more total capacity and I/O. Kafka assigns each new partition to the directory with the fewest partitions.

solid answer

~40 s

`log.dirs` accepts a comma-separated list of directories, typically one per physical disk — this is JBOD ('Just a Bunch Of Disks'), used instead of RAID. Each partition's log lives entirely in exactly one of these directories; Kafka does not stripe a single partition across disks. When a new partition replica is created, the broker places it in the log directory that currently holds the fewest partitions (a count-based round-robin), spreading load. JBOD gives more aggregate capacity and parallel I/O than a single disk, and avoids RAID write-amplification, but a single disk failure takes down only the partitions on that disk rather than the whole broker — in KRaft/modern Kafka the broker stays up and the affected replicas are marked offline. Recovery threads (`num.recovery.threads.per.data.dir`) operate per directory, so more dirs can parallelize startup recovery.

go deeper

for a junior

Know log.dirs can list multiple directories (one per disk) and that each partition lives in one of them.

for a middle

Explain JBOD vs RAID, count-based placement, and graceful per-disk failure handling.

for a senior

Reason about count-based imbalance, reassigning replicas between log dirs, and recovery parallelism per dir.

for a principal

Decide JBOD vs RAID at fleet scale, set disk-uniformity and rebalancing policy, and weigh replication factor against disk redundancy.

## What log.dirs is `log.dirs` is a broker configuration that takes a **comma-separated list of filesystem directories** where Kafka stores its partition log data. (The singular `log.dir` is the older single-directory form; `log.dirs` overrides it.) Pointing each directory at a separate physical disk is the **JBOD** ('Just a Bunch Of Disks') layout. ``` log.dirs=/data/disk1/kafka,/data/disk2/kafka,/data/disk3/kafka ``` ## JBOD vs RAID - **RAID** combines multiple disks into one logical volume at the OS/hardware level. RAID-10 gives redundancy but halves usable capacity and adds write amplification; RAID-0 gives no redundancy. - **JBOD** exposes each disk to Kafka separately and lets *Kafka* handle distribution and redundancy. Kafka already replicates partitions across *brokers* (via the replication factor), so disk-level RAID redundancy is often redundant. JBOD therefore gives full capacity, parallel I/O across spindles/SSDs, and avoids RAID write penalties. ## How partitions are placed A key rule: **one partition's log lives entirely inside one log directory.** Kafka never splits a single partition's segments across multiple directories — there is no striping at the Kafka layer. When a broker must create a new partition replica, it chooses the target directory using a **count-based balancing rule: pick the log directory currently hosting the fewest partitions.** This spreads partition count evenly but is *count*-based, not *size*- or *throughput*-based — so a few very hot or very large partitions can still unbalance disks. (Tools and Cruise Control can rebalance data, and `kafka-reassign-partitions` supports moving replicas between log directories on the same broker.) ## Failure semantics If one disk fails: - Only the partitions stored in that directory are affected. - In modern Kafka the broker does **not** crash; it marks the affected replicas/log dir **offline** and continues serving the partitions on the healthy disks. Leadership for the lost partitions moves to in-sync replicas on other brokers. - (In older versions, a single failed log dir could take the whole broker down — JBOD offline-dir tolerance, KIP-112/113, made this graceful.) ## Interaction with recovery and threads - `num.recovery.threads.per.data.dir` controls how many threads recover logs **per directory** at startup; more log dirs means more parallelism in startup recovery (and in flush at shutdown). - Each directory independently maintains its own `recovery-point-offset-checkpoint` and `replication-offset-checkpoint` files. ## Edge cases - **Imbalance:** count-based placement ignores size/throughput; monitor per-disk utilization. - **Adding a disk:** new partitions flow to the emptier new disk, but existing partitions don't auto-migrate — you reassign them. - **Mixed disk sizes:** count-based placement can overfill a smaller disk; keep disks uniform. ## Bottom line JBOD = one broker, many independent disks, one partition per dir, count-balanced placement, graceful per-disk failure — leaning on Kafka's cross-broker replication instead of RAID for redundancy.

  • Why is count-based partition placement sometimes insufficient for balancing disks?
    Because it balances the *number* of partitions, not their size or throughput. A few large or high-traffic partitions can fill or saturate one disk while others sit idle, requiring size/load-aware rebalancing (e.g., kafka-reassign-partitions between log dirs, or Cruise Control).
  • If you use JBOD with replication factor 3, do you still need RAID for redundancy?
    Usually no. Kafka already replicates each partition to multiple brokers, so a disk failure is covered by the replicas on other brokers. RAID would add write amplification and cost without much benefit; JBOD plus replication is the common production choice.

saying these in an interview costs you the question

  • Saying Kafka stripes a single partition across multiple log.dirs — it does not; one partition lives in one directory.
  • Claiming a single disk failure always crashes the whole broker — modern Kafka marks only that dir offline.
  • Assuming JBOD placement balances by data size or load — it balances by partition count.

context