skip to content

How do you estimate disk capacity for a Kafka cluster from throughput metrics and retention settings?

level: middleimportance: must knowfreq 60%

answer

  1. BytesIn × retentionSec × RF
  2. retention.bytes is PER PARTITION
  3. delete on time OR size, whichever first
  4. segments delete whole, add headroom
  5. compacted topics size by distinct keys

basics

~10 s

Disk needed = ingress bytes/sec × retention seconds × replication factor, summed across topics, plus headroom. Use BytesInPerSec for the rate and retention.ms (or retention.bytes) for how long data is kept.

solid answer

~40 s

Start from the steady-state ingress rate per topic (BytesInPerSec, which is already the compressed on-disk size). For a topic: storage = BytesInPerSec × retention.ms/1000 × replication.factor. Sum across topics for cluster total, then divide by broker count for per-broker need, plus 30-40% headroom for spikes, compaction tombstones, index files, and rebalancing. retention.bytes caps per-partition size independently of time and acts as a hard ceiling — Kafka deletes when either the time OR size limit is hit. Replication factor multiplies everything because every replica stores a full copy. Also account for segment granularity: a topic only frees space a closed segment at a time (segment.bytes / segment.ms), so actual usage slightly exceeds the theoretical retention window. Don't forget that compacted topics retain at least one record per key indefinitely.

go deeper

for a junior

Recall the basic formula: rate × retention × replication factor.

for a middle

Apply it per topic, sum to cluster, add headroom, and know retention.bytes is per partition.

for a senior

Account for segment granularity, indexes, compaction, and peak-vs-average sizing.

for a principal

Build a capacity model across topics with differing policies, headroom budgets, and growth, and tie disk-full to URP/availability risk.

**The core formula.** For one topic: ``` bytes_on_disk = BytesInPerSec × retention_seconds × replication_factor ``` Each term: - **BytesInPerSec** — measured ingress rate. Because Kafka stores batches as received, this is already the compressed on-disk size, so no extra compression factor is needed. Use a representative steady-state or peak value depending on whether you're sizing for average or worst case. - **retention_seconds** — from `retention.ms` (time-based). Data older than this becomes eligible for deletion. The alternative/complement is `retention.bytes`, a per-partition size cap; whichever limit is reached first triggers deletion. - **replication.factor** — every replica is a full copy on a different broker, so RF=3 triples storage. Sum across all topics for cluster-wide raw bytes, then divide by the number of brokers for per-broker storage (assuming balanced partition placement). **Headroom — why theoretical isn't enough:** - **Segment granularity.** A partition's log is split into segments (`segment.bytes`, default 1 GiB, or rolled by `segment.ms`). Kafka deletes whole *closed* segments, never partial ones, so the active segment plus the not-yet-expired tail mean real usage runs above the clean retention window. - **Indexes.** Each segment has `.index`, `.timeindex`, and `.txnindex` files — small but nonzero overhead. - **Rebalancing / reassignment.** Moving partitions temporarily stores two copies until the move completes. - **Spikes.** Producers burst; size for peak BytesIn, not just the mean. - Rule of thumb: keep 30-40% free disk, and alert well before the OS / Kafka log dir fills (a full disk can take a broker offline and cause under-replicated partitions). **retention.bytes vs retention.ms.** `retention.bytes` is *per partition*, not per topic. A topic with 50 partitions and `retention.bytes=1GB` can hold up to 50 GB per replica. Mixing the two: deletion fires when EITHER condition is met, so a size cap protects you from an ingest spike blowing the disk even if the time window hasn't elapsed. **Compaction.** For `cleanup.policy=compact`, time/size retention doesn't bound the topic the same way — compaction keeps the latest value per key indefinitely (unless also `delete`). Size such topics by *distinct key count × average record size*, not by ingest rate × time. **Worked example.** Topic at 50 MB/s BytesIn, 7-day retention, RF=3: 50e6 × 604800 × 3 ≈ 90.7 TB cluster-wide. On a 10-broker cluster ≈ 9.1 TB/broker; with 35% headroom provision ~12.3 TB usable per broker for this topic's share. **Gotcha.** People forget RF and size for one copy, then run out of disk at 1/3 the expected lifetime. They also forget that BytesInPerSec is already compressed and double-apply a compression ratio.

  • Is retention.bytes per topic or per partition?
    Per partition. A topic's max size from retention.bytes is retention.bytes × partition_count × replication_factor, which surprises people who treat it as a topic-level cap.
  • Why might actual disk usage exceed BytesIn × retention × RF?
    Segments are deleted whole only after they close, so the not-yet-expired tail plus the active segment, index files, and in-flight reassignments all add to the theoretical figure — hence headroom.

saying these in an interview costs you the question

  • Forgetting to multiply by replication factor.
  • Treating retention.bytes as a per-topic cap instead of per-partition.
  • Applying a compression ratio to BytesInPerSec (it's already compressed).
  • Assuming a compacted topic is bounded by retention.ms the same way a delete topic is.

context