What configuration controls when a segment rolls (closes and a new one opens), and what are the trade-offs of tuning it?
answer
- segment.bytes 1 GiB, segment.ms 7 days
- whichever fires first rolls
- broker: log.segment.bytes / log.roll.ms
- small = finer retention + more FDs
- segment.jitter.ms desyncs time rolls
basics
~20 sA segment rolls when it hits segment.bytes (default 1 GiB) or segment.ms (default 7 days) — whichever comes first. Smaller values roll more often, giving finer retention but more files and open handles; larger values mean fewer, bigger files.
solid answer
~50 sSegment rolling is driven mainly by two topic-level configs: `segment.bytes` (default 1073741824 = 1 GiB) closes the active segment once it reaches that many bytes, and `segment.ms` (default 604800000 = 7 days) closes it that long after creation even if it's not full. Broker defaults are `log.segment.bytes` and `log.roll.ms`/`log.roll.hours`. A roll can also be forced when the index is full or a timestamp would break index ordering. Trade-offs: small segments give finer retention/compaction granularity and free disk sooner, but inflate the number of files and open file descriptors and add roll overhead; large segments reduce file count and overhead but coarsen retention (you can't delete a partial segment) and delay reclaiming space. On low-throughput partitions, segment.ms is what eventually rolls an idle segment so retention can act. There's also `segment.jitter.ms` to desynchronize time-based rolls across partitions and avoid thundering-herd rolls.
go deeper
Know segment.bytes and segment.ms exist and roll the segment.
State the defaults (1 GiB / 7 days) and that whichever fires first wins.
Reason about the file-count vs retention-granularity trade-off and the retention interaction on idle partitions.
Set sizing policy across topics, account for FD limits, jitter, and compaction-unit size at fleet scale.
## What "rolling" means The **active segment** is the only file currently open for writes. **Rolling** = Kafka closes the current active segment (making it read-only/closed) and opens a fresh one with a new base offset. Retention and compaction only ever operate on closed segments, so roll timing directly controls how quickly old data becomes eligible for cleanup. ## The controlling configs Topic-level (per topic, override broker defaults): - **`segment.bytes`** — default `1073741824` (1 GiB). When the active segment's size reaches this, it rolls. This is the dominant trigger on busy partitions. - **`segment.ms`** — default `604800000` (7 days). Maximum age of an active segment; after this long since it was created it rolls even if far below `segment.bytes`. This is the dominant trigger on idle/low-traffic partitions. - **`segment.index.bytes`** — default 10 MiB; caps the .index/.timeindex size. If the index fills, the segment rolls early. - **`segment.jitter.ms`** — random subtracted from the time-based roll so many partitions created together don't all roll at the same instant (avoids a rolling/IO spike). Broker-level equivalents (defaults applied when topic doesn't override): `log.segment.bytes`, `log.roll.ms` / `log.roll.hours`, `log.index.size.max.bytes`, `log.roll.jitter.ms`. A roll can also be forced internally when appending a record whose timestamp/offset can't be represented relative to the current base (rare), or on broker restart for the previously active segment. ## Trade-offs of small vs large segments **Smaller segments (lower segment.bytes / segment.ms):** - + Finer retention granularity — space is reclaimed sooner because a whole segment becomes eligible faster. - + Compaction can run on more, smaller units. - - More files per partition -> more open file descriptors (watch the OS `ulimit -n`). - - More frequent rolls = more metadata/index churn and more page-cache fragmentation. **Larger segments:** - + Fewer files, fewer descriptors, less roll overhead. - - Coarser retention: you cannot delete part of a segment, so a single huge segment can pin a lot of data past `retention.ms` until it rolls. - - Slower space reclamation and potentially long compaction units. ## Interaction with retention Retention (`retention.ms` / `retention.bytes`) is evaluated per CLOSED segment. If `segment.ms` is large and a partition is idle, the active segment never rolls, so even very old records stay on disk. A common production tuning is to lower `segment.ms` (e.g. to a few hours) on low-throughput topics so retention can actually free old data on schedule. ## Edge cases - Setting `segment.ms` very low on a high-throughput topic creates an explosion of tiny segments and can exhaust file handles. - `segment.bytes` smaller than a single record/batch is invalid; a segment must hold at least one batch. - `segment.jitter.ms` only applies to time-based rolls, not size-based.
- A topic has retention.ms = 1 hour but you still see 2-day-old data on an idle partition. Why, and how do you fix it?Retention only deletes closed segments. With default segment.ms = 7 days and no new writes, the active segment never rolls, so old records can't be cleaned. Lower segment.ms (e.g. to 1 hour) so the idle segment rolls and retention can act.
- What's the risk of setting segment.ms to a few seconds on a busy topic?It rolls constantly, creating huge numbers of tiny segment files, exhausting open file descriptors and adding index/metadata overhead — potentially crashing the broker on 'Too many open files'.
saying these in an interview costs you the question
- Saying retention.ms alone deletes data — it acts on closed segments, gated by roll timing.
- Claiming only segment.bytes matters — segment.ms is what rolls idle partitions.
- Confusing segment.bytes (per-segment) with retention.bytes (per-partition total).
- Assuming smaller segments are always better — they cost file descriptors and overhead.