skip to content

Why does retention act on closed segments only, and how do segment.bytes and segment.ms affect how quickly data is actually deleted?

level: seniorimportance: should knowfreq 52%

answer

  1. delete whole files = O(1) unlink
  2. active segment never expires
  3. roll on segment.bytes (1GiB) or segment.ms (7d)
  4. time eligibility = max timestamp
  5. low throughput → lower segment.ms

basics

~20 s

Kafka deletes whole segment files, and only after a segment is closed (rolled) and fully past the limit. The active segment is never touched. segment.bytes and segment.ms control when segments roll, so they set how late data can actually be deleted versus its nominal retention.

solid answer

~50 s

Retention is applied at segment granularity because deleting an entire file is cheap and the log is immutable. A segment becomes deletable only after it closes (rolls) and every record in it is past the retention threshold — eligibility uses the segment's max timestamp for time, or total partition size for bytes. The currently-written active segment is always exempt. segment.bytes (default 1 GiB) and segment.ms (default 7 days) govern when the active segment rolls into a closed one. So real retention can overshoot the nominal value by roughly one segment: a low-throughput partition that takes days to fill 1 GiB, with segment.ms unset/large, can hold data well past retention.ms because the segment hasn't rolled. To make retention tight on low-volume topics, lower segment.ms (e.g. to retention.ms or a fraction of it) and/or segment.bytes so segments close and expire promptly.

go deeper

for a junior

Just know deletion is by whole files and the active one is safe.

for a middle

Explain rolling via segment.bytes/segment.ms and that closed segments are the deletion unit.

for a senior

Connect low throughput + large segments to retention overshoot and prescribe segment.ms tuning with trade-offs.

for a principal

Reason about segment sizing across the cluster: file-handle limits, recovery time, index memory vs retention tightness, and timestamp-skew hazards.

## Why segments, not records A Kafka partition log is an **append-only, immutable** sequence of records stored in fixed **segment files**. Deleting an arbitrary record from the middle would require rewriting files and shifting offsets — expensive and contrary to the append-only design. Instead Kafka deletes data by **unlinking whole segment files**, an O(1) filesystem operation. This is why retention granularity equals segment granularity. ## Segment lifecycle - The **active segment** is the one currently being appended to. It is **never** eligible for retention deletion. - The segment **rolls** (closes; a new active segment starts) when either: - it reaches `segment.bytes` (default **1 GiB**), or - `segment.ms` (default **7 days**) elapses since the segment was created. - Once rolled, the now-**closed** segment becomes a candidate the retention thread can evaluate. ## Eligibility check on a closed segment - **Time**: the segment is deletable when its **largest record timestamp** is older than `now - retention.ms`. Kafka uses the max timestamp so a single recent record keeps the whole segment alive. - **Size**: when the partition's total bytes exceed `retention.bytes`, the **oldest** closed segments are removed first until under the cap. The retention thread (`log.retention.check.interval.ms`, default 5 minutes) periodically performs these checks and renames/deletes (`log.segment.delete.delay.ms`, default 60 s, before the file is actually removed). ## How segment sizing controls real deletion latency Because only closed segments expire, the effective retention is `nominal retention + time-until-the-segment-rolls-and-the-checker-runs`. Two regimes: - **High throughput**: segments fill to `segment.bytes` quickly and roll often, so overshoot is small. - **Low throughput**: a 1 GiB segment may take days/weeks to fill. If `segment.ms` is large/unset, the active segment doesn't roll, so records sit in it far past `retention.ms` — retention appears "broken". ### Tuning for tight retention - Lower `segment.ms` (e.g. equal to or a fraction of `retention.ms`) so the active segment rolls on a time basis even when it isn't full. - Lower `segment.bytes` so segments close sooner. Trade-off: many small segments increase **open file handles**, index memory, and broker overhead, and slow recovery — so don't make segments tiny without reason. ## Edge cases - Setting `retention.ms` very low does little if segments never roll — you must also address segment rolling. - Producers writing records with skewed timestamps (e.g. backfills with old timestamps) can make a segment instantly eligible or, conversely, one future-dated record can pin a segment. - `log.retention.check.interval.ms` bounds how promptly the checker reacts; deletion is never instantaneous.

  • A low-volume topic has retention.ms=1h but data sticks around for days. Why, and how do you fix it?
    Its 1 GiB active segment never fills and segment.ms is large, so the segment doesn't roll and nothing becomes deletable. Lower segment.ms (e.g. to ~1h) and/or segment.bytes so segments close and expire.
  • Which timestamp does time-based retention use to judge a segment?
    The segment's largest (max) record timestamp. One recent record keeps the entire segment alive.
  • What's the downside of making segments very small?
    More open file handles, larger index footprint, more frequent rolling/checks, slower log recovery — broker overhead grows.

saying these in an interview costs you the question

  • Claiming retention deletes the active segment.
  • Saying retention.ms alone guarantees timely deletion regardless of segment rolling.
  • Using file mtime instead of max record timestamp for time eligibility.
  • Recommending tiny segments everywhere without noting file-handle/index overhead.

context