Compare time-based and size-based retention in Kafka. Which configs control each, and what happens when both are set?
answer
- time: retention.ms (.minutes/.hours), default 7d
- size: retention.bytes, default -1, per-partition
- both set = OR, stricter wins
- .ms beats .minutes beats .hours
- size = safety valve against spikes
basics
~20 sTime retention (log.retention.ms/.minutes/.hours, default 7 days) removes data older than a duration. Size retention (retention.bytes, default -1 = unlimited) caps bytes per partition. When both are set, a segment is removed if either limit is exceeded.
solid answer
~50 sKafka's delete policy enforces two independent caps per partition. Time-based retention uses log.retention.ms (or .minutes/.hours; .ms takes precedence) with a 7-day default — records older than that become eligible. Size-based retention uses retention.bytes (broker-default log.retention.bytes), default -1 meaning unlimited; it caps total bytes per partition, evicting oldest segments to fit. Crucially retention.bytes is per partition, not per topic — a 10-partition topic with retention.bytes=1GB can hold ~10GB. When both are configured, they are OR'd: a segment is eligible for deletion as soon as either the time limit or the size limit is breached, so the stricter constraint effectively wins. Both are enforced at segment granularity by the retention thread, and the active segment is exempt. Use time retention for predictable replay windows; add size retention as a safety valve against disk blowups during traffic spikes.
go deeper
Know there are two limits — a time one and a size one — with sensible defaults.
Name the exact configs, precedence of .ms/.minutes/.hours, defaults, and that both are OR'd.
Explain per-partition size scope, timestamp-based time eligibility, and the safety-valve design pattern.
Drive capacity planning: pick time for SLA, size for disk safety, and reason about overshoot from segment granularity.
## Two dimensions of retention Under `cleanup.policy=delete`, Kafka bounds each partition's log along two axes. ### Time-based retention - Configs: `log.retention.ms`, `log.retention.minutes`, `log.retention.hours` (broker level) and `retention.ms` (per-topic override). - Precedence: if more than one is set, the **finest unit wins** — `.ms` overrides `.minutes` overrides `.hours`. - Default: 168 hours = **7 days**. - Mechanism: a segment's eligibility is judged by the **largest record timestamp** in that segment (not file mtime). Once that max timestamp is older than `now - retention.ms`, the closed segment can be deleted. - A value of `-1` means **infinite** time retention. ### Size-based retention - Configs: `log.retention.bytes` (broker) / `retention.bytes` (per topic). - Default: `-1` = **unlimited**. - Scope: enforced **per partition**, not per topic. A topic with `retention.bytes=1073741824` (1 GiB) and 12 partitions can use ~12 GiB. - Mechanism: when a partition's total size exceeds the cap, the **oldest segments are deleted first** until the partition is back under the limit (the active segment is still exempt, so a partition can briefly exceed the cap). ## When both are set The two limits are combined with **OR**: a segment is removed when it violates the time limit **or** the size limit. Practically, whichever limit is hit first triggers deletion — the *stricter* one dominates. Example: `retention.ms=604800000` (7d) and `retention.bytes=50GB`. Under light traffic, data leaves after 7 days (time wins). During a spike that fills 50 GB in 2 days, size wins and data is evicted after ~2 days even though it is well under 7 days old. ## Why both matter operationally - **Time only**: predictable replay/compliance window, but disk usage is unbounded if throughput surges. - **Size only**: bounded disk, but the time window silently shrinks under load — bad for consumers that assume N days of replay. - **Both**: time gives the SLA window; size is a **safety valve** preventing a runaway partition from filling the disk. ## Edge cases & gotchas - Because enforcement is per **segment**, real retention overshoots the nominal value by up to one segment's worth of time/size (governed by `segment.bytes`/`segment.ms`). - The active segment never counts toward deletion, so a low-throughput partition that rarely rolls can retain data far past `retention.ms` until the segment closes. - `retention.bytes` being per-partition is a classic capacity-planning trap when partition counts are high.
- Is retention.bytes per topic or per partition?Per partition. Multiply by partition count to size disk: retention.bytes=1GB on a 20-partition topic can use ~20GB.
- If retention.ms=7d and retention.bytes=10GB, and a partition fills 10GB in one day, when is data deleted?After ~1 day — size retention triggers first. The limits are OR'd, so the stricter one wins.
saying these in an interview costs you the question
- Saying retention.bytes is per topic (it is per partition).
- Claiming both limits are AND'd (must violate both) — they are OR'd.
- Forgetting the 7-day default or that -1 means unlimited.
- Saying time is judged by file modification time rather than max record timestamp.