Design-wise, what risks does delete retention create for consumers and operations, and how do you mitigate retention deleting data a consumer still needs?
answer
- retention ignores consumer progress
- lapped consumer → OffsetOutOfRange
- auto.offset.reset=earliest silently skips
- size retention ≥ worst-case lag
- tiered storage (KIP-405) for long replay
basics
~20 sIf a consumer lags behind retention, Kafka can delete records before they're read, causing OffsetOutOfRange and data loss for that consumer. Mitigate by sizing retention above worst-case lag, monitoring lag, alerting, and using size limits as a disk safety valve.
solid answer
~40 sDelete retention is purely time/size-driven and ignores whether consumers have read the data. So a slow or stopped consumer whose committed offset falls behind the log-start offset (because retention deleted those segments) hits OffsetOutOfRange; with auto.offset.reset=earliest it silently jumps forward, skipping unread records — effective data loss. Mitigations: size retention.ms above worst-case consumer downtime/lag (e.g. a multi-day window for batch consumers); monitor consumer-group lag and log-start-offset deltas with alerting before lag approaches retention; set retention.bytes as a disk safety valve but understand it can shrink the effective time window under spikes; for durable replay needs use tiered storage (KIP-405) to keep history cheaply far beyond local disk; and for state topics use compaction instead. Capacity-plan: retention.bytes is per partition, and segment rolling adds overshoot, so leave headroom.
go deeper
Know that a consumer lagging past retention can lose data and gets reset.
Explain log-start offset, OffsetOutOfRange, and auto.offset.reset behavior.
Size retention against worst-case lag, monitor the start-offset gap, and use size limits/tiered storage deliberately.
Treat retention as an SLA/capacity contract: cost vs replay guarantees, tiered storage strategy, policy selection per data semantics, and org-wide monitoring standards.
## The core tension `cleanup.policy=delete` removes data on a **time/size schedule that is blind to consumption**. Kafka deliberately decouples retention from consumer progress (unlike a classic queue that holds a message until acknowledged). This gives multiple independent consumer groups a shared replayable log — but it means **retention can delete data a consumer hasn't read yet**. ## The failure mode: OffsetOutOfRange - Each partition has a **log-start offset** (oldest still-retained record) and a **log-end offset** (newest). - A consumer commits the offset of the next record it wants. If retention deletes segments so the **log-start offset advances past the consumer's committed offset**, that offset no longer exists. - On the next fetch the broker returns an **OffsetOutOfRange** error. The consumer applies `auto.offset.reset`: - `earliest` → jump to the new log-start, **silently skipping** all the deleted records (data loss for that consumer). - `latest` → jump to the end, skipping even more. - `none` → throw, failing the consumer (at least it's loud). ## Operational and design risks - **Slow/stopped consumers**: an outage longer than retention loses data on resume. - **Backfill/replay limits**: you can only reprocess as far back as retention keeps. - **Disk pressure**: with `retention.bytes=-1` and time-only retention, a traffic spike can fill disks and take brokers down. - **Shrinking effective window**: with size retention, heavy load silently reduces the real time window below the nominal `retention.ms`. ## Mitigations 1. **Size retention above worst-case lag/downtime.** For a batch consumer that runs nightly, keep at least a couple of days. Formula intuition: `retention.ms ≥ max expected consumer downtime + safety margin`. 2. **Monitor lag and start-offset deltas.** Alert when `(committed offset − log-start offset)` shrinks toward zero, i.e. the consumer is getting close to being lapped, not just on raw lag. 3. **Use `retention.bytes` as a safety valve**, sized per partition with headroom, to bound disk — but treat the resulting time window as a floor, not the SLA. 4. **Tiered storage (KIP-405)** offloads old segments to object storage, letting you keep long history (days/weeks/months of replay) without local disk cost; local retention can stay small. 5. **Right policy for the data**: use **compaction** for current-state-per-key topics so the latest value is never aged out; use delete for true event streams. 6. **Plan for overshoot**: segment rolling means actual retention exceeds nominal by up to a segment, and per-partition `retention.bytes` multiplies by partition count. ## What good answers emphasize Retention is a capacity and SLA decision, not just a config — you trade disk cost against replay/recovery guarantees, and you must monitor the gap between consumer progress and the log-start offset, because retention will never wait for a lagging consumer.
- A consumer was down 3 days; topic retention is 1 day. What happens on restart with auto.offset.reset=earliest?Its committed offset is now below the log-start offset, so it gets OffsetOutOfRange and resets to earliest, silently skipping the ~2 days of records that retention already deleted — data loss for that consumer.
- Why is monitoring raw lag insufficient compared to watching the gap to log-start offset?Raw lag tells you how far behind the head you are; the danger is being lapped from behind by retention. Watching committed-offset minus log-start-offset shows how close you are to data being deleted out from under you.
saying these in an interview costs you the question
- Saying Kafka keeps a message until it's consumed/acknowledged (it doesn't under delete).
- Believing auto.offset.reset=earliest prevents data loss (it just hides the skip).
- Ignoring that size retention can shrink the real time window under load.
- Forgetting tiered storage / compaction as alternatives to giant local retention.