When and why would you set cleanup.policy=compact,delete, and what are the semantics of combining both policies?
answer
- cleanup.policy is a comma list
- compact + retention.ms/bytes together
- retention can drop a key's LATEST value
- unbounded key space → need delete too
- Streams windowed changelog = compact,delete
basics
~20 scompact,delete applies compaction (keep latest per key) AND time/size retention (drop old segments) on the same topic. Use it when you want current state per key but also need to bound the topic's age or size so unbounded keys or stale data eventually age out.
solid answer
~50 scleanup.policy accepts a comma list. compact,delete runs both cleaners: the log cleaner compacts to the latest value per key, while the retention enforcer also deletes segments older than retention.ms or beyond retention.bytes. This is useful when a pure compacted topic would grow without bound because the key space is effectively unbounded (e.g. session IDs, request IDs) — you want fresh-per-key behavior but still need old data to expire. The key consequence: retention can delete a key's latest value even though compaction would have kept it, because retention operates on whole segments by age/size regardless of compaction status. So consumers cannot assume the latest value for every key is always present — it may have aged out. Kafka Streams windowed-store changelogs use compact,delete with retention.ms tied to the window+grace so old windows are reclaimed. Tune retention.ms/bytes alongside the compaction configs (min.cleanable.dirty.ratio, delete.retention.ms).
go deeper
Know that you can set both compact and delete on one topic.
Explain that it keeps latest-per-key while still aging out old segments and the unbounded-key motivation.
Articulate that retention can delete a key's latest value and tune retention vs compaction configs together.
Design windowed-changelog and bounded-state architectures, set bootstrap-safe retention, and reason about independent-cleaner failure modes.
## The two policies - **delete** (default): age/size retention — drops whole segments older than `retention.ms` or once the partition exceeds `retention.bytes`. - **compact**: keeps the latest value per key, removing superseded values via the log cleaner. `cleanup.policy` is a **list**, so `compact,delete` enables **both** simultaneously on one topic. ## Why combine them A pure compacted topic never expires data by time — it keeps the latest value of **every key forever**. If your key space is effectively unbounded (per-session keys, per-request keys, time-bucketed keys), that latest-per-key set itself grows without limit and the topic balloons. Combining with delete lets you say: keep the latest value per key, **but** also let segments older than retention.ms (or beyond retention.bytes) age out entirely. You get current-state semantics *and* a hard bound on storage/age. ## Critical semantic: retention can remove a 'latest' value The two cleaners are independent. Retention deletes **whole segments by age/size**, with no awareness of whether a record is the latest for its key. So under compact,delete, **a key's most recent value can be deleted** simply because its segment aged out. This breaks the naive compaction assumption that 'the latest value per key is always readable.' Consumers and materialized stores must tolerate keys disappearing due to retention, not just due to tombstones. ## Canonical use case: Kafka Streams windowed changelogs Kafka Streams configures **windowed** state-store changelog topics as `compact,delete` with `retention.ms` set to the window size plus grace period (and a retention slack). Compaction keeps the latest per windowed key; retention reclaims segments for windows that have fully expired, so the changelog doesn't accumulate dead windows forever. (Non-windowed KTable changelogs use plain `compact`.) ## Tuning interplay - `retention.ms` / `retention.bytes` — the delete side. - `min.cleanable.dirty.ratio`, `min.compaction.lag.ms`, `max.compaction.lag.ms` — the compact side. - `delete.retention.ms` — tombstone visibility; still applies since compact is active. - `segment.ms` / `segment.bytes` — smaller segments let retention reclaim space at finer granularity (retention only drops *closed* segments). ## Operational cautions - Don't set retention.ms shorter than the time consumers need to bootstrap full state, or new consumers will start from an incomplete snapshot. - Monitor that both cleaners keep up; a stuck log cleaner plus active retention can yield a topic that loses old keys but never compacts dirty ones. - Switching an existing topic from compact to compact,delete will start aging out historical data — irreversible for already-deleted segments.
- Under compact,delete, can a consumer assume the latest value for every key is always present?No. Retention deletes whole segments by age/size irrespective of compaction, so a key's most recent value can age out. Consumers must tolerate keys vanishing, not just tombstone deletes.
- Which Kafka Streams store type uses compact,delete and why?Windowed state-store changelogs. retention.ms is set to window size + grace so expired windows' segments are reclaimed, while compaction keeps the latest value per windowed key. Non-windowed KTable changelogs use plain compact.
- Why must retention.ms be large enough for new-consumer bootstrap?A new consumer reconstructs full state by reading from the log start. If retention already deleted segments holding some keys' latest values, the bootstrapped snapshot is incomplete.
saying these in an interview costs you the question
- Saying compact,delete still guarantees the latest value per key is always present — retention can delete it.
- Treating cleanup.policy as a single value rather than a comma-separated list.
- Confusing the use case with plain compact (KTable) vs windowed (compact,delete).
- Assuming the two cleaners coordinate to protect latest values — they operate independently.