What does log compaction do to a Kafka-style topic, and how does it differ from simple time/size-based retention?
answer
- keeps latest value per key
- tombstone = null value deletes key
- backs KTable/state-store changelogs
- delete policy = age/size; compact policy = per-key latest
- background cleaner thread rewrites segments
basics
~20 sCompaction keeps only the newest record for each key and throws away older ones, instead of deleting everything older than some age. It's like keeping only the latest version of each row in a table, not a full history.
solid answer
~40 sTime/size retention deletes whole segments once they age out or the topic exceeds a size cap, regardless of key - eventually all data disappears. Compaction instead runs a background process per partition that, for every key, keeps only the most recent record and removes earlier ones with the same key (a null-valued record acts as a 'tombstone' that deletes the key entirely after a grace period). The result is a topic that always holds at least the latest state for every key seen, functioning like a durable, replayable key-value snapshot rather than a bounded event history. This is exactly what backs changelog topics for stream-processing state stores and KTables: on recovery, replaying the compacted topic rebuilds the current state without replaying every historical update.
go deeper
Should know compaction keeps the latest value per key rather than deleting by age, at a conceptual level.
Should be able to explain tombstones, when to choose compact vs delete policy, and why compaction suits state-store changelogs.
Should reason about the async nature of compaction, key-design implications, and the operational cost (disk/CPU) of running the cleaner.
Should design system-wide data-retention strategy choosing compact vs delete vs combined policies per topic based on downstream consumption patterns and compliance/storage constraints.
## Why a partition needs a cleanup policy Every topic partition needs a policy for what happens to old data, because disk isn't infinite. Kafka exposes **two cleanup policies**, and they solve different problems. | The 'delete' policy | The 'compact' policy | |---|---| | Segments (the physical log files a partition is split into) are removed once they exceed a configured age (e.g., 7 days) or the partition exceeds a configured size. | A background thread called the **log cleaner** periodically scans each partition and, for every distinct key, keeps only the record with the highest offset (the most recent write) and removes every earlier record sharing that key. | | It doesn't look at the content of the records at all - once a segment ages out, everything in it is gone, key or no key. | The mechanical result is that a compacted topic converges toward holding exactly one record per key: the latest known value. | ## The 'delete' policy - a history of what happened The 'delete' policy is the default. This is the right policy for genuine event streams: a stream of 'OrderPlaced' events where you care about the full sequence of what happened, and you're fine losing events older than your retention window because downstream consumers have already processed them and the business doesn't need year-old raw events sitting in the log. ## The 'compact' policy - current state The 'compact' policy works differently and solves a different problem: representing current state rather than a history of events. Every record in a topic is a **key-value pair**. - If a producer writes a record with a **null value** for a given key, that acts as a **'tombstone'** - after a configurable delay (`delete.retention.ms`), the compactor removes the key entirely, including the tombstone itself, so deletions eventually get cleaned up too. - It behaves less like an event log and more like a durable, append-only-on-the-wire representation of a key-value table, one you can always fully rebuild by reading the topic from the start. ## Why it exists: backing a state store The reason this exists is that stream-processing frameworks (Kafka Streams' **KTable**, **ksqlDB** tables, **Flink's** changelog streams) need a way to durably back an in-memory state store - things like 'current running total per customer' or 'latest known inventory count per SKU' - so that state survives a process crash or gets rebuilt when a task moves to a different machine. - A raw event-retention topic is the **wrong backing store** for this: if you tried to rebuild state by replaying seven days of raw update events, you'd replay stale intermediate values you no longer care about, and eventually old updates would age out and be lost, silently corrupting the rebuilt state. - A compacted **'changelog' topic** solves both problems: replaying it always yields the current state, and it never loses a key's latest value to a time-based retention cutoff, because compaction is keyed, not age-based. ## The trade-off The trade-off is that compaction costs continuous background I/O and disk churn: - The cleaner thread has to periodically rewrite segments to physically remove superseded records, which is CPU and disk work you don't pay for with simple delete-based retention. - It also changes your **data-modeling contract**: every record must carry a meaningful key, because compaction is entirely key-driven; a topic full of null or duplicate-nonsense keys either compacts uselessly or, worse, silently collapses records you actually wanted to keep distinct. - There's also a subtlety around **timing** - compaction runs asynchronously and isn't instantaneous, so a consumer reading a 'compacted' topic can still transiently see multiple records for the same key if it reads between compaction passes; consumers of compacted topics need to be written expecting 'last write for this key wins' semantics, not 'there is exactly one record per key at all times.' ## Failure modes The most common failure mode in production is choosing the wrong policy for the job: 1. Using plain delete-based retention for what is actually a state changelog, so historical intermediate updates silently age out and a task recovering after a crash rebuilds an incomplete or wrong state store. 2. Or, the reverse, applying compaction to what should be a genuine event history, silently collapsing away events a downstream analytics consumer actually needed (e.g., every price change for an item, not just the latest price). 3. Another common issue is forgetting to send tombstones for deleted entities, so compacted topics accumulate keys forever with no way to signal 'this entity no longer exists.' ## Where it shows up A concrete, widely used example is Kafka Streams itself: every KTable and every stateful operator's local **RocksDB** store is backed by a compacted changelog topic under the hood. When a stream-processing task is reassigned to a new machine (say, after a rebalance or a crash), it rebuilds its local state store by replaying that compacted changelog from the beginning - which is fast and correct precisely because compaction guarantees the replay yields only current values, not the full history of every update ever made.
- Why would using delete-based retention instead of compaction break a Kafka Streams state store's recovery?Delete-based retention removes data purely by age, so once the retention window passes, early updates to a key are gone even if that key still needs to be reflected in the rebuilt state. A task recovering by replaying such a topic would miss updates that aged out, producing an incomplete or wrong state store instead of the correct current value.
- What's the purpose of a tombstone record in a compacted topic, and what could go wrong if producers never send them?A tombstone (a record with a null value) tells the compactor 'this key is deleted, remove it entirely after the grace period.' Without tombstones, deleted entities' keys and their last known value linger in the compacted topic forever, so anything rebuilding state from it will incorrectly resurrect entities that should no longer exist.
- Can a consumer reading a compacted topic ever see more than one record for the same key?Yes - compaction runs asynchronously in the background, so between compaction passes a partition can still contain several superseded records for a key. Consumers of compacted topics should be written to apply 'last write wins per key' logic rather than assume strict one-record-per-key at read time.
It's like a whiteboard where instead of erasing the whole thing every week, someone walks around and erases only outdated sticky notes, replacing them with nothing once a newer note with the same label appears - the board always shows the latest note per label, never a full history of every note ever posted.
saying these in an interview costs you the question
- Thinks compaction deletes data based on age
- Doesn't know what a tombstone is
- Applies compaction to a topic where full event history is required downstream
- Assumes a compacted topic always has exactly one record per key at any instant
- Confuses compaction with simple deduplication of identical records