Topics, Partitions and Log Storage
Kafka's data model: topics split into partitioned append-only logs, how keys route records, and how retention or compaction reclaims space. The most-asked Kafka area, because every design question lands on partitioning and ordering.
part ofApache Kafkaoverview, primer and where to startread it →on this pageshowhide
explore
- Topics and Partition Fundamentals5 questions
- Partition Keys and Hash Routing5 questions
- Ordering Guarantees5 questions
- Choosing and Changing Partition Count5 questions
- Retention Policies and Cleanup (Delete)5 questions
- Log Compaction and Tombstones5 questions
- Offset Addressing Model6 questions
- Record Batches and Message Format6 questions
- Topic Configuration and Administration5 questions
questions
page 2 of 2How does the log cleaner decide which partitions to compact, and what role does min.cleanable.dirty.ratio play?
basics
~20 sThe log cleaner splits each partition into a 'clean' head (already compacted) and a 'dirty' tail (new records). It computes the dirty ratio = dirty bytes / total bytes and only cleans a partition when that ratio exceeds min.cleanable.dirty.ratio (default 0.5). The dirtiest eligible log is cleaned first.
What problem does min.compaction.lag.ms solve, and how does it interact with consumers that need to see every update to a key?
basics
~20 smin.compaction.lag.ms sets a minimum age a record must reach before the cleaner may compact it away. It guarantees that every version of a key stays in the log for at least that long, so consumers reading within that window can observe intermediate updates, not just the latest value.
How does Kafka resolve an offset from a timestamp (offsetsForTimes), and what are its limitations?
basics
~20 sKafka can look up, per partition, the offset of the earliest record whose timestamp is >= a given time, using the consumer's offsetsForTimes (ListOffsets API) and the per-segment .timeindex files. It returns null if no record has a timestamp at or after that time.
You key correctly and Kafka delivers records in order, yet your consumer still processes events out of order. What consumer-side assumptions break ordering?
basics
~20 sKafka only delivers a partition in order; ordering breaks if the consumer hands records to a thread pool, processes multiple partitions concurrently per record-key, retries/dead-letters async, or commits offsets in a way that lets reprocessing reorder side effects.
What are the per-partition costs that make over-partitioning harmful, and what cluster limits should you weigh when choosing a high partition count?
basics
~20 sEach partition costs open file handles (log segments + index files), broker memory, replication threads, and controller/metadata load. More partitions also slow leader-election failover and lengthen rebalances. Open file descriptor limits and end-to-end latency cap how many partitions a cluster can hold.
How does compression.type work at the record-batch level, and what happens when producer, topic, and broker compression settings differ?
basics
~20 scompression.type sets the codec (none, gzip, snappy, lz4, zstd) applied to the whole batch's record section, recorded in the batch header. The producer normally compresses; the broker keeps it as-is unless the topic's compression.type forces a different codec, which makes the broker recompress.
What does the batch-level CRC protect in RecordBatch v2, and how does that format design enable broker-side zero-copy (sendfile) of compressed data to consumers?
basics
~20 sEach batch has one CRC-32C checksum covering the batch body (records + most of the header) to detect corruption. Because the whole compressed batch is a self-contained, opaque blob, the broker can stream it straight from the OS page cache to the network with sendfile — no decompression or JVM-heap copy.
Why does retention act on closed segments only, and how do segment.bytes and segment.ms affect how quickly data is actually deleted?
basics
~20 sKafka deletes whole segment files, and only after a segment is closed (rolled) and fully past the limit. The active segment is never touched. segment.bytes and segment.ms control when segments roll, so they set how late data can actually be deleted versus its nominal retention.
Design-wise, what risks does delete retention create for consumers and operations, and how do you mitigate retention deleting data a consumer still needs?
basics
~20 sIf a consumer lags behind retention, Kafka can delete records before they're read, causing OffsetOutOfRange and data loss for that consumer. Mitigate by sizing retention above worst-case lag, monitoring lag, alerting, and using size limits as a disk safety valve.
Walk through the describe/alter-config workflow with kafka-configs.sh and AdminClient, including incrementalAlterConfigs versus the legacy alterConfigs.
basics
~10 sUse kafka-configs.sh --describe to read a topic/broker's configs and their source, and --alter --add-config/--delete-config to change them. Programmatically, prefer AdminClient.incrementalAlterConfigs (SET/DELETE/APPEND ops) over the deprecated alterConfigs, which replaced the whole config set.
Explain the append-only log structure of a partition and the role of offsets, log-start offset, and the high watermark.
basics
~20 sEach partition is an append-only log: records are added at the tail, each getting a monotonically increasing offset. The log-start offset is the earliest still-retained offset; the high watermark is the highest offset replicated to all in-sync replicas — consumers can only read up to it.
A team keys Kafka records by customer_id and a few large customers cause severe partition skew. Walk through the trade-offs and options for fixing it.
basics
~20 sKeying by customer_id keeps each customer ordered but routes all of a hot customer's traffic to one partition, overloading it and its consumer. Options: composite keys to spread, a custom partitioner, more partitions, or accepting weaker per-key ordering — each trades ordering against balance.
When and why would you set cleanup.policy=compact,delete, and what are the semantics of combining both policies?
basics
~20 scompact,delete applies compaction (keep latest per key) AND time/size retention (drop old segments) on the same topic. Use it when you want current state per key but also need to bound the topic's age or size so unbounded keys or stale data eventually age out.
A teammate wants a single monotonic sequence number across all partitions of a topic to order events globally. Why is this impossible with offsets, and what are the alternatives?
basics
~20 sOffsets are independent per-partition counters, so they can't order events across partitions. Kafka gives no global ordering by design. Alternatives: use a single partition, key-partition only what must be ordered, or add an application-level timestamp/sequence and reorder downstream.
When is a single-partition topic justified for ordering, and what alternatives give ordering without sacrificing throughput?
basics
~10 sA single partition gives total topic order but caps throughput to one consumer and one broker. Prefer keyed multi-partition topics so you keep per-entity order while scaling, reserving single-partition for genuinely global-order, low-volume streams.
Walk through how you would repartition a high-traffic keyed topic — for example to change the partition count or partitioning key — given Kafka cannot do it in place. What ordering and cutover concerns matter?
basics
~20 sCreate a new topic with the desired partition count/key scheme, then republish: a bridge consumer reads the old topic and a producer writes into the new one with the new partitioning. Migrate consumers, then producers, to the new topic, drain the old, and retire it.
A topic's effective config value can come from several sources. What is the full precedence order Kafka uses to resolve a dynamic config, and how would you diagnose an unexpected effective value?
basics
~20 sOrder, highest to lowest: per-topic dynamic override, then per-broker dynamic config, then cluster-wide dynamic default broker config, then static broker config (server.properties), then the built-in default. To diagnose, describe the topic and read each entry's source tag.
showing 31–47 of 47