skip to content

Topics, Partitions and Log Storage

Kafka's data model: topics split into partitioned append-only logs, how keys route records, and how retention or compaction reclaims space. The most-asked Kafka area, because every design question lands on partitioning and ordering.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

How does the log cleaner decide which partitions to compact, and what role does min.cleanable.dirty.ratio play?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The log cleaner splits each partition into a 'clean' head (already compacted) and a 'dirty' tail (new records). It computes the dirty ratio = dirty bytes / total bytes and only cleans a partition when that ratio exceeds min.cleanable.dirty.ratio (default 0.5). The dirtiest eligible log is cleaned first.

open as a page

What problem does min.compaction.lag.ms solve, and how does it interact with consumers that need to see every update to a key?

level: seniorimportance: should knowfreq 35%

basics

~20 s

min.compaction.lag.ms sets a minimum age a record must reach before the cleaner may compact it away. It guarantees that every version of a key stays in the log for at least that long, so consumers reading within that window can observe intermediate updates, not just the latest value.

open as a page

How does Kafka resolve an offset from a timestamp (offsetsForTimes), and what are its limitations?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Kafka can look up, per partition, the offset of the earliest record whose timestamp is >= a given time, using the consumer's offsetsForTimes (ListOffsets API) and the per-segment .timeindex files. It returns null if no record has a timestamp at or after that time.

open as a page

You key correctly and Kafka delivers records in order, yet your consumer still processes events out of order. What consumer-side assumptions break ordering?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Kafka only delivers a partition in order; ordering breaks if the consumer hands records to a thread pool, processes multiple partitions concurrently per record-key, retries/dead-letters async, or commits offsets in a way that lets reprocessing reorder side effects.

open as a page

What are the per-partition costs that make over-partitioning harmful, and what cluster limits should you weigh when choosing a high partition count?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each partition costs open file handles (log segments + index files), broker memory, replication threads, and controller/metadata load. More partitions also slow leader-election failover and lengthen rebalances. Open file descriptor limits and end-to-end latency cap how many partitions a cluster can hold.

open as a page

How does compression.type work at the record-batch level, and what happens when producer, topic, and broker compression settings differ?

level: seniorimportance: should knowfreq 28%

basics

~20 s

compression.type sets the codec (none, gzip, snappy, lz4, zstd) applied to the whole batch's record section, recorded in the batch header. The producer normally compresses; the broker keeps it as-is unless the topic's compression.type forces a different codec, which makes the broker recompress.

open as a page

What does the batch-level CRC protect in RecordBatch v2, and how does that format design enable broker-side zero-copy (sendfile) of compressed data to consumers?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Each batch has one CRC-32C checksum covering the batch body (records + most of the header) to detect corruption. Because the whole compressed batch is a self-contained, opaque blob, the broker can stream it straight from the OS page cache to the network with sendfile — no decompression or JVM-heap copy.

open as a page

Why does retention act on closed segments only, and how do segment.bytes and segment.ms affect how quickly data is actually deleted?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Kafka deletes whole segment files, and only after a segment is closed (rolled) and fully past the limit. The active segment is never touched. segment.bytes and segment.ms control when segments roll, so they set how late data can actually be deleted versus its nominal retention.

open as a page

Design-wise, what risks does delete retention create for consumers and operations, and how do you mitigate retention deleting data a consumer still needs?

level: seniorimportance: should knowfreq 44%

basics

~20 s

If a consumer lags behind retention, Kafka can delete records before they're read, causing OffsetOutOfRange and data loss for that consumer. Mitigate by sizing retention above worst-case lag, monitoring lag, alerting, and using size limits as a disk safety valve.

open as a page

Walk through the describe/alter-config workflow with kafka-configs.sh and AdminClient, including incrementalAlterConfigs versus the legacy alterConfigs.

level: seniorimportance: should knowfreq 45%

basics

~10 s

Use kafka-configs.sh --describe to read a topic/broker's configs and their source, and --alter --add-config/--delete-config to change them. Programmatically, prefer AdminClient.incrementalAlterConfigs (SET/DELETE/APPEND ops) over the deprecated alterConfigs, which replaced the whole config set.

open as a page

Explain the append-only log structure of a partition and the role of offsets, log-start offset, and the high watermark.

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each partition is an append-only log: records are added at the tail, each getting a monotonically increasing offset. The log-start offset is the earliest still-retained offset; the high watermark is the highest offset replicated to all in-sync replicas — consumers can only read up to it.

open as a page

A team keys Kafka records by customer_id and a few large customers cause severe partition skew. Walk through the trade-offs and options for fixing it.

level: principalimportance: should knowfreq 34%

basics

~20 s

Keying by customer_id keeps each customer ordered but routes all of a hot customer's traffic to one partition, overloading it and its consumer. Options: composite keys to spread, a custom partitioner, more partitions, or accepting weaker per-key ordering — each trades ordering against balance.

open as a page

When and why would you set cleanup.policy=compact,delete, and what are the semantics of combining both policies?

level: principalimportance: should knowfreq 30%

basics

~20 s

compact,delete applies compaction (keep latest per key) AND time/size retention (drop old segments) on the same topic. Use it when you want current state per key but also need to bound the topic's age or size so unbounded keys or stale data eventually age out.

open as a page

A teammate wants a single monotonic sequence number across all partitions of a topic to order events globally. Why is this impossible with offsets, and what are the alternatives?

level: principalimportance: should knowfreq 35%

basics

~20 s

Offsets are independent per-partition counters, so they can't order events across partitions. Kafka gives no global ordering by design. Alternatives: use a single partition, key-partition only what must be ordered, or add an application-level timestamp/sequence and reorder downstream.

open as a page

When is a single-partition topic justified for ordering, and what alternatives give ordering without sacrificing throughput?

level: principalimportance: should knowfreq 40%

basics

~10 s

A single partition gives total topic order but caps throughput to one consumer and one broker. Prefer keyed multi-partition topics so you keep per-entity order while scaling, reserving single-partition for genuinely global-order, low-volume streams.

open as a page

Walk through how you would repartition a high-traffic keyed topic — for example to change the partition count or partitioning key — given Kafka cannot do it in place. What ordering and cutover concerns matter?

level: principalimportance: should knowfreq 45%

basics

~20 s

Create a new topic with the desired partition count/key scheme, then republish: a bridge consumer reads the old topic and a producer writes into the new one with the new partitioning. Migrate consumers, then producers, to the new topic, drain the old, and retire it.

open as a page

A topic's effective config value can come from several sources. What is the full precedence order Kafka uses to resolve a dynamic config, and how would you diagnose an unexpected effective value?

level: principalimportance: should knowfreq 30%

basics

~20 s

Order, highest to lowest: per-topic dynamic override, then per-broker dynamic config, then cluster-wide dynamic default broker config, then static broker config (server.properties), then the built-in default. To diagnose, describe the topic and read each entry's source tag.

open as a page

showing 31–47 of 47