What is a TopicPartition, and how is a single record uniquely addressed in Kafka?
answer
- TopicPartition = (topic, partition)
- Record address = (topic, partition, offset)
- Key type for seek/assign/commit
- Offset meaningless without partition
- One leader per TopicPartition
basics
~10 sA TopicPartition is the identity of one specific partition: the pair (topic name, partition number). A single record is uniquely addressed by adding its offset, giving the triple (topic, partition, offset).
solid answer
~40 sIn the Kafka client, `org.apache.kafka.common.TopicPartition` is a small value object holding a topic name and a partition index — it names exactly one partition's log. It is the key type used throughout the consumer API: `assign(Collection<TopicPartition>)`, `seek(TopicPartition, offset)`, `position(TopicPartition)`, and offset-commit maps (`Map<TopicPartition, OffsetAndMetadata>`) are all keyed by it. A single record is uniquely addressed by the triple **(topic, partition, offset)**: the TopicPartition pins down which log, and the offset pins down the position within that log. Offsets are only meaningful relative to a TopicPartition because each partition has its own independent offset sequence. This identity is also how Kafka tracks consumer progress — committed offsets are stored per (group, topic, partition) — and how the producer reports `RecordMetadata` with the topic, partition, and assigned offset after a successful write.
go deeper
Know that a partition is named by (topic, partition) and a record by (topic, partition, offset).
Recognize TopicPartition as the API key for assign/seek/commit and explain why the partition is needed to interpret an offset.
Connect TopicPartition to leader routing, committed-offset storage, and replay/seek mechanics.
Use the (TopicPartition, offset) addressing model to design replay, lag monitoring, and exactly-once positioning across a platform.
## What identifies a partition A topic is just a name shared by many partitions, so a topic name alone is ambiguous. To point at one specific log you need **two** things: 1. the **topic name** (e.g. `orders`), and 2. the **partition index** (e.g. `2`). The pair `(orders, 2)` is a **TopicPartition** — in the Java/Kotlin client this is the class `org.apache.kafka.common.TopicPartition`, an immutable value object with `topic()` and `partition()` accessors and value-based `equals`/`hashCode` so it works as a map key. ## What identifies a single record Within a partition, every record has an **offset** — its monotonically increasing position (0, 1, 2, …). Offsets reset per partition, so an offset is meaningless without knowing which partition it belongs to. Therefore the globally unique address of any record is the triple: ``` (topic, partition, offset) ``` or equivalently `(TopicPartition, offset)`. ## Where TopicPartition shows up in the API It is the universal key for partition-scoped operations: - **Manual assignment:** `consumer.assign(List.of(new TopicPartition("orders", 0)))`. - **Seeking:** `consumer.seek(tp, 1234L)` to reposition; `seekToBeginning(tps)` / `seekToEnd(tps)`. - **Position & committed:** `consumer.position(tp)` (next offset to read), `consumer.committed(Set<TopicPartition>)`. - **Commit maps:** `consumer.commitSync(Map<TopicPartition, OffsetAndMetadata>)`. - **Rebalance callbacks:** `ConsumerRebalanceListener.onPartitionsAssigned/Revoked(Collection<TopicPartition>)`. - **End offsets / lag:** `consumer.endOffsets(Collection<TopicPartition>)`. ## Leader assignment Each TopicPartition has, at any moment, exactly one **leader** broker that serves its reads and writes, plus follower replicas. Clients use metadata to map a TopicPartition to its current leader broker; when leadership moves (failover), the client refreshes metadata and re-routes. So a TopicPartition is also the granularity at which leadership and routing are tracked. ## Producer side After a successful send, the broker returns `RecordMetadata` exposing `topic()`, `partition()`, and `offset()` — i.e. it tells you the exact `(topic, partition, offset)` where your record landed. ## Why this matters Consumer progress (committed offsets), lag monitoring, exactly-once positioning, and manual replay all operate on `(TopicPartition, offset)`. Internally, committed offsets are stored per **(consumer group, topic, partition)** in the `__consumer_offsets` topic — again partition-scoped, never topic-scoped.
- Why can't a record be identified by (topic, offset) alone?Because offsets are per-partition: every partition independently has an offset 0, 1, 2, …. The same offset value exists in every partition. You must include the partition to disambiguate which log the offset refers to.
- Where does Kafka store the per-partition consumer progress?In the internal compacted topic `__consumer_offsets`, keyed by (group id, topic, partition) → committed offset. Progress is tracked per TopicPartition per group, never per whole topic.
saying these in an interview costs you the question
- Identifying a record by (topic, offset) without the partition
- Treating offsets as comparable across partitions
- Thinking TopicPartition includes the offset (it does not — that's the separate position)
- Assuming a partition's leader never changes