skip to content

Topics and Partition Fundamentals

What a topic and a partition really are, and why the partition is the unit of parallelism, ordering, and replication. The baseline question every Kafka interview starts from.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a Kafka topic, and what is a partition?

level: juniorimportance: must knowfreq 90%

answer

  1. Topic = logical name; partition = physical log
  2. Partition = append-only, ordered
  3. Offset is per-partition
  4. Partition = unit of parallelism/storage/replication
  5. Order guaranteed within a partition only

basics

~20 s

A topic is a named stream of records (like a category or feed name). Each topic is split into one or more partitions. A partition is an ordered, append-only log of records that Kafka stores and serves.

solid answer

~40 s

A topic is the logical, named channel producers write to and consumers read from — e.g. `payments` or `clicks`. Kafka does not store a topic as one big sequence; it splits the topic into N partitions. A partition is a physical, append-only, ordered log: records are appended at the end and never modified in place. Each record in a partition gets a monotonically increasing offset (its position in that partition). Partitions are the real unit of storage, parallelism, and replication — the topic is just the umbrella name over them. Producers append to partitions; consumers read partitions sequentially by advancing their offset. So 'topic' is the logical abstraction your application talks about, while 'partition' is the concrete log Kafka actually writes, replicates, and serves.

go deeper

for a junior

Define topic (named stream) and partition (ordered append-only log); know offsets exist.

for a middle

Explain why topics are split: parallelism, storage, replication; ordering only within a partition.

for a senior

Articulate offset semantics (per-partition, monotonic), partition as unit of replication, and the ordering trade-off it forces on system design.

for a principal

Frame topic/partition split as the core scalability primitive and teach the downstream consequences (ordering, keying, consumer parallelism) to a team.

## First principles Kafka is a distributed **append-only log** system. To understand topics and partitions you only need two ideas: a *log* (a sequence you can only append to and read in order) and *naming/splitting* (how Kafka labels and scales those logs). ### Topic A **topic** is a human-chosen name that groups related records — for example `orders`, `user-signups`, or `sensor-readings`. Applications publish (produce) to a topic and subscribe (consume) from a topic. The topic itself is a **logical** concept: it has no single physical file. Think of it as a category label. ### Partition Under the hood, every topic is divided into a fixed number **N** of **partitions**, numbered `0` to `N-1`. A partition is the *physical* thing Kafka actually stores: an **append-only, ordered log**. - **Append-only:** new records are only added at the tail. Existing records are never edited or inserted in the middle. (Old records eventually disappear via retention or compaction, but they are not mutated.) - **Ordered:** within a single partition, records are kept in the exact order they were appended. - **Offset:** each record is assigned a monotonically increasing 64-bit integer called the **offset**, which is its position in *that* partition (0, 1, 2, …). Offsets are per-partition, not per-topic — partition 0 and partition 1 both have an offset 5, referring to different records. ### Why split a topic into partitions? A partition is the **unit of parallelism, storage, and replication**: - **Parallelism:** different partitions can be produced to and consumed from independently and in parallel, on different brokers. More partitions → more throughput and more consumers can work at once. - **Storage:** each partition's log is stored as a set of files (segments) on a broker's disk. Splitting a topic spreads its data across machines. - **Replication:** Kafka replicates each partition (not the whole topic) to multiple brokers for fault tolerance. ### Ordering caveat Kafka guarantees ordering **within a partition**, not across the whole topic. If record A and record B land in different partitions, Kafka makes no promise about which a consumer sees first. This is the single most important consequence of the topic→partition split. ### Mental model ``` Topic "orders" (logical name) ├─ Partition 0: [r0][r1][r2][r3] ... (append at the end →) ├─ Partition 1: [r0][r1][r2] ... └─ Partition 2: [r0][r1][r2][r3][r4] ... ``` The topic is the box drawn around the three logs; each log is a partition with its own independent offset sequence.

  • Does Kafka guarantee ordering across a whole topic?
    No. Ordering is guaranteed only within a single partition. Records in different partitions have no defined relative order, since each partition is an independent log with its own offsets.
  • Are offsets unique across a topic?
    No, offsets are per-partition. Each partition starts at offset 0 and increases independently, so the same offset value exists in every partition pointing to different records. A record is globally identified by (topic, partition, offset).

saying these in an interview costs you the question

  • Saying Kafka guarantees ordering across an entire topic
  • Treating offset as a topic-wide (global) sequence number
  • Claiming a topic is stored as a single file/log
  • Confusing partitions with replicas (a partition is the log; replicas are copies of it)

context

open as a page

Why is the partition — not the topic — the fundamental unit of parallelism and replication in Kafka?

level: middleimportance: must knowfreq 78%

basics

~20 s

Because Kafka stores, replicates, and serves data per partition. Each partition lives on a broker, has its own replicas, and can be consumed by exactly one consumer in a group at a time — so partitions, not topics, drive parallelism.

open as a page

How does a producer decide which partition a record goes to, and how does that interact with ordering guarantees?

level: seniorimportance: must knowfreq 72%

basics

~20 s

If a record has a key, the default partitioner hashes the key (murmur2 mod partition count) so the same key always maps to the same partition, preserving per-key order. Keyless records are spread across partitions (sticky/round-robin), so their relative order across partitions isn't guaranteed.

open as a page

What is a TopicPartition, and how is a single record uniquely addressed in Kafka?

level: middleimportance: should knowfreq 60%

basics

~10 s

A TopicPartition is the identity of one specific partition: the pair (topic name, partition number). A single record is uniquely addressed by adding its offset, giving the triple (topic, partition, offset).

open as a page

Explain the append-only log structure of a partition and the role of offsets, log-start offset, and the high watermark.

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each partition is an append-only log: records are added at the tail, each getting a monotonically increasing offset. The log-start offset is the earliest still-retained offset; the high watermark is the highest offset replicated to all in-sync replicas — consumers can only read up to it.

open as a page