What is a Kafka topic, and what is a partition?
answer
- Topic = logical name; partition = physical log
- Partition = append-only, ordered
- Offset is per-partition
- Partition = unit of parallelism/storage/replication
- Order guaranteed within a partition only
basics
~20 sA topic is a named stream of records (like a category or feed name). Each topic is split into one or more partitions. A partition is an ordered, append-only log of records that Kafka stores and serves.
solid answer
~40 sA topic is the logical, named channel producers write to and consumers read from — e.g. `payments` or `clicks`. Kafka does not store a topic as one big sequence; it splits the topic into N partitions. A partition is a physical, append-only, ordered log: records are appended at the end and never modified in place. Each record in a partition gets a monotonically increasing offset (its position in that partition). Partitions are the real unit of storage, parallelism, and replication — the topic is just the umbrella name over them. Producers append to partitions; consumers read partitions sequentially by advancing their offset. So 'topic' is the logical abstraction your application talks about, while 'partition' is the concrete log Kafka actually writes, replicates, and serves.
go deeper
Define topic (named stream) and partition (ordered append-only log); know offsets exist.
Explain why topics are split: parallelism, storage, replication; ordering only within a partition.
Articulate offset semantics (per-partition, monotonic), partition as unit of replication, and the ordering trade-off it forces on system design.
Frame topic/partition split as the core scalability primitive and teach the downstream consequences (ordering, keying, consumer parallelism) to a team.
## First principles Kafka is a distributed **append-only log** system. To understand topics and partitions you only need two ideas: a *log* (a sequence you can only append to and read in order) and *naming/splitting* (how Kafka labels and scales those logs). ### Topic A **topic** is a human-chosen name that groups related records — for example `orders`, `user-signups`, or `sensor-readings`. Applications publish (produce) to a topic and subscribe (consume) from a topic. The topic itself is a **logical** concept: it has no single physical file. Think of it as a category label. ### Partition Under the hood, every topic is divided into a fixed number **N** of **partitions**, numbered `0` to `N-1`. A partition is the *physical* thing Kafka actually stores: an **append-only, ordered log**. - **Append-only:** new records are only added at the tail. Existing records are never edited or inserted in the middle. (Old records eventually disappear via retention or compaction, but they are not mutated.) - **Ordered:** within a single partition, records are kept in the exact order they were appended. - **Offset:** each record is assigned a monotonically increasing 64-bit integer called the **offset**, which is its position in *that* partition (0, 1, 2, …). Offsets are per-partition, not per-topic — partition 0 and partition 1 both have an offset 5, referring to different records. ### Why split a topic into partitions? A partition is the **unit of parallelism, storage, and replication**: - **Parallelism:** different partitions can be produced to and consumed from independently and in parallel, on different brokers. More partitions → more throughput and more consumers can work at once. - **Storage:** each partition's log is stored as a set of files (segments) on a broker's disk. Splitting a topic spreads its data across machines. - **Replication:** Kafka replicates each partition (not the whole topic) to multiple brokers for fault tolerance. ### Ordering caveat Kafka guarantees ordering **within a partition**, not across the whole topic. If record A and record B land in different partitions, Kafka makes no promise about which a consumer sees first. This is the single most important consequence of the topic→partition split. ### Mental model ``` Topic "orders" (logical name) ├─ Partition 0: [r0][r1][r2][r3] ... (append at the end →) ├─ Partition 1: [r0][r1][r2] ... └─ Partition 2: [r0][r1][r2][r3][r4] ... ``` The topic is the box drawn around the three logs; each log is a partition with its own independent offset sequence.
- Does Kafka guarantee ordering across a whole topic?No. Ordering is guaranteed only within a single partition. Records in different partitions have no defined relative order, since each partition is an independent log with its own offsets.
- Are offsets unique across a topic?No, offsets are per-partition. Each partition starts at offset 0 and increases independently, so the same offset value exists in every partition pointing to different records. A record is globally identified by (topic, partition, offset).
saying these in an interview costs you the question
- Saying Kafka guarantees ordering across an entire topic
- Treating offset as a topic-wide (global) sequence number
- Claiming a topic is stored as a single file/log
- Confusing partitions with replicas (a partition is the log; replicas are copies of it)