skip to content

Architecture and Internals

How a Kafka broker actually stores and serves data: append-only segment files, indexes, page cache, and the request-handling threads. Interviewers dig here to separate people who have run Kafka from people who have only used the client API.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What is the purpose of bootstrap.servers, and why don't you need to list every broker?

level: juniorimportance: must knowfreq 80%

answer

  1. seed, not full list
  2. first Metadata request
  3. client learns all brokers + leaders
  4. 2-3 entries for redundancy
  5. real addresses come from advertised.listeners

basics

~20 s

bootstrap.servers is the initial list of broker host:port pairs a client contacts to discover the full cluster. After the first metadata fetch the client learns all brokers, so the list only needs a few entries for redundancy.

solid answer

~40 s

bootstrap.servers tells a Kafka client where to make its first connection. The client connects to any one of the listed brokers and issues a Metadata request; the broker responds with the full cluster topology: every broker's id and advertised address, plus which broker leads each partition. From then on the client connects directly to the relevant leaders and ignores bootstrap.servers except to re-bootstrap if all connections drop. That's why you don't enumerate the whole cluster: the list is only a discovery seed. You still provide 2-3 entries (often a load balancer or several brokers) so the client can bootstrap even if one seed broker is down. Crucially, the addresses brokers return come from advertised.listeners, not from bootstrap.servers, so the bootstrap host being reachable doesn't guarantee the advertised addresses are.

go deeper

for a junior

Know it's the initial connection seed and a partial list suffices.

for a middle

Explain the Metadata request and that real addresses come from advertised.listeners.

for a senior

Diagnose the advertised-address mismatch failure and bootstrap-vs-data-path distinction.

for a principal

Design client connectivity across NAT/LB/multi-network topologies and reason about SPOFs.

## The discovery problem A Kafka cluster has many brokers, and partition leadership moves around (failover, rebalancing). A client must always send each produce/fetch to the *current leader* of the target partition. Hardcoding all brokers and leaders would be brittle. Kafka solves this with **bootstrap + metadata discovery**. ## bootstrap.servers `bootstrap.servers` is a client config: a comma-separated list of `host:port` entries, e.g. `b1:9092,b2:9092`. On startup the client: 1. Connects to one reachable entry. 2. Sends a **Metadata** request. 3. Receives the full picture: all broker ids and their **advertised** addresses, and the leader broker for every partition of the topics it cares about. 4. Opens direct connections to the leaders it needs. After this, bootstrap.servers is essentially unused unless the client loses all connections and must re-bootstrap. ## Why a partial list is enough Because any broker can answer the Metadata request with the *entire* cluster view, you only need one reachable seed. You list 2-3 for **availability** (so bootstrapping survives a single broker outage), not for completeness. Listing all brokers is unnecessary and adds maintenance burden as the cluster grows. ## The advertised-address gotcha The addresses the client uses *after* bootstrap come from each broker's **advertised.listeners**, not from bootstrap.servers. A classic failure: bootstrap.servers points at a reachable address (e.g. via a NAT or load balancer), the Metadata request succeeds, but the advertised addresses returned are internal hostnames the client cannot resolve/reach. The client then fails on the first produce/fetch despite a "successful" bootstrap. Always ensure advertised.listeners are reachable by clients. ## Edge cases - A single load-balancer DNS name as the only bootstrap entry works but is a SPOF for *bootstrapping* if the LB fails; the data path still uses advertised addresses directly. - Wrong port or listener (e.g. pointing at an inter-broker listener) bootstraps to the wrong security protocol and fails.

  • After bootstrapping, where does the client get the addresses it actually connects to?
    From each broker's advertised.listeners, returned in the Metadata response; bootstrap.servers is only the initial seed.
  • Why list more than one bootstrap server?
    Redundancy: if one seed broker is down, the client can still bootstrap from another. It is not for completeness.

saying these in an interview costs you the question

  • Saying you must list every broker
  • Claiming the client keeps using bootstrap.servers for all requests
  • Confusing bootstrap.servers with advertised.listeners
  • Thinking a reachable bootstrap address guarantees the data path works

context

open as a page

What is a Kafka broker, and what role does broker.id play in a cluster?

level: juniorimportance: must knowfreq 75%

basics

~10 s

A broker is a single Kafka server that stores topic partition data and serves produce/fetch requests. broker.id is its unique numeric identifier within the cluster; no two brokers may share the same id.

open as a page

What is a Kafka partition's commit log, and why is it split into segments on disk?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Each partition is an append-only log: records are only added to the end, never changed in place. Kafka splits that log into fixed-size files called segments so old data can be deleted or compacted one whole file at a time instead of editing one giant file.

open as a page

What is the Kafka log cleaner, and how do you enable it for a topic?

level: juniorimportance: must knowfreq 60%

basics

~20 s

The log cleaner is a background process that runs compaction: it scans a topic's log and keeps only the latest record per key, deleting older duplicates. You enable it by setting the topic config cleanup.policy=compact.

open as a page

Why does Kafka rely on the operating system page cache instead of maintaining its own in-process (JVM heap) record cache?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Kafka writes data to files and lets the OS keep recently used file pages in RAM (the page cache). It avoids a JVM heap cache to dodge GC pressure, double-buffering, and to reuse the OS cache that survives broker restarts.

open as a page

What are the .index and .timeindex files that Kafka maintains alongside each log segment, and what does each one map?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Each Kafka log segment has a .index file mapping a message offset to its physical byte position in the .log file, and a .timeindex file mapping a timestamp to an offset. They let Kafka jump near a record without scanning the whole segment.

open as a page

What is Kafka Tiered Storage (KIP-405), and what problem does it solve at a high level?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Tiered Storage (KIP-405) moves old, closed log segments from local broker disks to cheaper remote object storage (like S3). Recent data stays local for fast reads; older data is fetched from remote when needed.

open as a page

Explain the difference between listeners and advertised.listeners, and the role of listener.security.protocol.map.

level: middleimportance: must knowfreq 65%

basics

~20 s

listeners is where the broker binds and accepts connections. advertised.listeners is the address it publishes in metadata for clients to connect back. listener.security.protocol.map maps each named listener to a security protocol (PLAINTEXT, SSL, SASL_SSL, etc.).

open as a page

How should you size a Kafka broker's JVM heap versus the OS page cache, and why does Kafka deliberately keep its heap small?

level: middleimportance: must knowfreq 72%

basics

~20 s

Give the Kafka broker a small JVM heap (commonly 6 GB) and leave most of the machine's RAM free for the OS page cache, which caches log file data. Kafka relies on the page cache, not heap, for read/write throughput.

open as a page

How does the cleaner handle tombstones, and what role does delete.retention.ms play?

level: middleimportance: must knowfreq 50%

basics

~20 s

A tombstone is a record with a real key but a null value, signaling 'this key is deleted'. The cleaner keeps tombstones around for delete.retention.ms (default 24h) so consumers can observe the deletion, then removes them in a later pass.

open as a page

How do fetch.min.bytes and fetch.max.wait.ms drive DelayedFetch, and when does the broker decide to park a fetch versus answer it immediately?

level: middleimportance: must knowfreq 50%

basics

~20 s

If a fetch can already return at least fetch.min.bytes of data, the broker replies immediately. If not, it parks a DelayedFetch and waits up to fetch.max.wait.ms for more data to accumulate; if that timer fires first, it returns whatever it has (possibly empty).

open as a page

Walk through how a produce request with acks=all flows through DelayedProduce and what triggers its completion.

level: middleimportance: must knowfreq 48%

basics

~20 s

The leader writes the records to its log, then parks a DelayedProduce in purgatory keyed by each target partition. When followers replicate and the high watermark advances past the produced offsets for all partitions, the operation completes and the broker acknowledges the producer.

open as a page

Walk through the full lifecycle of a single client request inside a Kafka broker, from the TCP socket to the response being written back. Name each thread pool and queue it passes through.

level: middleimportance: must knowfreq 70%

basics

~20 s

An acceptor thread accepts the TCP connection and hands it to a network (processor) thread, which reads the request and puts it on a shared request queue. A request-handler (I/O) thread takes it, KafkaApis processes it, and the response goes back through the network thread to the client.

open as a page

What are the RemoteStorageManager (RSM) and RemoteLogMetadataManager (RLMM) plugins, and how do their responsibilities differ?

level: middleimportance: must knowfreq 60%

basics

~10 s

RemoteStorageManager moves the actual segment bytes to/from remote storage (the data plane). RemoteLogMetadataManager tracks metadata — which segments exist remotely, their offset/epoch ranges, and their state (the control plane).

open as a page

What is the Kafka controller, and how does its role differ between ZooKeeper mode and KRaft mode?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The controller is the broker/node responsible for cluster-wide management: partition leader election, tracking broker liveness, and propagating metadata. In ZooKeeper mode one elected broker is the controller; in KRaft, dedicated controller nodes run a Raft quorum that stores metadata in a log.

open as a page

What are the checkpoint files in a Kafka log directory (recovery-point-offset-checkpoint and replication-offset-checkpoint), and what does each track?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Each log directory holds small checkpoint files. recovery-point-offset-checkpoint records, per partition, the offset up to which data has been flushed to disk (so recovery can skip it). replication-offset-checkpoint records the high watermark, the offset up to which messages are replicated/committed.

open as a page

What configuration controls when a segment rolls (closes and a new one opens), and what are the trade-offs of tuning it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A segment rolls when it hits segment.bytes (default 1 GiB) or segment.ms (default 7 days) — whichever comes first. Smaller values roll more often, giving finer retention but more files and open handles; larger values mean fewer, bigger files.

open as a page

How does the cleaner decide which log to compact next, and what is min.cleanable.dirty.ratio?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Each cleaner thread picks the log with the highest 'dirty ratio' — the fraction of the cleanable log that hasn't been compacted yet. A log only becomes eligible once that ratio reaches min.cleanable.dirty.ratio (default 0.5).

open as a page

What is zero-copy in Kafka, which syscall provides it, and what copies/context-switches does it eliminate on the consumer read path?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Zero-copy lets Kafka send file data from the page cache straight to the network socket using the sendfile syscall (via Java's FileChannel.transferTo). It skips copying the bytes into the application (user space) and back, saving CPU and memory bandwidth.

open as a page

How do num.io.threads and num.network.threads affect broker throughput, and how would you decide whether to increase one of them?

level: seniorimportance: must knowfreq 60%

basics

~20 s

num.network.threads sizes the network (processor) pool that does socket I/O; num.io.threads sizes the request-handler pool that does the actual work (log reads/writes). Use the idle-percent metrics: if a pool's average idle percent is low, it is saturated and you raise its thread count.

open as a page

Trace the remote-fetch read path: what happens inside a broker when a consumer requests an offset that lives only in remote storage?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The broker sees the requested offset is below the local log's start, asks the RemoteLogMetadataManager which remote segment holds it, uses the remote offset index to find the byte position, streams the bytes back via the RemoteStorageManager, and returns them in the fetch response — transparently to the consumer.

open as a page

How does a Kafka broker know on startup whether it was shut down cleanly, and why does that distinction matter?

level: juniorimportance: should knowfreq 35%

basics

~20 s

On a clean shutdown the broker flushes all data, updates the checkpoint files, and writes a clean-shutdown marker file in each log dir. On startup, if that marker is present it skips log recovery; if it is missing (a crash), it runs recovery.

open as a page

What is the request purgatory in a Kafka broker, and why does it exist?

level: juniorimportance: should knowfreq 35%

basics

~20 s

Purgatory is a broker-side holding area for requests that can't be answered immediately. Instead of blocking a thread, the broker parks the request there and completes it later when a condition is met or a timeout expires.

open as a page

What is the difference between the acceptor thread and the network (processor) threads in a Kafka broker, and how many of each exist?

level: juniorimportance: should knowfreq 45%

basics

~20 s

The acceptor thread only accepts new client connections (one per listener). Network/processor threads handle reading and writing data on existing connections; there are num.network.threads of them (default 3), and the acceptor spreads new connections across them round-robin.

open as a page

What does configuring multiple log.dirs (JBOD) give you on a Kafka broker, and how are partitions placed across the directories?

level: middleimportance: should knowfreq 55%

basics

~20 s

Setting multiple paths in log.dirs (JBOD = Just a Bunch Of Disks) lets one broker spread partition data across several independent disks for more total capacity and I/O. Kafka assigns each new partition to the directory with the fewest partitions.

open as a page

How are segment files named, and what does the filename tell you about the records inside?

level: middleimportance: should knowfreq 45%

basics

~20 s

A segment's files are named after its base offset — the offset of the first record in that segment — zero-padded to 20 digits, e.g. 00000000000000006000.log. So the filename tells you the lowest offset that segment can contain.

open as a page

What guarantees does Kafka provide about offset ordering within a partition, and how do segments preserve it?

level: middleimportance: should knowfreq 40%

basics

~20 s

Within a single partition, offsets are strictly increasing and assigned in append order — record N+1 always has a higher offset than record N. Across partitions there is no ordering. Segments preserve this because each new segment's base offset continues where the previous one ended.

open as a page

Explain how Kafka's append-only log and sequential disk writes make it fast even on spinning disks.

level: middleimportance: should knowfreq 55%

basics

~20 s

Kafka only appends to the end of partition log files instead of writing in random places. Sequential writes are far faster than random writes (no disk seeks), so even HDDs can sustain high throughput, and writes go through the page cache first.

open as a page

What is the role of KafkaApis in the request pipeline, and how does it decide what to do with a request?

level: middleimportance: should knowfreq 35%

basics

~10 s

KafkaApis is the broker class that actually handles requests. A request-handler thread calls KafkaApis.handle(), which looks at the request's API key (Produce, Fetch, Metadata, etc.) and dispatches to the matching handler method.

open as a page

What is index.interval.bytes, and how does it govern the density of Kafka's sparse indexes? What are the tradeoffs of changing it?

level: middleimportance: should knowfreq 40%

basics

~20 s

index.interval.bytes (default 4096) is how many bytes of log Kafka appends before adding a new index entry. Smaller means denser indexes and faster lookups but larger files; larger means sparser indexes, smaller files, slightly slower lookups.

open as a page

showing 1–30 of 45