skip to content

What is the fundamental architectural difference between how Apache Pulsar and Apache Kafka store and serve data?

level: juniorimportance: must knowfreq 70%

answer

  1. Kafka = coupled compute+storage
  2. Pulsar = stateless brokers + BookKeeper
  3. bookies store ledgers
  4. broker death = no data move in Pulsar
  5. extra hop, more components

basics

~10 s

Kafka couples storage and serving: each broker owns its partition data on local disk. Pulsar separates them: brokers are stateless serving nodes, and storage lives in a separate layer (Apache BookKeeper).

solid answer

~40 s

Kafka is a single-tier system: a broker both serves clients and stores the partition logs on its own local disk, so compute and storage scale together and are tightly bound. Pulsar is two-tier: brokers are stateless compute nodes that handle producers/consumers but persist nothing locally; durable storage is delegated to Apache BookKeeper, a distributed write-ahead-log service made of nodes called bookies. Because brokers hold no data, you can add brokers to scale throughput or replace a failed broker instantly without moving any data. Conversely you scale storage by adding bookies independently. The cost is more moving parts (brokers + bookies + metadata in ZooKeeper/etcd) and an extra network hop on the read/write path. Kafka's coupled design is simpler and has fewer hops but rebalancing a partition means physically copying its log between brokers.

go deeper

for a junior

Know the one-liner: Kafka brokers store their own data; Pulsar brokers are stateless and BookKeeper stores the data.

for a middle

Explain the consequences: cheap broker failover, independent scaling of compute vs storage, extra hop and more components.

for a senior

Compare against Kafka tiered storage, discuss when disaggregation actually helps, and reason about which tier saturates first.

for a principal

Frame the cost/benefit for a fleet: operational surface, failure-domain isolation, and whether the elasticity justifies the added moving parts for the org's workload.

**The core question is: where does the data live, and who serves it?** **Kafka (single-tier / coupled compute+storage).** A Kafka cluster is a set of *brokers*. A topic is split into *partitions*; each partition is an append-only log file. Every partition has a *leader broker* that owns the authoritative copy on that broker's local disk, plus *follower* replicas on other brokers. Producers and consumers talk to the leader broker, and that same broker reads/writes the log from its own disk. So in Kafka a broker is simultaneously the compute (serving) node and the storage node for its partitions. Consequences: (1) to grow storage you add brokers and rebalance partitions, which physically copies log segments across the network; (2) a broker failure means its partitions must elect a new leader among replicas that already hold copies; (3) compute and storage capacity are bought together even if you only need one. **Pulsar (two-tier / disaggregated).** Pulsar splits the two jobs: - **Brokers** are *stateless* serving nodes. They accept producer writes and feed consumers, manage subscriptions and dispatching, but keep **no** topic data on local disk. - **Apache BookKeeper** is the storage tier. It is a distributed, low-latency write-ahead log service whose nodes are called **bookies**. The unit of storage is a **ledger** (an append-only sequence of entries). A broker writes incoming messages to a ledger spread across several bookies. - **Metadata** (which broker owns which topic, ledger locations, cursors) lives in ZooKeeper (older) or etcd. Because a Pulsar broker stores nothing, you can: kill a broker and have another broker pick up its topics in milliseconds (no data copy — the data is in BookKeeper already); add brokers purely to add serving throughput; add bookies purely to add storage/IO. This is the classic *separation of compute and storage* benefit. **Trade-offs.** Pulsar's design adds an extra network hop (client → broker → bookies) and more components to operate (brokers, bookies, metadata store, plus a function worker if you use Pulsar Functions). Kafka's tighter coupling means fewer hops and a simpler ops surface, at the price of expensive partition rebalancing and tying storage growth to broker growth. Modern Kafka has narrowed the gap with **tiered storage (KIP-405)** which offloads cold segments to object storage, but the *hot* path remains broker-local; it is not full storage disaggregation like Pulsar's. **Edge cases to know.** A Pulsar broker failure is cheap (stateless), but a *bookie* failure triggers BookKeeper's auto-recovery/re-replication to restore the configured replication. In Kafka, the analogous event is under-replicated partitions and leader re-election. Also, Pulsar's two tiers can be scaled in the wrong direction relative to your bottleneck (e.g. adding brokers when bookies are the IO bottleneck), so you must know which tier is saturated.

  • If a Pulsar broker crashes, what has to happen for its topics to keep serving, and why is it fast?
    Topic ownership is reassigned to another broker via the metadata store. It is fast because brokers are stateless — the new broker just starts reading/writing the existing BookKeeper ledgers; no log data is copied, unlike Kafka leader election which relies on replicas that already hold the data but is still partition-by-partition.
  • Does Kafka's tiered storage (KIP-405) make it equivalent to Pulsar's separation?
    No. KIP-405 offloads only cold/closed log segments to object storage; the active segment and serving still happen on the local broker disk. It is not full compute-storage disaggregation — the hot path stays broker-coupled.

saying these in an interview costs you the question

  • Saying Pulsar brokers store partition data locally
  • Claiming BookKeeper is just Pulsar's replication of Kafka's ISR
  • Saying Pulsar has no metadata store / is simpler to operate than Kafka
  • Confusing bookies with brokers

context