skip to content

Ecosystem and Platform Choices

What surrounds Kafka: ksqlDB, the non-JVM client ecosystem, managed offerings, Kafka-compatible alternatives, and where a queue beats a log. Interviewers ask to see whether you can choose a platform, not just operate one.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What is the fundamental architectural difference between how Apache Pulsar and Apache Kafka store and serve data?

level: juniorimportance: must knowfreq 70%

answer

  1. Kafka = coupled compute+storage
  2. Pulsar = stateless brokers + BookKeeper
  3. bookies store ledgers
  4. broker death = no data move in Pulsar
  5. extra hop, more components

basics

~10 s

Kafka couples storage and serving: each broker owns its partition data on local disk. Pulsar separates them: brokers are stateless serving nodes, and storage lives in a separate layer (Apache BookKeeper).

solid answer

~40 s

Kafka is a single-tier system: a broker both serves clients and stores the partition logs on its own local disk, so compute and storage scale together and are tightly bound. Pulsar is two-tier: brokers are stateless compute nodes that handle producers/consumers but persist nothing locally; durable storage is delegated to Apache BookKeeper, a distributed write-ahead-log service made of nodes called bookies. Because brokers hold no data, you can add brokers to scale throughput or replace a failed broker instantly without moving any data. Conversely you scale storage by adding bookies independently. The cost is more moving parts (brokers + bookies + metadata in ZooKeeper/etcd) and an extra network hop on the read/write path. Kafka's coupled design is simpler and has fewer hops but rebalancing a partition means physically copying its log between brokers.

go deeper

for a junior

Know the one-liner: Kafka brokers store their own data; Pulsar brokers are stateless and BookKeeper stores the data.

for a middle

Explain the consequences: cheap broker failover, independent scaling of compute vs storage, extra hop and more components.

for a senior

Compare against Kafka tiered storage, discuss when disaggregation actually helps, and reason about which tier saturates first.

for a principal

Frame the cost/benefit for a fleet: operational surface, failure-domain isolation, and whether the elasticity justifies the added moving parts for the org's workload.

**The core question is: where does the data live, and who serves it?** **Kafka (single-tier / coupled compute+storage).** A Kafka cluster is a set of *brokers*. A topic is split into *partitions*; each partition is an append-only log file. Every partition has a *leader broker* that owns the authoritative copy on that broker's local disk, plus *follower* replicas on other brokers. Producers and consumers talk to the leader broker, and that same broker reads/writes the log from its own disk. So in Kafka a broker is simultaneously the compute (serving) node and the storage node for its partitions. Consequences: (1) to grow storage you add brokers and rebalance partitions, which physically copies log segments across the network; (2) a broker failure means its partitions must elect a new leader among replicas that already hold copies; (3) compute and storage capacity are bought together even if you only need one. **Pulsar (two-tier / disaggregated).** Pulsar splits the two jobs: - **Brokers** are *stateless* serving nodes. They accept producer writes and feed consumers, manage subscriptions and dispatching, but keep **no** topic data on local disk. - **Apache BookKeeper** is the storage tier. It is a distributed, low-latency write-ahead log service whose nodes are called **bookies**. The unit of storage is a **ledger** (an append-only sequence of entries). A broker writes incoming messages to a ledger spread across several bookies. - **Metadata** (which broker owns which topic, ledger locations, cursors) lives in ZooKeeper (older) or etcd. Because a Pulsar broker stores nothing, you can: kill a broker and have another broker pick up its topics in milliseconds (no data copy — the data is in BookKeeper already); add brokers purely to add serving throughput; add bookies purely to add storage/IO. This is the classic *separation of compute and storage* benefit. **Trade-offs.** Pulsar's design adds an extra network hop (client → broker → bookies) and more components to operate (brokers, bookies, metadata store, plus a function worker if you use Pulsar Functions). Kafka's tighter coupling means fewer hops and a simpler ops surface, at the price of expensive partition rebalancing and tying storage growth to broker growth. Modern Kafka has narrowed the gap with **tiered storage (KIP-405)** which offloads cold segments to object storage, but the *hot* path remains broker-local; it is not full storage disaggregation like Pulsar's. **Edge cases to know.** A Pulsar broker failure is cheap (stateless), but a *bookie* failure triggers BookKeeper's auto-recovery/re-replication to restore the configured replication. In Kafka, the analogous event is under-replicated partitions and leader re-election. Also, Pulsar's two tiers can be scaled in the wrong direction relative to your bottleneck (e.g. adding brokers when bookies are the IO bottleneck), so you must know which tier is saturated.

  • If a Pulsar broker crashes, what has to happen for its topics to keep serving, and why is it fast?
    Topic ownership is reassigned to another broker via the metadata store. It is fast because brokers are stateless — the new broker just starts reading/writing the existing BookKeeper ledgers; no log data is copied, unlike Kafka leader election which relies on replicas that already hold the data but is still partition-by-partition.
  • Does Kafka's tiered storage (KIP-405) make it equivalent to Pulsar's separation?
    No. KIP-405 offloads only cold/closed log segments to object storage; the active segment and serving still happen on the local broker disk. It is not full compute-storage disaggregation — the hot path stays broker-coupled.

saying these in an interview costs you the question

  • Saying Pulsar brokers store partition data locally
  • Claiming BookKeeper is just Pulsar's replication of Kafka's ISR
  • Saying Pulsar has no metadata store / is simpler to operate than Kafka
  • Confusing bookies with brokers

context

open as a page

What is librdkafka, and what is its relationship to the Kafka clients available for Python, Go, .NET, and Node.js?

level: juniorimportance: must knowfreq 70%

basics

~20 s

librdkafka is a C/C++ implementation of the Kafka client protocol. The official Python, Go, .NET, and Node.js clients are thin language bindings that wrap librdkafka, so they share one battle-tested core instead of each reimplementing Kafka.

open as a page

What is Kafka Connect, and how do sink connectors get data from Kafka topics into a data lake or warehouse like S3, BigQuery, or Snowflake?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Kafka Connect is a framework for moving data between Kafka and external systems without custom code. A sink connector reads records from Kafka topics and writes them to a destination like S3, BigQuery, or Snowflake using ready-made plugins and config.

open as a page

What is the Azure Event Hubs Kafka endpoint, and how does an existing Kafka application connect to it?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Azure Event Hubs exposes a Kafka-compatible endpoint on port 9093, so a normal Kafka client can produce and consume by just changing bootstrap.servers and using SASL/SSL auth with a connection string — no Kafka broker is actually run.

open as a page

What are Redpanda and WarpStream, and what does it mean that they are 'Kafka-compatible'?

level: juniorimportance: must knowfreq 55%

basics

~20 s

They are alternative streaming systems that speak the Kafka wire protocol, so existing Kafka clients work unchanged. Redpanda is a C++ broker with no JVM or ZooKeeper; WarpStream stores data in S3 with stateless brokers.

open as a page

What is the fundamental difference between Kafka's log-based model and a traditional message queue like RabbitMQ or ActiveMQ?

level: juniorimportance: must knowfreq 85%

basics

~10 s

Kafka stores messages as an append-only log that consumers read by position (offset); messages stay for a retention period. A traditional queue deletes each message once a consumer acknowledges it, so reading is destructive.

open as a page

In ksqlDB, what is the difference between a STREAM and a TABLE, and when would you use each?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A STREAM is an unbounded, append-only sequence of independent events (an immutable log). A TABLE represents the latest value per key (a mutable, upserted view). Use a STREAM for facts/events, a TABLE for current state.

open as a page

What is a managed Kafka offering, and what kinds of operational work does it take off your hands compared to running Apache Kafka yourself?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A managed Kafka offering is Kafka run for you by a cloud vendor (e.g. Confluent Cloud, Amazon MSK, Aiven). The vendor handles installing brokers, patching, scaling, backups, and hardware failures, so you mostly just create topics and connect clients.

open as a page

What is the fundamental architectural difference between Kafka Streams and engines like Apache Flink or Spark Structured Streaming, and why does it matter operationally?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Kafka Streams is a Java library you embed inside your own application, with no separate cluster. Flink and Spark are standalone clusters you deploy, scale, and operate on the side, each with its own master and worker processes.

open as a page

How do Pulsar's subscription modes let one system act as both a message queue and a streaming log, and how does that compare to Kafka consumer groups?

level: middleimportance: must knowfreq 55%

basics

~20 s

Pulsar offers four subscription modes — Exclusive, Failover, Shared, and Key_Shared. Shared/Key_Shared fan messages out across consumers like a work queue; Exclusive/Failover preserve per-partition order like a streaming log. Kafka only has the streaming-log model via consumer groups.

open as a page

How does Kafka achieve fan-out to multiple independent consumers, and how does that differ from a fanout exchange in RabbitMQ or SNS-to-SQS?

level: middleimportance: must knowfreq 70%

basics

~10 s

In Kafka each consumer group has its own offsets, so many groups read the same topic without copying data. RabbitMQ/SNS fan out by duplicating each message into multiple queues at publish time.

open as a page

What do CREATE STREAM AS SELECT (CSAS) and CREATE TABLE AS SELECT (CTAS) do, and what runs underneath when you execute one?

level: middleimportance: must knowfreq 70%

basics

~20 s

CSAS and CTAS create a new, continuously-updated derived stream or table from a SELECT over existing ones. Each launches a persistent query — a long-running Kafka Streams job — that writes results into a new backing Kafka topic.

open as a page

Explain the difference between a push query (EMIT CHANGES) and a pull query in ksqlDB.

level: middleimportance: must knowfreq 72%

basics

~20 s

A push query (EMIT CHANGES) subscribes to a continuous stream of updates and never finishes until cancelled — it pushes new results as they arrive. A pull query is a point-in-time lookup against a materialized table's current state and returns immediately.

open as a page

How does a schema-aware Kafka serializer (e.g. the Avro/Protobuf serializer with Schema Registry) lay out bytes on the wire, and why does the format matter for cross-client interoperability?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A schema-aware serializer registers the schema in Schema Registry, gets back an integer schema ID, and writes a small header — a 0x00 magic byte plus the 4-byte big-endian schema ID — in front of the serialized payload. Any client (JVM or librdkafka) that follows this same wire format can look up the ID and deserialize, which is what makes Avro/Protobuf data interoperable across languages.

open as a page

How does a Debezium CDC pipeline move database changes through Kafka into an analytics lakehouse, and what makes it preferable to periodic batch extracts?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Debezium is a Kafka Connect source connector that reads a database's transaction log and emits each insert/update/delete as a Kafka event. Sink connectors then land those change events in the lake, giving near-real-time, low-load, complete change history instead of nightly full-table dumps.

open as a page

Where do vendor Kafka-protocol endpoints like Azure Event Hubs diverge from Apache Kafka in transactions, compaction, and admin operations, and how would you detect this before migrating?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Protocol-compatible endpoints implement the produce/consume path but often lack full Kafka transactions/exactly-once, log compaction (cleanup.policy=compact), and parts of the AdminClient. Detect gaps by testing transactional.id, compacted topics, and admin calls against the endpoint before migrating.

open as a page

How does each engine achieve exactly-once processing in a Kafka pipeline, and what is the scope of that guarantee?

level: seniorimportance: must knowfreq 70%

basics

~10 s

Kafka Streams uses Kafka transactions (set processing.guarantee=exactly_once_v2) so reads, state updates, and writes commit atomically. Flink uses checkpoint barriers plus a two-phase-commit Kafka sink. Spark uses checkpointed offsets with idempotent/transactional sinks.

open as a page

You are choosing a stream processing engine for a Kafka-centric platform. What concrete factors drive the choice between Kafka Streams, Flink, and Spark Structured Streaming, and when would you pick each?

level: principalimportance: must knowfreq 65%

basics

~20 s

Pick Kafka Streams for Kafka-to-Kafka microservice logic owned by one team with no extra cluster. Pick Flink for large, complex, low-latency stateful jobs and multi-source joins on a shared platform. Pick Spark when you already run Spark and want unified batch plus streaming SQL.

open as a page

What is the Confluent REST Proxy, and when would you use it to produce or consume over HTTP instead of a native Kafka client?

level: middleimportance: should knowfreq 50%

basics

~20 s

The Confluent REST Proxy is an HTTP service that sits in front of a Kafka cluster, letting clients produce and consume via REST/JSON instead of the binary Kafka protocol. Use it for environments that can't run a native client — restricted languages, serverless/edge, or firewalled HTTP-only networks — at the cost of higher latency and lower throughput.

open as a page

Explain the pattern of using Kafka as the streaming backbone for ELT into a lakehouse, and how Iceberg/lake table formats fit into it.

level: middleimportance: should knowfreq 45%

basics

~20 s

Kafka acts as the central event highway: producers and CDC feed raw events in, sink connectors land them in cheap lake storage, and transformations run afterward in the warehouse/lake (the 'T' of ELT). Table formats like Apache Iceberg make those landed files behave like real, queryable, transactional tables.

open as a page

How does the Confluent S3 sink connector lay out files in object storage, and how can it achieve exactly-once delivery to S3?

level: middleimportance: should knowfreq 50%

basics

~20 s

The S3 sink groups records by partitioner (by Kafka partition or by time) and writes batched objects (Parquet/Avro/JSON) when flush.size or a rotate interval is hit. It achieves exactly-once by naming files deterministically from offsets, so replays overwrite the same object instead of duplicating.

open as a page

How do AMQP and MQTT protocol edges fronting Kafka work, and when would you choose them over the native Kafka protocol?

level: middleimportance: should knowfreq 28%

basics

~20 s

An AMQP or MQTT edge is a gateway that accepts those protocols at the boundary and forwards messages into Kafka. Use them when clients can't speak Kafka — e.g. IoT/MQTT devices or AMQP messaging apps — trading native Kafka features for broad client reach.

open as a page

When would you put a Google Pub/Sub-to-Kafka bridge (or the Pub/Sub Kafka shim) between a publisher and a Kafka cluster, and what are the trade-offs?

level: middleimportance: should knowfreq 30%

basics

~20 s

Use a Pub/Sub-to-Kafka bridge to connect Google Pub/Sub and Kafka ecosystems without rewriting apps — e.g. fan messages from Pub/Sub into Kafka topics. Trade-offs: extra hop adds latency, the bridge can break ordering/exactly-once, and it's another component to run.

open as a page

How does the Event Hubs Throughput-Unit / Processing-Unit capacity model differ from sizing an Apache Kafka cluster, and how does that interact with partitions?

level: middleimportance: should knowfreq 40%

basics

~20 s

On Apache Kafka you size brokers, disks, and partitions. On Event Hubs you buy Throughput Units (or Processing/Capacity Units) that cap MB/s and events/s regardless of partition count; partitions are fixed at creation and mainly affect parallelism, not capacity.

open as a page

How does Redpanda's architecture (C++ thread-per-core, no JVM, no ZooKeeper, Raft) differ from Apache Kafka's, and what does each choice buy you?

level: middleimportance: should knowfreq 45%

basics

~20 s

Redpanda is one C++ binary using a thread-per-core model, with no JVM (no GC pauses) and no ZooKeeper. It uses Raft for both metadata and partition replication, giving lower, more predictable tail latency and simpler operations than JVM-based Kafka.

open as a page

Explain offset-based replay in Kafka: how does a consumer reprocess past data, and why can't a destructive-consume broker do the same?

level: middleimportance: should knowfreq 50%

basics

~20 s

A Kafka consumer tracks a position (offset) and can seek backward to an earlier offset or timestamp to re-read records that are still within retention. A queue deletes messages on ack, so there's nothing left to re-read.

open as a page

Compare Amazon MSK (provisioned) with MSK Serverless: what does each abstract, and when would you choose one over the other?

level: middleimportance: should knowfreq 55%

basics

~20 s

Provisioned MSK gives you sized broker instances you pick and pay for by the hour; you control instance type and broker count. MSK Serverless hides brokers and capacity entirely and bills on throughput/storage, auto-scaling. Choose Serverless for spiky/unknown load, provisioned for steady high throughput where per-unit cost is lower.

open as a page

Where does stateful windowing state live in Kafka Streams versus Flink versus Spark, and how is it made fault-tolerant?

level: middleimportance: should knowfreq 55%

basics

~20 s

Kafka Streams keeps state in embedded RocksDB on each instance, backed up to compacted changelog topics in Kafka. Flink keeps state in a state backend (heap or RocksDB) snapshotted to durable storage via checkpoints. Spark stores state in a checkpointed, versioned state store.

open as a page

Describe Pulsar's native multi-tenancy and geo-replication model, and contrast it with how Kafka achieves the same goals.

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pulsar bakes in a tenant/namespace/topic hierarchy with per-tenant quotas, auth, and isolation, and supports cross-region replication by configuring clusters on a namespace. Kafka has a flat topic namespace and relies on external tools (ACLs, MirrorMaker 2) to approximate both.

open as a page

Explain Pulsar's segment-centric storage versus Kafka's partition-centric log, and why it matters operationally.

level: seniorimportance: should knowfreq 45%

basics

~20 s

In Kafka a partition's whole log lives on the broker that leads it, so the partition is the storage unit. In Pulsar a topic is a sequence of small segments (BookKeeper ledgers) spread across many bookies, so no single node holds the whole topic.

open as a page

showing 1–30 of 48