skip to content

Ecosystem and Platform Choices

What surrounds Kafka: ksqlDB, the non-JVM client ecosystem, managed offerings, Kafka-compatible alternatives, and where a queue beats a log. Interviewers ask to see whether you can choose a platform, not just operate one.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

What feature-parity gaps should you expect between the JVM Kafka client and non-JVM (librdkafka-based) clients, and how do you decide which to use?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The JVM client usually gets new Kafka features first and most completely — notably Kafka Streams, the full transactional/EOS API, and incremental cooperative rebalancing landed there earliest. librdkafka clients catch up later and historically lag on some advanced features, so pick the JVM client when you need the latest stream-processing or transactional capabilities.

open as a page

What is Kafka Tiered Storage (KIP-405), and how does it change the economics and architecture of using Kafka as a long-retention backbone for analytics?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Tiered Storage (KIP-405) lets a Kafka broker offload older log segments to cheap remote object storage (like S3) while keeping recent data on local disk. This makes long retention affordable and lets brokers store far more history without huge local disks.

open as a page

Explain WarpStream's S3-backed, zero-disk, stateless-broker architecture and the latency-versus-cost trade-off it makes.

level: seniorimportance: should knowfreq 40%

basics

~20 s

WarpStream brokers (Agents) keep no local data — they batch writes and flush them straight to S3, which provides durability. This removes disks and cross-AZ replication cost, but S3 round-trips raise end-to-end latency to hundreds of milliseconds instead of single-digit ms.

open as a page

Compare the ordering and retention guarantees of Kafka versus a traditional queue, and explain the trade-offs each model forces.

level: seniorimportance: should knowfreq 55%

basics

~20 s

Kafka guarantees order only within a partition and keeps data for a retention window, enabling replay. Queues with competing consumers generally don't preserve order and delete on ack, so there's no replay but messages can't pile up indefinitely.

open as a page

A team asks whether to use Kafka or a traditional message queue (RabbitMQ/SQS) for a new system. How do you decide, and what workloads favor each?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Choose Kafka for high-throughput event streams that need replay, multi-consumer fan-out, and ordered-per-key data. Choose a queue for task/work dispatch needing per-message acks, elastic workers, priorities, and rich retry/DLQ, with no need to replay history.

open as a page

What are the join types and co-partitioning requirements in ksqlDB, and how do stream-stream, stream-table, and table-table joins differ?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Stream-stream joins are windowed (you join events within a time window) and produce a stream. Stream-table joins are non-windowed enrichment lookups against the table's current state. Table-table joins maintain a continuously-updated joined table. All require co-partitioned inputs (same key, same partition count).

open as a page

How does ksqlDB materialize state for aggregations and pull queries, and what role do RocksDB state stores and changelog topics play in fault tolerance?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Aggregations build a materialized view stored in a local RocksDB state store on disk, partitioned across ksqlDB nodes. Each store is backed by a compacted Kafka changelog topic, so on failure or restart the state is rebuilt by replaying that changelog.

open as a page

How do windowed aggregations work in ksqlDB, and what are the differences between tumbling, hopping, and session windows?

level: seniorimportance: should knowfreq 55%

basics

~10 s

Windowed aggregations group events into time buckets before aggregating. Tumbling windows are fixed-size, non-overlapping; hopping windows are fixed-size but overlap by a smaller advance; session windows are activity-based, closing after a gap of inactivity.

open as a page

Azure Event Hubs exposes a 'Kafka protocol surface.' What does that mean, and what Kafka features or behaviors should you NOT assume are present?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Event Hubs isn't Apache Kafka; it's Azure's own event service that speaks the Kafka wire protocol. Existing Kafka clients can produce/consume by changing the bootstrap endpoint. But it's not a full Kafka broker, so some features (Kafka Connect-style admin, transactions, compacted topics, certain configs) may be limited or absent.

open as a page

How do pricing models differ across Confluent Cloud, MSK provisioned, MSK Serverless, and Aiven, and what cost drivers should architects watch?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They bill differently: MSK provisioned charges per broker-hour plus storage (capacity-based). MSK Serverless and Confluent Cloud charge mostly on usage — throughput in/out, partitions, storage. Aiven charges a flat plan per cluster size. Watch data-transfer/egress, partition counts, retention/storage, and idle capacity.

open as a page

Explain event-time versus processing-time and the role of watermarks in stateful windowing. How do Kafka Streams, Flink, and Spark differ in handling late and out-of-order data?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Event-time is when an event actually happened; processing-time is when the engine sees it. Watermarks estimate how far event-time has progressed so windows know when to close. Flink uses explicit watermarks; Spark uses withWatermark; Kafka Streams uses a grace period instead of true watermarks.

open as a page

As an architect, when would you choose Apache Pulsar over Kafka, and when does Kafka's partition-centric design remain the better choice?

level: principalimportance: should knowfreq 50%

basics

~20 s

Choose Pulsar when you need unified queue+stream semantics, native multi-tenancy, built-in geo-replication, or independent compute/storage scaling and huge topic counts. Stick with Kafka for its mature ecosystem, simpler ops, and when partition-based streaming with a rich connector/stream-processing toolchain fits.

open as a page

How does Kafka's wire-protocol API versioning let clients of different versions and languages interoperate with brokers, and what is the role of ApiVersions negotiation?

level: principalimportance: should knowfreq 35%

basics

~20 s

Each Kafka request type (Produce, Fetch, Metadata, etc.) has its own API key and an independently incrementing version number. On connect, a client sends an ApiVersions request; the broker replies with the min/max version it supports per API key. The client then picks the highest version both sides support. This per-API negotiation lets old/new and JVM/non-JVM clients interoperate with brokers across versions.

open as a page

As an architect, how do you decide between a vendor Kafka-protocol endpoint (e.g. Event Hubs) and self-managed/managed Apache Kafka, and how do you minimize lock-in?

level: principalimportance: should knowfreq 35%

basics

~20 s

Pick a vendor Kafka endpoint when you want zero ops and your workload uses only the supported produce/consume subset; pick real Kafka when you need transactions, compaction, full admin, or portability. Minimize lock-in by isolating vendor specifics, avoiding native-only features, and testing portability.

open as a page

As an architect, how would you decide between Apache Kafka, Redpanda, and WarpStream for a given workload?

level: principalimportance: should knowfreq 35%

basics

~20 s

Match the architecture to the dominant constraint: pick Redpanda when low/predictable latency and simple ops matter; pick WarpStream when cost (cross-AZ + disk) and elasticity dominate and you can tolerate hundreds-of-ms latency; stay on Kafka for maximum ecosystem maturity and full feature/KIP coverage.

open as a page

What delivery and schema semantics should you expect from the Snowflake and BigQuery Kafka Connect sinks, and how do they differ from a file-based S3 sink?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Warehouse sinks (Snowflake, BigQuery) write rows directly into managed tables rather than files in a bucket. They use the warehouse's streaming-ingest APIs, can offer exactly-once or near-exactly-once via offset tracking, and auto-map record schemas to table columns — whereas the S3 sink just writes batched files you must catalog separately.

open as a page

What practical compatibility issues should you check before migrating a Kafka application to Redpanda or WarpStream?

level: seniorimportance: nice to knowfreq 25%

basics

~10 s

Although clients are protocol-compatible, verify the supported Kafka protocol/API version range, transactions/exactly-once support, admin and Connect/Streams behavior, and operational differences (metrics, configs, durability model). Test with your real client versions before cutting over.

open as a page

Managed Kafka offerings advertise hiding ZooKeeper/KRaft and providing tiered storage. As a principal, how do these abstractions change capacity planning, scaling, and lock-in decisions?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Hiding ZooKeeper/KRaft removes metadata-cluster operations and raises practical partition limits, so you scale partitions/topics more freely. Tiered storage offloads old data to cheap object storage, so retention is no longer bounded by broker disk. Both reduce ops but deepen reliance on vendor-specific behavior and pricing.

open as a page

showing 31–48 of 48