You are choosing a stream processing engine for a Kafka-centric platform. What concrete factors drive the choice between Kafka Streams, Flink, and Spark Structured Streaming, and when would you pick each?
answer
- ownership: app team vs platform team
- latency: Streams/Flink per-record vs Spark micro-batch
- Flink = big state + flexible event-time
- Spark = lakehouse + batch+stream SQL
- Streams = Kafka-to-Kafka, no cluster
basics
~20 sPick Kafka Streams for Kafka-to-Kafka microservice logic owned by one team with no extra cluster. Pick Flink for large, complex, low-latency stateful jobs and multi-source joins on a shared platform. Pick Spark when you already run Spark and want unified batch plus streaming SQL.
solid answer
~50 sDrive the decision from operational ownership, latency, state size, source/sink breadth, and existing investment. Kafka Streams fits when processing is part of an application a single team deploys, the data flow is Kafka-to-Kafka, latency must be low (per-record, no micro-batch), and you do not want to run a separate cluster — the cost is being capped by partition count and lacking rich multi-source connectors and SQL. Flink fits large, complex, very stateful, low-latency jobs: rich event-time/watermark control, true streaming, many connectors, SQL, savepoint-based upgrades and rescaling — but you must operate a Flink platform (HA, Kubernetes operator, savepoint lifecycle). Spark Structured Streaming fits organizations already on Spark/Databricks wanting one engine for batch and streaming with strong SQL and lakehouse integration (Delta/Iceberg), accepting micro-batch latency for many use cases. The meta-factor is org topology: who owns and operates the runtime.
go deeper
Know the headline picks: Streams = embedded simple, Flink = powerful cluster, Spark = batch+stream SQL.
List a few decision factors (latency, ownership, connectors) and match each engine to a use case.
Give a structured trade-off across latency, state, breadth, and existing investment with concrete picks.
Frame the choice around org topology and platform strategy, anticipate anti-patterns, and justify context-dependent decisions rather than a universal winner.
## Decision factors 1. **Operational ownership / org topology** — The biggest factor. Kafka Streams keeps processing inside an application a product team already builds, ships, and on-calls. Flink/Spark are typically a **shared platform** run by a data/platform team. Choosing the wrong boundary creates either bottleneck platform teams or product teams drowning in cluster ops. 2. **Latency** — Kafka Streams and Flink are **true (per-record) streaming**: low, predictable latency. Spark Structured Streaming is historically **micro-batch** (latency ~ trigger interval, often hundreds of ms to seconds); fine for many analytics use cases, less so for tight real-time SLAs (its low-latency modes narrow but don't fully erase this). 3. **State size and complexity** — Flink shines with very large state (RocksDB backend, incremental checkpoints) and complex event-time logic (CEP, flexible lateness). Kafka Streams handles substantial state but is bounded by partition-count parallelism and local-disk restore concerns. Spark handles big aggregations well, especially batch-shaped. 4. **Sources and sinks** — Flink and Spark have broad connector ecosystems (files, S3, JDBC, Iceberg/Delta/Hudi, many message systems) and rich SQL. Kafka Streams is **Kafka-first**; non-Kafka I/O typically goes through Kafka Connect or app code. 5. **Existing investment** — If you already run a Spark/Databricks lakehouse, Spark Structured Streaming unifies batch and streaming with one skillset and storage. If you already run Flink, extend it. 6. **Scaling and upgrades** — Flink's **savepoints** give clean rescale and code-change semantics; Spark restarts from checkpoints; Streams rescales via partition count and consumer-group rebalance. 7. **SQL and analyst access** — Flink SQL and Spark SQL let non-Java users build pipelines; Kafka Streams is code-centric (ksqlDB sits on top for SQL on Kafka but is a separate product). ## When to pick each - **Kafka Streams**: enrichment/aggregation/joins **within a microservice**, Kafka-to-Kafka, low latency, single-team ownership, no appetite for a cluster. E.g., per-order enrichment in an order service. - **Apache Flink**: large-scale, complex, low-latency stateful analytics, multi-source joins, CEP, strict event-time/lateness control, on a platform team's runtime. E.g., real-time fraud detection across many feeds. - **Spark Structured Streaming**: organizations standardized on Spark/lakehouse wanting unified batch+stream and SQL, tolerant of micro-batch latency. E.g., streaming ingestion into Delta/Iceberg tables alongside batch ETL. ## Anti-patterns / edge cases - Standing up a **Flink cluster for one small Kafka-to-Kafka transform** that Kafka Streams would do with zero extra infra. - Forcing **sub-100ms latency** onto Spark micro-batch when Flink or Streams is the right tool. - Using **Kafka Streams for joins across many heterogeneous external systems** it can't natively reach. - Ignoring **who operates the runtime** — the cheapest engine to run is the one your org already operates well. ## The meta-point There is rarely a single 'best' engine; the right answer is contextual. A senior/principal answer names the **factors and trade-offs**, ties them to **org topology and existing platforms**, and gives **concrete picks** rather than declaring one engine universally superior.
- A team needs sub-100ms latency for a Kafka-to-Kafka enrichment owned by one squad. Which engine and why?Kafka Streams (or Flink): both are true per-record streaming with low latency. Streams wins here because it needs no separate cluster and stays within the squad's app, matching single-team ownership; Flink would add cluster ops the use case doesn't justify.
- Why might an organization standardized on Databricks pick Spark Structured Streaming even for some real-time work?Operational leverage: one engine, one skillset, and direct lakehouse (Delta/Iceberg) integration for unified batch and streaming, plus rich SQL. The micro-batch latency is acceptable for many analytics use cases, and avoiding a second runtime reduces total cost of ownership.
- What is the main risk of adopting Flink for a small workload?Operating a whole distributed runtime (HA JobManager, Kubernetes operator, savepoint lifecycle, upgrades) for a job that Kafka Streams could run as a library with no extra infrastructure — disproportionate operational cost.
saying these in an interview costs you the question
- Declaring one engine universally 'best' without naming factors or context.
- Ignoring operational ownership / org topology, which is usually the dominant factor.
- Claiming Spark Structured Streaming is always as low-latency as Flink — it is historically micro-batch.
- Recommending a full Flink/Spark cluster for a trivial Kafka-to-Kafka transform Kafka Streams handles with no infra.