Compare Amazon MSK (provisioned) with MSK Serverless: what does each abstract, and when would you choose one over the other?
answer
- Provisioned = pick instance type + broker count
- Serverless = no brokers, auto-scale, quotas
- Serverless bills per-partition-hour + GB
- Serverless = IAM auth only
- Steady/high → provisioned; spiky → serverless
basics
~20 sProvisioned MSK gives you sized broker instances you pick and pay for by the hour; you control instance type and broker count. MSK Serverless hides brokers and capacity entirely and bills on throughput/storage, auto-scaling. Choose Serverless for spiky/unknown load, provisioned for steady high throughput where per-unit cost is lower.
solid answer
~50 sProvisioned **Amazon MSK** runs actual Kafka brokers on EC2 instance types you choose (e.g. kafka.m5.large), with a broker count you set; AWS manages patching, ZooKeeper/KRaft, and node replacement, but you size and pay per broker-hour plus storage. You can tune many broker configs and use provisioned/tiered storage. **MSK Serverless** removes capacity planning: no instance types, no broker count, automatic partition/throughput scaling, billed per cluster-hour plus per-partition-hour, ingress/egress GB, and storage. It enforces quotas and uses IAM auth only. Choose **provisioned** for steady, high, predictable throughput where reserved capacity is cheaper and you need config control or specific storage tuning. Choose **Serverless** for spiky, unpredictable, or low-baseline workloads, dev/test, or teams wanting zero capacity planning. The trade-off is control and per-unit cost (provisioned) versus elasticity and operational simplicity (serverless), with serverless carrying per-partition pricing that punishes very high partition counts.
go deeper
Know that provisioned means you pick servers and serverless means AWS scales for you.
Explain the billing models (broker-hour vs per-partition/per-GB) and the steady-vs-spiky decision.
Reason about auth restrictions, quotas, partition-count cost, and config control when choosing a flavor.
Model TCO across utilization curves and govern when teams should default to serverless vs reserve provisioned capacity.
## Background Amazon **MSK** (Managed Streaming for Apache Kafka) is AWS's managed Kafka. It comes in two flavors that abstract different amounts. ## MSK Provisioned You create a cluster by choosing: - **Broker instance type** (e.g. `kafka.t3.small`, `kafka.m5.large`, `kafka.m7g` Graviton) — this fixes CPU/RAM/network per broker. - **Number of brokers** (a multiple of the number of AZs). - **Storage** per broker (EBS), with optional **provisioned throughput** and **tiered storage** (hot data on EBS, older data offloaded to a cheaper tier, reducing local disk needs). AWS manages the rest: OS/Kafka patching, the metadata layer (ZooKeeper historically; newer versions use **KRaft**), broker replacement, and monitoring via CloudWatch. You retain a good deal of **config control** through MSK configuration objects (many `server.properties` settings). **Billing** is per **broker-hour** by instance type plus storage (and provisioned-throughput if enabled). Cost is decoupled from actual message volume — you pay for capacity whether or not you use it, which makes it cheap per unit at high steady utilization and wasteful when idle. ## MSK Serverless Serverless removes capacity planning entirely: - **No instance type, no broker count** — you just create a cluster and topics. - **Automatic scaling** of throughput and partitions up to account/cluster quotas (e.g. ingress/egress and partition limits). - **IAM authentication only** (no SASL/SCRAM or mTLS options). - **Billing** is usage-based: per **cluster-hour**, per **partition-hour**, **ingress and egress GB**, and **storage GB-month**. Because partitions are billed individually, a design with thousands of partitions can become expensive on Serverless even at modest throughput. ## How to choose - **Steady, high, predictable throughput**: provisioned is usually cheaper per unit and gives config/storage control. Reserve capacity, run hot. - **Spiky, bursty, or unknown load; dev/test; low baseline**: serverless avoids over-provisioning and the auto-scaling absorbs spikes. - **Need specific broker configs or non-IAM auth**: provisioned (serverless restricts both). - **Very high partition counts**: watch serverless per-partition pricing; provisioned may be cheaper. ## Edge cases / gotchas - Serverless has hard **quotas** (max partitions, max throughput) — a workload can outgrow them. - Provisioned requires you to right-size; under-sizing causes throttling and over-sizing wastes money. - Both integrate with **MSK Connect** (managed Kafka Connect) and IAM, but feature parity is not perfect (serverless is more restricted).
- Why can a high-partition-count workload be surprisingly expensive on MSK Serverless?Serverless bills per partition-hour in addition to throughput and storage. Thousands of partitions cost money even at low traffic, so an over-partitioned design that is fine on provisioned (paid by broker-hour) can be costly on serverless.
- What authentication options does MSK Serverless support?IAM authentication only. Provisioned MSK additionally supports SASL/SCRAM and mTLS, so workloads requiring those must use provisioned.
saying these in an interview costs you the question
- Claiming MSK Serverless is always cheaper — per-partition and per-GB billing can exceed reserved provisioned capacity at steady high load or high partition counts.
- Saying provisioned MSK lets you ignore capacity planning — you must size instance type and broker count.
- Assuming serverless supports all auth mechanisms (it is IAM-only).