skip to content

Managed Kafka Offerings

What Confluent Cloud, Amazon MSK, Aiven and Azure Event Hubs actually manage, and how their pricing and limits differ. A build-versus-buy question that comes up in nearly every platform interview.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a managed Kafka offering, and what kinds of operational work does it take off your hands compared to running Apache Kafka yourself?

level: juniorimportance: must knowfreq 70%

answer

  1. Vendor runs brokers + patching + failures
  2. You still own topics/partitions/clients/cost
  3. Never see ZooKeeper/KRaft
  4. Confluent Cloud, MSK, Aiven, Event Hubs
  5. Less toil, more cost + lock-in

basics

~20 s

A managed Kafka offering is Kafka run for you by a cloud vendor (e.g. Confluent Cloud, Amazon MSK, Aiven). The vendor handles installing brokers, patching, scaling, backups, and hardware failures, so you mostly just create topics and connect clients.

solid answer

~40 s

Self-managed Kafka means you provision servers, install and patch Kafka and its metadata layer (ZooKeeper, now KRaft), size and rebalance the cluster, monitor disks, handle broker failures, and run upgrades. A managed offering (Confluent Cloud, Amazon MSK, Aiven, Azure Event Hubs' Kafka surface) moves most of that to the vendor: broker provisioning, OS/Kafka patching, replacing failed nodes, security defaults (TLS, IAM/SASL), monitoring, and often scaling. You still own application-level concerns: topic and partition design, client configuration, consumer-group tuning, schema/data governance, and cost from throughput and storage. The trade-off is less operational toil and faster onboarding in exchange for higher per-unit cost, less low-level control (some configs are locked), and vendor lock-in. Offerings differ in how much they abstract: MSK still exposes brokers and instance types, while MSK Serverless and Confluent Cloud hide brokers entirely.

go deeper

for a junior

Know that the vendor runs the servers and handles failures/patching, and that you still create topics and connect apps.

for a middle

Articulate the concrete operational list that shifts to the vendor versus what stays with the application team.

for a senior

Distinguish degrees of abstraction across offerings (plain MSK vs serverless vs Confluent) and the lock-in/control trade-off.

for a principal

Frame managed-vs-self as a TCO and risk decision: operational toil, SLA, governance, exit cost, and where the abstraction leaks.

## What Apache Kafka is Apache Kafka is a distributed event-streaming platform: producers write records to **topics**, which are split into **partitions** spread across **brokers** (server processes). Consumers read from partitions, and Kafka keeps a durable, replicated log on disk. Running it well requires real operational skill. ## What 'self-managed' costs you If you run Kafka yourself you are responsible for: - **Provisioning** the broker machines (VMs/containers) and storage. - **Metadata layer**: historically a separate **ZooKeeper** ensemble for cluster coordination; modern Kafka replaces this with **KRaft** (Kafka Raft, KIP-500), where brokers/controllers manage metadata internally. Either way you must operate it. - **Patching and upgrades** of the OS and Kafka itself, ideally with no downtime. - **Scaling**: adding brokers and **rebalancing partitions** (e.g. via Cruise Control) so load is even. - **Failure handling**: replacing dead brokers, re-replicating data, watching for under-replicated partitions. - **Security**: TLS, SASL/mTLS auth, ACLs, secrets. - **Monitoring**: lag, disk, ISR (in-sync replicas), throughput. ## What 'managed' abstracts A managed offering is Kafka operated **as a service** by a vendor. The vendor takes over infrastructure and operations: provisioning, patching, node replacement, the metadata layer (you never see ZooKeeper/KRaft), secure defaults, and monitoring. The major options are **Confluent Cloud** (from Kafka's original creators, most fully managed), **Amazon MSK** and **MSK Serverless** (AWS), **Aiven for Apache Kafka** (multi-cloud), and **Azure Event Hubs**, which exposes a **Kafka-protocol surface** (a Kafka-compatible API on top of Event Hubs, not Kafka brokers). ## What stays yours Managed does **not** mean responsibility-free. You still design topics and partition counts, configure clients (acks, retries, idempotence), tune consumer groups and lag, manage **schemas** and governance, and control **cost**, which on most offerings tracks throughput, partition count, storage, and data transfer. ## Trade-offs - **Pros**: dramatically less operational toil, fast onboarding, vendor-backed SLAs, secure defaults. - **Cons**: higher per-unit price, **less control** (some broker configs are locked or unavailable), and **lock-in** (proprietary features, IAM models, connectors). ## Edge cases Degree of abstraction varies: plain MSK still shows you broker instance types and counts (you size the cluster); MSK Serverless and Confluent Cloud hide brokers entirely and bill on usage. Event Hubs is API-compatible but not feature-complete Kafka (e.g. some admin operations and configs behave differently).

  • Name one responsibility that stays with you even on a fully managed Kafka.
    Plenty: topic/partition design, client tuning (acks/idempotence), consumer-group/lag management, schema governance, and controlling cost from throughput/storage. The vendor manages infrastructure, not your data model or application behavior.
  • Does 'managed' always mean you never see brokers?
    No. Plain Amazon MSK still exposes broker instance types and counts that you size. MSK Serverless and Confluent Cloud hide brokers and bill on usage. Degree of abstraction differs per offering.

saying these in an interview costs you the question

  • Saying managed Kafka removes all your responsibilities, including partition design and client tuning.
  • Claiming Azure Event Hubs is full Apache Kafka rather than a Kafka-compatible protocol surface.
  • Assuming every managed offering hides brokers (plain MSK does not).

context

open as a page

Compare Amazon MSK (provisioned) with MSK Serverless: what does each abstract, and when would you choose one over the other?

level: middleimportance: should knowfreq 55%

basics

~20 s

Provisioned MSK gives you sized broker instances you pick and pay for by the hour; you control instance type and broker count. MSK Serverless hides brokers and capacity entirely and bills on throughput/storage, auto-scaling. Choose Serverless for spiky/unknown load, provisioned for steady high throughput where per-unit cost is lower.

open as a page

Azure Event Hubs exposes a 'Kafka protocol surface.' What does that mean, and what Kafka features or behaviors should you NOT assume are present?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Event Hubs isn't Apache Kafka; it's Azure's own event service that speaks the Kafka wire protocol. Existing Kafka clients can produce/consume by changing the bootstrap endpoint. But it's not a full Kafka broker, so some features (Kafka Connect-style admin, transactions, compacted topics, certain configs) may be limited or absent.

open as a page

How do pricing models differ across Confluent Cloud, MSK provisioned, MSK Serverless, and Aiven, and what cost drivers should architects watch?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They bill differently: MSK provisioned charges per broker-hour plus storage (capacity-based). MSK Serverless and Confluent Cloud charge mostly on usage — throughput in/out, partitions, storage. Aiven charges a flat plan per cluster size. Watch data-transfer/egress, partition counts, retention/storage, and idle capacity.

open as a page

Managed Kafka offerings advertise hiding ZooKeeper/KRaft and providing tiered storage. As a principal, how do these abstractions change capacity planning, scaling, and lock-in decisions?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Hiding ZooKeeper/KRaft removes metadata-cluster operations and raises practical partition limits, so you scale partitions/topics more freely. Tiered storage offloads old data to cheap object storage, so retention is no longer bounded by broker disk. Both reduce ops but deepen reliance on vendor-specific behavior and pricing.

open as a page