What is a managed Kafka offering, and what kinds of operational work does it take off your hands compared to running Apache Kafka yourself?
answer
- Vendor runs brokers + patching + failures
- You still own topics/partitions/clients/cost
- Never see ZooKeeper/KRaft
- Confluent Cloud, MSK, Aiven, Event Hubs
- Less toil, more cost + lock-in
basics
~20 sA managed Kafka offering is Kafka run for you by a cloud vendor (e.g. Confluent Cloud, Amazon MSK, Aiven). The vendor handles installing brokers, patching, scaling, backups, and hardware failures, so you mostly just create topics and connect clients.
solid answer
~40 sSelf-managed Kafka means you provision servers, install and patch Kafka and its metadata layer (ZooKeeper, now KRaft), size and rebalance the cluster, monitor disks, handle broker failures, and run upgrades. A managed offering (Confluent Cloud, Amazon MSK, Aiven, Azure Event Hubs' Kafka surface) moves most of that to the vendor: broker provisioning, OS/Kafka patching, replacing failed nodes, security defaults (TLS, IAM/SASL), monitoring, and often scaling. You still own application-level concerns: topic and partition design, client configuration, consumer-group tuning, schema/data governance, and cost from throughput and storage. The trade-off is less operational toil and faster onboarding in exchange for higher per-unit cost, less low-level control (some configs are locked), and vendor lock-in. Offerings differ in how much they abstract: MSK still exposes brokers and instance types, while MSK Serverless and Confluent Cloud hide brokers entirely.
go deeper
Know that the vendor runs the servers and handles failures/patching, and that you still create topics and connect apps.
Articulate the concrete operational list that shifts to the vendor versus what stays with the application team.
Distinguish degrees of abstraction across offerings (plain MSK vs serverless vs Confluent) and the lock-in/control trade-off.
Frame managed-vs-self as a TCO and risk decision: operational toil, SLA, governance, exit cost, and where the abstraction leaks.
## What Apache Kafka is Apache Kafka is a distributed event-streaming platform: producers write records to **topics**, which are split into **partitions** spread across **brokers** (server processes). Consumers read from partitions, and Kafka keeps a durable, replicated log on disk. Running it well requires real operational skill. ## What 'self-managed' costs you If you run Kafka yourself you are responsible for: - **Provisioning** the broker machines (VMs/containers) and storage. - **Metadata layer**: historically a separate **ZooKeeper** ensemble for cluster coordination; modern Kafka replaces this with **KRaft** (Kafka Raft, KIP-500), where brokers/controllers manage metadata internally. Either way you must operate it. - **Patching and upgrades** of the OS and Kafka itself, ideally with no downtime. - **Scaling**: adding brokers and **rebalancing partitions** (e.g. via Cruise Control) so load is even. - **Failure handling**: replacing dead brokers, re-replicating data, watching for under-replicated partitions. - **Security**: TLS, SASL/mTLS auth, ACLs, secrets. - **Monitoring**: lag, disk, ISR (in-sync replicas), throughput. ## What 'managed' abstracts A managed offering is Kafka operated **as a service** by a vendor. The vendor takes over infrastructure and operations: provisioning, patching, node replacement, the metadata layer (you never see ZooKeeper/KRaft), secure defaults, and monitoring. The major options are **Confluent Cloud** (from Kafka's original creators, most fully managed), **Amazon MSK** and **MSK Serverless** (AWS), **Aiven for Apache Kafka** (multi-cloud), and **Azure Event Hubs**, which exposes a **Kafka-protocol surface** (a Kafka-compatible API on top of Event Hubs, not Kafka brokers). ## What stays yours Managed does **not** mean responsibility-free. You still design topics and partition counts, configure clients (acks, retries, idempotence), tune consumer groups and lag, manage **schemas** and governance, and control **cost**, which on most offerings tracks throughput, partition count, storage, and data transfer. ## Trade-offs - **Pros**: dramatically less operational toil, fast onboarding, vendor-backed SLAs, secure defaults. - **Cons**: higher per-unit price, **less control** (some broker configs are locked or unavailable), and **lock-in** (proprietary features, IAM models, connectors). ## Edge cases Degree of abstraction varies: plain MSK still shows you broker instance types and counts (you size the cluster); MSK Serverless and Confluent Cloud hide brokers entirely and bill on usage. Event Hubs is API-compatible but not feature-complete Kafka (e.g. some admin operations and configs behave differently).
- Name one responsibility that stays with you even on a fully managed Kafka.Plenty: topic/partition design, client tuning (acks/idempotence), consumer-group/lag management, schema governance, and controlling cost from throughput/storage. The vendor manages infrastructure, not your data model or application behavior.
- Does 'managed' always mean you never see brokers?No. Plain Amazon MSK still exposes broker instance types and counts that you size. MSK Serverless and Confluent Cloud hide brokers and bill on usage. Degree of abstraction differs per offering.
saying these in an interview costs you the question
- Saying managed Kafka removes all your responsibilities, including partition design and client tuning.
- Claiming Azure Event Hubs is full Apache Kafka rather than a Kafka-compatible protocol surface.
- Assuming every managed offering hides brokers (plain MSK does not).