skip to content

Cloud Event-Hubs and Protocol-Edge Choices

Kafka-protocol endpoints on other platforms and AMQP or MQTT edges in front of Kafka, plus where the compatibility ends. Interviewers ask because 'Kafka-compatible' rarely includes transactions or compaction.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is the Azure Event Hubs Kafka endpoint, and how does an existing Kafka application connect to it?

level: juniorimportance: must knowfreq 55%

answer

  1. port 9093, SASL_SSL + PLAIN
  2. username = $ConnectionString
  3. event hub == topic, namespace == cluster
  4. Throughput Units, not brokers
  5. protocol re-impl, not real Kafka

basics

~20 s

Azure Event Hubs exposes a Kafka-compatible endpoint on port 9093, so a normal Kafka client can produce and consume by just changing bootstrap.servers and using SASL/SSL auth with a connection string — no Kafka broker is actually run.

solid answer

~40 s

Event Hubs is Azure's managed streaming service that speaks the Kafka wire protocol on an endpoint like `NAMESPACE.servicebus.windows.net:9093`. An existing Kafka producer/consumer connects by setting `bootstrap.servers` to that host, `security.protocol=SASL_SSL`, `sasl.mechanism=PLAIN`, and a JAAS config whose username is the literal `$ConnectionString` and password is the namespace's connection string (or OAuth/Entra ID token). An Event Hub maps to a Kafka topic; throughput is billed in Throughput Units / Processing Units rather than brokers. No ZooKeeper/KRaft, no broker management. The key caveat: it implements the protocol, not the whole Apache Kafka surface — so some admin, transactional, and compaction features differ or are missing. It is a drop-in for the common produce/consume path, not a full Kafka cluster.

go deeper

for a junior

Know it's a Kafka-compatible endpoint on 9093 you reach by changing bootstrap.servers and auth.

for a middle

Know the exact SASL_SSL/PLAIN/$ConnectionString config and the topic↔event-hub mapping.

for a senior

Articulate that it's a protocol re-implementation with feature gaps and a TU billing model, not a broker fleet.

for a principal

Frame migration trade-offs: zero-ops vs feature subset and Azure-specific scaling/lock-in implications.

## What this is Apache **Kafka** is an open-source distributed log: producers write records to **topics** (split into **partitions**), consumers read them, and brokers store the data. Running Kafka yourself means operating brokers plus a metadata layer (ZooKeeper historically, KRaft now). **Azure Event Hubs** is Microsoft's fully managed event-streaming service. To ease migration, it offers a **Kafka-compatible endpoint**: it re-implements the Kafka *wire protocol* (the binary request/response format Kafka clients speak) so that a standard Kafka client library (e.g. the Java `kafka-clients` jar, librdkafka, etc.) can talk to Event Hubs as if it were a Kafka broker — without Microsoft running any actual Kafka broker code. ## How you connect You point an existing app at it with config only: - `bootstrap.servers=<namespace>.servicebus.windows.net:9093` (port **9093**, TLS-only). - `security.protocol=SASL_SSL` (always encrypted in transit). - `sasl.mechanism=PLAIN` (or `OAUTHBEARER` for Microsoft Entra ID / Azure AD tokens). - JAAS: username is the *literal string* `$ConnectionString`, password is the namespace's connection string (a Shared Access Signature key), OR a bearer token via OAuth. ## Mapping of concepts - A **namespace** ≈ a Kafka cluster boundary. - An **event hub** ≈ a Kafka **topic**. - **Partitions** exist but are fixed at creation (you choose 1–32 typically, up to higher with dedicated) and historically could not be increased on standard tiers. - Capacity is sold as **Throughput Units (TU)** (Standard tier) or **Processing Units (PU)** (Premium) or **Capacity Units (CU)** (Dedicated) — *not* as a number of brokers. 1 TU ≈ 1 MB/s or 1000 events/s ingress, 2 MB/s egress. ## Why it matters / edges Because it is a *protocol re-implementation*, it is a drop-in for the **common produce/consume path** but diverges on advanced features: the full **AdminClient** surface, **Kafka transactions / exactly-once (idempotent producer with `transactional.id`)**, **log compaction** (`cleanup.policy=compact`), and per-topic config knobs are partially or not supported depending on tier and date. So you get fast migration and zero broker ops, at the cost of being limited to the subset Microsoft chose to implement, plus the operational/billing model (TUs) being Azure-specific — a source of lock-in. ## Edge cases - Consumer group offsets are stored, but some clients expect the `__consumer_offsets` internal topic semantics; Event Hubs handles this server-side. - Default minimum supported Kafka protocol versions are gated (very old clients are rejected). - Throttling appears as standard Kafka quota/throttle responses when you exceed your TUs.

  • What username do you put in the SASL/PLAIN JAAS config for the connection-string auth method?
    The literal string `$ConnectionString`; the password is the namespace's Shared Access connection string. Alternatively use `OAUTHBEARER` with a Microsoft Entra ID token.
  • Does connecting via the Kafka endpoint mean Azure runs Apache Kafka brokers for you?
    No. Event Hubs re-implements the Kafka wire protocol on its own service; no broker, ZooKeeper, or KRaft is run. That is why some Kafka features are unsupported.

saying these in an interview costs you the question

  • Saying Event Hubs runs managed Apache Kafka brokers under the hood (it implements the protocol, not Kafka itself).
  • Claiming you must rewrite the app with an Azure SDK — the whole point is config-only migration.
  • Saying capacity is scaled by adding brokers (it's Throughput/Processing Units).

context

open as a page

Where do vendor Kafka-protocol endpoints like Azure Event Hubs diverge from Apache Kafka in transactions, compaction, and admin operations, and how would you detect this before migrating?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Protocol-compatible endpoints implement the produce/consume path but often lack full Kafka transactions/exactly-once, log compaction (cleanup.policy=compact), and parts of the AdminClient. Detect gaps by testing transactional.id, compacted topics, and admin calls against the endpoint before migrating.

open as a page

How do AMQP and MQTT protocol edges fronting Kafka work, and when would you choose them over the native Kafka protocol?

level: middleimportance: should knowfreq 28%

basics

~20 s

An AMQP or MQTT edge is a gateway that accepts those protocols at the boundary and forwards messages into Kafka. Use them when clients can't speak Kafka — e.g. IoT/MQTT devices or AMQP messaging apps — trading native Kafka features for broad client reach.

open as a page

When would you put a Google Pub/Sub-to-Kafka bridge (or the Pub/Sub Kafka shim) between a publisher and a Kafka cluster, and what are the trade-offs?

level: middleimportance: should knowfreq 30%

basics

~20 s

Use a Pub/Sub-to-Kafka bridge to connect Google Pub/Sub and Kafka ecosystems without rewriting apps — e.g. fan messages from Pub/Sub into Kafka topics. Trade-offs: extra hop adds latency, the bridge can break ordering/exactly-once, and it's another component to run.

open as a page

How does the Event Hubs Throughput-Unit / Processing-Unit capacity model differ from sizing an Apache Kafka cluster, and how does that interact with partitions?

level: middleimportance: should knowfreq 40%

basics

~20 s

On Apache Kafka you size brokers, disks, and partitions. On Event Hubs you buy Throughput Units (or Processing/Capacity Units) that cap MB/s and events/s regardless of partition count; partitions are fixed at creation and mainly affect parallelism, not capacity.

open as a page

As an architect, how do you decide between a vendor Kafka-protocol endpoint (e.g. Event Hubs) and self-managed/managed Apache Kafka, and how do you minimize lock-in?

level: principalimportance: should knowfreq 35%

basics

~20 s

Pick a vendor Kafka endpoint when you want zero ops and your workload uses only the supported produce/consume subset; pick real Kafka when you need transactions, compaction, full admin, or portability. Minimize lock-in by isolating vendor specifics, avoiding native-only features, and testing portability.

open as a page