Where do vendor Kafka-protocol endpoints like Azure Event Hubs diverge from Apache Kafka in transactions, compaction, and admin operations, and how would you detect this before migrating?
answer
- EOS: transactional.id / read_committed often missing
- cleanup.policy=compact unsupported
- AdminClient partial (alterConfigs no-op, ACLs, partitions)
- conformance harness > trusting docs
- gaps push you to native APIs = lock-in
basics
~20 sProtocol-compatible endpoints implement the produce/consume path but often lack full Kafka transactions/exactly-once, log compaction (cleanup.policy=compact), and parts of the AdminClient. Detect gaps by testing transactional.id, compacted topics, and admin calls against the endpoint before migrating.
solid answer
~40 sEvent Hubs and similar Kafka-API services re-implement a subset of Kafka. The classic divergences: (1) **Transactions / exactly-once** — the idempotent+transactional producer (`enable.idempotence`, `transactional.id`, `initTransactions/commitTransaction`) and `read_committed` consumers are historically unsupported or limited, so EOS pipelines (and Kafka Streams' transactional state) may not work. (2) **Log compaction** — `cleanup.policy=compact` (keep latest value per key) often isn't supported; only time/size retention is. That breaks changelog/KTable and compacted-topic patterns. (3) **AdminClient** — topic create/alter/configs, partition increase, ACLs, and describe operations are partial; many config knobs are read-only or ignored. Detect before migrating by running a conformance pass: attempt `initTransactions`, create a `compact` topic and verify dedup, and exercise `AdminClient.createTopics/alterConfigs/incrementalAlterConfigs/createAcls` against the endpoint, plus run your client's full integration suite. Check the vendor's documented Kafka API support matrix and protocol-version floor.
go deeper
Know that 'compatible' means a subset; some features are missing.
Name the three big gaps: transactions/EOS, compaction, admin.
Design a conformance test that probes transactional.id, compact topics, and AdminClient before migrating.
Tie feature gaps to architectural lock-in and decide go/no-go on EOS-dependent workloads.
## Why divergence exists A vendor that exposes a **Kafka-compatible endpoint** (Azure Event Hubs, and conceptually similar gateways) re-implements the Kafka *wire protocol* on top of its own storage engine. It implements the high-value, high-frequency request types (Produce, Fetch, offset commit, metadata, group coordination) faithfully, but advanced subsystems are expensive to replicate and are often **partially or not implemented**. The three usual gaps: ### 1. Transactions / Exactly-Once Semantics (EOS) Kafka's EOS rests on: the **idempotent producer** (`enable.idempotence=true`, dedup via producer id + sequence numbers) and **transactions** (`transactional.id`, `initTransactions()`, `beginTransaction()`, `sendOffsetsToTransaction()`, `commitTransaction()`), plus consumers reading with `isolation.level=read_committed`. This powers atomic read-process-write and Kafka Streams' fault-tolerant state. On Event Hubs, transactional producers and `read_committed` semantics are historically **not supported** (or narrowly limited), so any EOS pipeline silently degrades or errors. This is the single biggest landmine. ### 2. Log compaction Kafka supports `cleanup.policy=compact`: instead of (or alongside) deleting by time/size, the broker keeps the **latest record per key**, enabling KTable/changelog/state-restore patterns and "topic as a key-value snapshot." Many vendor endpoints support only **retention-based** cleanup (time/size) and reject or ignore `compact`. Pipelines relying on compacted topics (Kafka Streams changelogs, Debezium snapshots, config topics) break or grow unbounded. ### 3. AdminClient surface The Kafka `AdminClient` covers `createTopics`, `deleteTopics`, `createPartitions`, `describeConfigs`, `alterConfigs`/`incrementalAlterConfigs`, `createAcls`, `describeCluster`, consumer-group admin, etc. Vendors implement this **partially**: some topic CRUD works, but altering per-topic configs, increasing partitions, or managing ACLs may be unsupported, no-ops, or routed through the cloud control plane (ARM/portal) instead of the Kafka API. Tooling that auto-provisions topics or tweaks configs at runtime can fail. ## How to detect BEFORE migrating 1. **Read the vendor's Kafka API support matrix** and the **minimum supported protocol/API version** — old clients may be rejected outright. 2. **Run a conformance harness against the endpoint**, not just docs: - Producer: call `initTransactions()` / `commitTransaction()`; confirm it doesn't throw `UnsupportedVersionException`/timeout. - Compaction: create a topic with `cleanup.policy=compact`, write duplicate keys, verify only latest survives. - Admin: exercise `createTopics`, `createPartitions`, `incrementalAlterConfigs`, `createAcls`, `describeConfigs` and assert results. 3. **Run your application's full integration test suite** pointed at the endpoint (especially Kafka Streams apps — they use transactions + compacted changelogs heavily). 4. **Check ecosystem tools**: Connect, Schema Registry expectations, MirrorMaker offset translation. ## The lock-in angle Because the gaps push you toward the vendor's *native* APIs (Event Hubs SDK, AMQP, capture-to-blob, ARM-based admin) to fill them, designs tend to accrete Azure-specific dependencies — that is the real lock-in, more than the protocol itself. ## Edge cases - A transactional producer may *initialize* but fail at commit, so test the full cycle. - `alterConfigs` may return success yet silently ignore the change. - Partition increase may be a control-plane-only operation, invisible to `AdminClient`.
- A team relies on Kafka Streams with stateful aggregations. Why is moving to a transactions-less endpoint risky?Kafka Streams uses exactly-once (transactional producer + read_committed) and compacted changelog topics for fault-tolerant state. Without transactions and compaction, EOS and reliable state restoration break, risking duplicates and lost/unbounded state.
- Why isn't reading the vendor docs enough to verify compatibility?Support matrices lag and some ops succeed-but-no-op (e.g. alterConfigs). A conformance harness exercising transactions, compaction, and admin against the live endpoint catches silent divergence.
saying these in an interview costs you the question
- Assuming a 'Kafka-compatible' endpoint supports the full Kafka feature set including EOS and compaction.
- Believing AdminClient calls that return success actually applied (some are silent no-ops).
- Migrating a Kafka Streams EOS app without first testing transactions and compacted changelogs.