What changed between the original exactly_once and exactly_once_v2 (KIP-447), and why was v2 introduced?
answer
- v1: producer per task/partition; v2: producer per thread
- v1 fences by per-partition transactional.id
- v2 fences via ConsumerGroupMetadata (KIP-447)
- needs brokers >= 2.5
- was exactly_once_beta in 2.5
basics
~20 sThe original exactly_once used one transactional producer per task (input partition), which didn't scale — many producers and slow rebalances. exactly_once_v2 (KIP-447) uses one producer per stream thread and fences zombies via consumer-group metadata, so it scales far better. v2 needs brokers >= 2.5.
solid answer
~40 sOriginal `exactly_once` (KIP-98/129) created one transactional producer per task, i.e. per input partition. With hundreds of partitions you get hundreds of producers, hundreds of broker connections, lots of transactional.ids, and expensive rebalances because each producer must re-init transactions and the broker fences by transactional.id. KIP-447 (`exactly_once_v2`, originally exposed as `exactly_once_beta` in 2.5) reduces this to **one transactional producer per StreamThread (per instance)**. Zombie fencing no longer relies on a unique transactional.id per partition; instead `sendOffsetsToTransaction` carries `ConsumerGroupMetadata`, letting the broker fence stale producers using the consumer group's generation/epoch. Benefits: far fewer producers/connections, much faster and cheaper rebalances, better throughput. Requirement: brokers (and the transaction coordinator) must be >= 2.5. The original `exactly_once` is deprecated; new apps should use `exactly_once_v2`.
go deeper
Know v2 is the modern recommended value and v1 (exactly_once) is deprecated.
Explain producer-per-task vs producer-per-thread and that v2 scales better with fewer producers.
Articulate the per-partition transactional.id fencing problem and how ConsumerGroupMetadata replaces it (KIP-447), plus the 2.5 broker requirement.
Discuss migration procedure, blast-radius tradeoff of per-thread producers, and coordinator-load implications at scale.
## Background: how the original EOS fenced zombies Kafka transactions identify a logical producer by a stable **transactional.id**. The broker's transaction coordinator tracks a **producer epoch** per transactional.id; when a new producer with the same id calls `initTransactions()`, the epoch bumps and any older producer (a ‘zombie’) is **fenced** (its sends/commits fail with `ProducerFencedException`). To make this fencing correct across rebalances, the original Streams EOS gave **each task a unique transactional.id** derived from the application id + task id (partition). Reason: a partition can move to another instance during a rebalance, and the only way to guarantee the old owner can't still commit was to tie the transactional.id to the partition, so whoever owns the partition next bumps the epoch and fences the previous owner. ## Why that didn't scale One transactional.id per partition means **one transactional producer per partition** (a producer can run only one transactional.id). Consequences: - N partitions → N producers per instance → N× broker connections, memory, threads. - Each rebalance reassigns partitions, so producers must be **closed and recreated** and re-run `initTransactions()` — slow, and it touches the transaction coordinator a lot. - Transaction state on the coordinator grows with the number of transactional.ids. This made EOS painful at high partition counts and made rebalances a latency cliff. ## What KIP-447 changed KIP-447 decouples zombie fencing from per-partition transactional.ids. The insight: the **consumer group** already coordinates partition ownership and has a generation/epoch. By passing **`ConsumerGroupMetadata`** (group id, generation id, member id, instance id) into `producer.sendOffsetsToTransaction(...)`, the transaction coordinator can ask the **group coordinator** whether this producer's consumer is still the current owner. A producer whose consumer has been rebalanced out (stale generation) is fenced when it tries to commit offsets — without needing a partition-scoped transactional.id. Result: Streams uses **one transactional.id per StreamThread** (thread-level), so **one producer per thread**, regardless of how many partitions/tasks that thread handles. Far fewer producers, far cheaper rebalances. ## Naming / version timeline - 2.5: shipped as `exactly_once_beta`. - 2.6: renamed `exactly_once_v2`; `exactly_once_beta` deprecated as an alias. - 3.0: original `exactly_once` deprecated; v2 is the recommended value. - Requirement: brokers and the transaction/group coordinator must be **>= 2.5** because the broker side must support the metadata-based fencing API. ## Migration notes - Upgrade brokers to >= 2.5 first. - Because the transactional.id scheme changes, an in-place rolling upgrade from `exactly_once` to `exactly_once_v2` must follow the documented two-rolling-bounce procedure to avoid mixing fencing models; otherwise zombie fencing could be momentarily incorrect. ## Edge cases - v2 still requires single-cluster transactions and durable output config. - Throughput improves but EOS still adds commit-marker and read_committed latency. - Per-thread producer means a single producer error can affect all tasks on that thread (broader blast radius than per-task producers) — a deliberate tradeoff for scalability.
- Why did the original EOS need a transactional.id per partition?Because a partition can move to another instance on rebalance; tying the transactional.id to the partition guarantees the new owner bumps the epoch and fences the previous owner from committing.
- What broker-side requirement does v2 impose, and why?Brokers and coordinators must be >= 2.5 because v2's zombie fencing uses consumer-group-metadata validation between the transaction and group coordinators, an API added in 2.5.
saying these in an interview costs you the question
- Claiming v2 just renames v1 with no behavioral change.
- Saying v2 uses a producer per partition (that's v1).
- Forgetting the broker >= 2.5 requirement.
- Asserting you can hot-swap exactly_once to exactly_once_v2 without the documented rolling-bounce migration.