skip to content

How do you preserve per-aggregate message ordering when publishing outbox events to Kafka, and where can ordering break?

level: seniorimportance: should knowfreq 55%

answer

  1. Order only within a partition
  2. Key = aggregate id -> same partition
  3. enable.idempotence=true preserves order on retry
  4. Monotonic outbox id; CDC = commit order
  5. Per-aggregate order, never global (single partition kills throughput)

basics

~20 s

Kafka only orders messages within a single partition. Use the aggregate id (e.g. order id) as the Kafka message key so all events for one entity land on the same partition in order. Publish them in commit/sequence order from the relay.

solid answer

~50 s

Kafka guarantees ordering only **within a partition**, and the producer routes by key (hash of key -> partition). So to keep events for one aggregate in order, set the Kafka **key = aggregate id** (order id, customer id); every event for that entity hashes to the same partition. The relay must also emit rows in the right order: use a monotonic outbox sequence/id, and with multiple poller threads partition the work by key (don't let two workers interleave the same aggregate). With CDC, commit-log order naturally preserves this. Ordering can still break: (a) `max.in.flight.requests.per.connection > 1` without idempotence can reorder on retry — set `enable.idempotence=true` (which caps in-flight at 5 and preserves order); (b) repartitioning/adding partitions changes key->partition mapping; (c) parallel consumers across partitions see no global order — only per-key order; (d) careless multi-threaded relay claiming. Global total order would require a single partition, which kills throughput, so design for per-aggregate order, not global.

go deeper

for a junior

Know that the message key (aggregate id) decides the partition and that order holds within a partition.

for a middle

Explain keying by aggregate id plus producing in monotonic outbox order; mention idempotent producer.

for a senior

Enumerate where ordering breaks (in-flight retries, multi-thread relay, repartitioning) and the configs that fix each.

for a principal

Trade off per-aggregate vs global ordering against throughput; set partitioning/keying conventions org-wide.

## What Kafka actually guarantees A Kafka **topic** is split into **partitions**. Kafka guarantees message order **only within a single partition** — there is no global ordering across partitions. The producer decides a message's partition by hashing its **key**: `partition = hash(key) % numPartitions` (default partitioner). Messages with the same key go to the same partition; messages with no key are spread round-robin. **Consequence**: to keep all events for one business entity (an *aggregate* — e.g. one order) in order, give them the **same key**. The natural choice is the **aggregate id**. Then `OrderCreated`, `OrderPaid`, `OrderShipped` for order #42 all land on one partition, in the order they were produced, and a single consumer of that partition reads them in order. ## Two ordering responsibilities Ordering in the outbox pattern has two halves: 1. **Produce in order**: the relay must publish outbox rows for a given aggregate in the order they were committed. Store a monotonic column (auto-increment `id` or a sequence) and process ascending. The Debezium **Outbox Event Router** sets the Kafka key from the aggregate-id column for you; with **CDC** the transaction log already reflects commit order. 2. **Route by key**: the message key must be the aggregate id so order is preserved on the partition. ## Where ordering breaks - **Producer retries with parallel in-flight requests**: if `max.in.flight.requests.per.connection > 1` and a batch is retried, a later batch can land before the retried one — **reordering on the same partition**. Fix: `enable.idempotence=true`. The idempotent producer tags records with a producer id + sequence number; the broker rejects out-of-order/duplicate sequences, so order is preserved even with up to 5 in-flight requests, and duplicates from retries are dropped broker-side. - **Multi-threaded polling relay**: if two workers grab events for the *same* aggregate and race to Kafka, they can invert order. Partition the claim by key (e.g. hash the aggregate id to a worker), or keep per-key publishing single-threaded. - **Changing partition count**: adding partitions rehashes keys, so a key that used to go to partition 2 may now go to partition 5 — historical vs new events for that key can sit on different partitions and lose relative order. Plan partition counts up front. - **Consuming across partitions**: even with perfect per-partition order, a consumer group reading many partitions has no cross-partition order. You only get **per-key (per-aggregate) order**, never global order — unless you use a single partition (a throughput bottleneck rarely worth it). ## Rule of thumb Design for **per-aggregate ordering**: key by aggregate id, enable idempotent producer, keep the relay's per-key publishing serial, and fix partition count early. Reserve single-partition 'global order' only for truly low-volume control streams.

  • Why can max.in.flight.requests.per.connection > 1 reorder messages, and how does the idempotent producer fix it?
    With multiple in-flight batches, a retried earlier batch can be appended after a later one, inverting order on the partition. enable.idempotence=true assigns each record a producer-id + monotonic sequence number; the broker enforces sequence order and rejects duplicates/out-of-order writes, so order holds with up to 5 in-flight requests.
  • If you need strict global ordering across all events, what's the cost?
    You must use a single partition, which serializes all traffic through one broker leader and one consumer thread — eliminating Kafka's parallelism. It's almost always better to redesign for per-aggregate ordering keyed by entity id.

saying these in an interview costs you the question

  • Saying Kafka orders messages globally across a topic (it's per-partition only).
  • Keying by event type instead of aggregate id (breaks per-entity order).
  • Ignoring max.in.flight / idempotence and assuming retries can't reorder.
  • Claiming you can add partitions freely without affecting key->partition mapping.
  • Believing more consumer threads preserve order across partitions.

context