skip to content

Which outbox table columns does Debezium's EventRouter SMT read, and how does it pick the topic and key?

level: middleimportance: should knowfreq 52%

answer

  1. five conventional columns, all remappable
  2. one column decides the destination topic
  3. one column decides the partition key
  4. the body travels untouched as the value
  5. extras can ride in message headers

basics

~20 s

Debezium's EventRouter reads an outbox row's id, aggregate type, aggregate id, event type and payload columns. The aggregate type selects the destination topic, the aggregate id becomes the message key, and the payload becomes the message value.

solid answer

~50 s

`io.debezium.transforms.outbox.EventRouter` turns captured inserts on an outbox table into domain events. Its conventional columns are `id`, `aggregatetype`, `aggregateid`, `type` and `payload`, and each is remappable — `table.field.event.id`, `table.field.event.key`, `table.field.event.type`, `table.field.event.payload` — so the SMT fits a table you already have. Routing is driven by `route.by.field`, which defaults to `aggregatetype`: the value of that column is substituted into `route.topic.replacement` to produce the destination topic, so orders and customers land on separate topics from one captured table. The value of the key column, conventionally the aggregate id, becomes the Kafka message key, which is what keeps all events for one aggregate on one partition and in order. The payload column becomes the value. Extra columns can be projected into the value or into headers via `table.fields.additional.placement`, and the row's id travels as a header so consumers can deduplicate.

code

sql · 7 lines
sql
CREATE TABLE outbox (
  id            UUID         PRIMARY KEY,
  aggregatetype VARCHAR(255) NOT NULL,
  aggregateid   VARCHAR(255) NOT NULL,
  type          VARCHAR(255) NOT NULL,
  payload       JSONB        NOT NULL
);

go deeper

for a junior

Know that an outbox table holds one row per domain event with a few well-known columns, and that Debezium captures that table so the service never publishes its business tables directly.

for a middle

Explain which column drives the topic, which becomes the message key, and which becomes the value — and that every one of those roles is remappable through configuration rather than fixed by name.

for a senior

Show why keying by aggregate id is the ordering guarantee, how the propagated event id supports consumer deduplication under at-least-once delivery, and what an insert-then-delete outbox row means for table growth.

for a principal

Treat data-driven routing as a governance surface: a new aggregate type creates a topic at runtime, and the published payload is an organisational contract that nothing in the transform enforces.

## Why the router exists Raw change events publish your table schema. Every consumer then depends on column names you wanted to be free to rename, and on rows that mean nothing outside your service. An outbox table inverts that: the service writes a *designed event* into a table inside the same local transaction as the business change, and the connector captures that table instead of the business tables. The router is the piece that converts a captured outbox row into a message that looks nothing like a row. ## The conventional shape A minimal outbox table has five columns. `id` is a unique event identifier, typically a UUID. `aggregatetype` names the kind of thing the event is about — `Order`, `Customer` — and is the routing dimension. `aggregateid` identifies the specific instance, and becomes the message key. `type` names the event itself — `OrderCreated`, `OrderCancelled`. `payload` holds the event body, usually JSON. None of these names are mandatory. `table.field.event.id`, `table.field.event.key`, `table.field.event.type`, `table.field.event.payload` and `table.field.event.timestamp` each point the router at whichever column plays that role, which matters when you are retrofitting an existing table or when your schema conventions forbid all-lowercase compound names. ## Routing `route.by.field` names the column whose value selects the topic; it defaults to the aggregate type column. Its value is interpolated into `route.topic.replacement`, whose template expands the routed value into a topic name. The practical effect is one captured table fanning out into one topic per aggregate type, which is usually the granularity consumers want to subscribe at: a fulfilment service consumes order events without also receiving every customer-profile change. A subtlety candidates miss: the routed topic must exist or auto-creation must be permitted, and because routing is data-driven, inserting a new aggregate type creates a new topic at runtime. That is a governance question, not just a config one. ## Keys, ordering and deduplication The key column's value becomes the record key. Since a Kafka producer partitions by key, every event for aggregate `order-42` lands on the same partition, and the broker preserves order within a partition — so consumers see that order's events in the sequence the service wrote them, without any global ordering guarantee across aggregates. Choosing the aggregate id as the key is therefore not cosmetic; it is the entire ordering story. Delivery is at-least-once, so consumers must tolerate a repeat. The router propagates the outbox row's identifier as a message header exactly so a consumer can keep a seen-ids set and discard duplicates cheaply, without parsing the payload. ## Extra fields `table.fields.additional.placement` projects further outbox columns into the message. An entry names the column, whether it goes into the value or into a header, and optionally the alias to use — putting the event type into a header, for example, lets a consumer filter or dispatch without deserialising the payload at all. Trace identifiers, tenant ids and a schema version are the columns teams most often add. ## Insert-only semantics and cleanup The router is built for insert events. The service inserts an outbox row and never updates it; many teams delete the row in the very same transaction, because log-based capture reads the transaction log rather than the table, so the insert is captured even though the row never becomes visible to a query — and the table stays empty. On engines where that trick does not apply, a periodic purge job keeps the table small. Either way, updates to outbox rows are not part of the model, and a pipeline that produces them is a design smell to raise rather than a case to configure around. ## Operating notes Capture the outbox table only, not the business tables, or you have published the schema you were trying to hide. Watch that a low-traffic outbox does not leave the capture position stalled behind an idle log. And treat the payload as a contract with its own versioning story: the router happily ships whatever string the service wrote, so nothing in this transform stops a producer from breaking every consumer in one deploy.

  • Why does the aggregate id, rather than the event id, belong in the message key?
    The key decides the partition, and order is only guaranteed within a partition. Keying by aggregate id puts every event for one order on one partition, so a consumer sees created-then-cancelled in the right sequence. Keying by event id would scatter an aggregate's events across partitions and destroy that ordering entirely.
  • How does a consumer deduplicate when the pipeline is at-least-once?
    Use the outbox row's identifier, which the router propagates as a message header. A consumer records processed ids — in its own database, ideally in the same transaction as the effect — and discards repeats. Because the id is generated by the producer at write time, it survives connector restarts and re-deliveries.
  • What happens if a service inserts an outbox row with a brand-new aggregate type?
    Routing is data-driven, so the router computes a topic name that may not exist yet. Depending on cluster policy the topic is auto-created with default partitioning and retention, or the produce fails. Either outcome is a reason to govern aggregate types as part of the event catalogue rather than letting any deploy invent one.

The outbox row is a pre-addressed envelope: one field is the mailbox it goes to, one is the recipient it must stay in order with, and the rest is the letter the postal service never opens.

saying these in an interview costs you the question

  • Says the router reads the business table rather than the outbox table
  • Thinks the topic is fixed in config rather than derived from a column value
  • Uses the event id as the message key, destroying per-aggregate ordering
  • Assumes the router updates or deletes outbox rows itself
  • Claims the outbox guarantees exactly-once delivery to consumers

context