Describe the structure of a Debezium change event envelope (op/before/after/source) and explain how deletes and tombstones work.
answer
- op: c / u / d / r
- before + after + source
- delete = op:d then tombstone (null value)
- tombstone enables log compaction to remove key
- tombstones.on.delete default true
basics
~20 sEach Debezium event has a value with op (c/u/d/r), before (old row), after (new row), and source metadata. A delete sends an event with op='d', before populated, after null — followed by a separate tombstone message (same key, null value) so log-compacted topics can drop the key.
solid answer
~40 sDebezium wraps every change in an *envelope* whose value struct has four key fields: `op` — the operation: `c` (create/insert), `u` (update), `d` (delete), `r` (read, emitted during snapshots); `before` — the row state prior to the change (null for inserts); `after` — the row state after (null for deletes); and `source` — metadata (connector, db, table, binlog file/position or LSN, transaction id, timestamp). The Kafka message *key* is the table's primary key. For a delete, Debezium emits a `d` event (before populated, after null), then immediately a second record with the same key and a **null value** — the *tombstone*. The tombstone exists so that on a **log-compacted** topic Kafka's compaction will eventually delete the key entirely, preventing the deleted row from lingering forever. Tombstone emission is controlled by `tombstones.on.delete` (default true).
go deeper
Recognize the four op codes and that before/after carry old and new row state.
Explain the full envelope and the two-message delete (event + tombstone) and why the tombstone exists.
Tie tombstones to log compaction semantics and explain REPLICA IDENTITY / binlog_row_image effects on 'before'.
Design downstream contracts around the envelope, decide on ExtractNewRecordState flattening, and reason about compaction-based materialized state.
**The envelope.** Debezium does not emit raw rows; it wraps each change in a standard structure called the *envelope* so consumers always get the same shape. The Kafka record has two parts: - **Key**: a struct of the source table's **primary key** (or a configured key). This determines partitioning and, crucially, log-compaction identity. - **Value**: the envelope struct, with these fields: - **`op`** — a single character naming the operation: - `c` = create (INSERT) - `u` = update (UPDATE) - `d` = delete (DELETE) - `r` = read — used only for rows emitted during the **initial snapshot** (they are not real-time changes, just the current state being read out) - **`before`** — the full row *before* the change. `null` for inserts (nothing existed before). Whether `before` is fully populated on updates/deletes depends on DB config: MySQL needs `binlog_row_image=FULL`; Postgres needs `REPLICA IDENTITY FULL` on the table, otherwise only key columns appear. - **`after`** — the full row *after* the change. `null` for deletes (nothing remains). - **`source`** — provenance metadata: connector name, database, schema, table, server id, binlog filename + position (MySQL) or LSN (Postgres), transaction id, source timestamp `ts_ms`, and whether it came from a snapshot. - **`ts_ms`** — when Debezium processed the event; plus `transaction` block when transaction metadata is enabled. **Deletes — two messages.** A single DELETE in the source produces **two** Kafka records: 1. A **delete event**: same key, value envelope with `op:'d'`, `before` = the deleted row, `after` = null. 2. A **tombstone**: same key, **value = null** (the entire value, not just `after`). **Why the tombstone?** Kafka **log compaction** keeps only the *latest* value per key and, when it sees a key whose latest value is `null`, it treats that as a deletion marker and eventually removes the key from the topic entirely (after `delete.retention.ms`). On a compacted topic acting as a 'current state' table, the tombstone is what lets the deleted key actually disappear instead of persisting forever as a `d` event. The first (non-null) `d` event carries the *information* about the delete; the tombstone carries the *compaction semantics*. **Config knob:** `tombstones.on.delete` (default `true`). Set it `false` if your downstream cannot handle null-value records (some sinks choke on them) or you are not using compaction — but then deleted keys won't be compacted away. **Common SMT:** The `ExtractNewRecordState` single-message transform (a.k.a. the 'new record state extraction' / unwrap SMT) flattens the envelope to just the `after` row for consumers that want a plain row, and can be configured to re-add chosen metadata and to handle deletes via `delete.handling.mode` (rewrite, drop, or none).
- What does op='r' mean and when is it produced?'r' = read; it tags rows emitted during the initial snapshot. They represent existing state read out of the table, not a live change captured from the log.
- Why might the 'before' field be missing on a Postgres update or delete?Postgres only logs old row values fully when the table's REPLICA IDENTITY is FULL. With the default (DEFAULT/primary key), 'before' contains only the key columns, not all old values.
- What breaks if you set tombstones.on.delete=false on a compacted topic?Compaction never removes deleted keys, so the 'd' event remains as the last value for that key indefinitely, and consumers reading the compacted topic as current state will still see the (deleted) key.
saying these in an interview costs you the question
- Saying a delete produces only one message (it produces a 'd' event plus a tombstone).
- Confusing the tombstone (whole value null) with the 'd' event (after is null but value is a struct).
- Claiming 'before' is always fully populated regardless of REPLICA IDENTITY / binlog_row_image.
- Thinking op='r' means a read query against the DB rather than a snapshot row.