What does Debezium's ExtractNewRecordState SMT do to a change event, and how does it treat deletes?
answer
- the envelope collapses to one row image
- before, op and source stop travelling
- a delete has nothing to unwrap
- default behaviour makes deletions vanish
- re-attach metadata with add.fields
basics
~20 sDebezium's ExtractNewRecordState SMT flattens the change-event envelope down to the after row image, so a sink sees an ordinary record instead of before/after/op. Deletes arrive with a null after image and are dropped by default.
solid answer
~50 sDebezium emits a nested envelope per change: `before`, `after`, `op`, `source`, `ts_ms`. Most sinks — a JDBC sink writing to a table, a search indexer — want a flat row, not that envelope. `io.debezium.transforms.ExtractNewRecordState` (historically nicknamed the *unwrap* SMT) replaces the value with the `after` struct, so the record looks like a plain row image. Deletes are the trap. A delete has `after: null`, so there is nothing to unwrap. By default the SMT drops the delete event and drops the tombstone that follows it, which means deleted rows silently survive in the sink. You choose the behaviour explicitly: forward the delete as a null value so a delete-aware sink acts on it, or rewrite it into a row carrying a `__deleted` marker. `add.fields` and `add.headers` put the metadata you still need — `op`, `source.ts_ms`, the LSN — back onto the flattened record.
code
properties · 5 linestransforms=unwrap
transforms.unwrap.type=io.debezium.transforms.ExtractNewRecordState
transforms.unwrap.drop.tombstones=false
transforms.unwrap.delete.handling.mode=rewrite
transforms.unwrap.add.fields=op,source.ts_msgo deeper
Be able to say that a Debezium change event is nested — before, after, op, source — and that this transform hands the sink just the current row image instead of the whole envelope.
Explain where the SMT runs, what happens to the value schema after flattening, and why a delete event and its tombstone need an explicit handling decision rather than the default.
Diagnose the live symptom: deleted source rows still present in the target. Show that you would choose the delete mode deliberately, re-add op and source timestamps, and treat the shape change as a consumer-visible contract change.
Own the decision of whether a topic publishes row images at all. Weigh flattening against keeping the envelope or moving to designed outbox events, given that the choice is topic-wide and expensive to reverse once consumers exist.
## The envelope you start with Every Debezium change event carries a nested value: `before` (the row as it was), `after` (the row as it now is), `op` (`c` create, `u` update, `d` delete, `r` read-during-snapshot), a `source` block with connector, database, table, timestamp and log position, and `ts_ms`. That envelope is exactly what you want if you are building a stream processor that needs to see the transition. It is exactly what you do *not* want if the consumer is a sink connector that expects one field per column, or a downstream table that mirrors the source. ## What the SMT does `io.debezium.transforms.ExtractNewRecordState` is a single message transform that replaces the record value with the contents of `after`. A create or update event becomes a flat struct whose fields are the table's columns; the schema handed to the converter is the row schema rather than the envelope schema. Snapshot reads (`op: r`) flatten the same way, so a sink cannot tell a snapshot row from a streamed insert unless you deliberately keep the `op` field. Because flattening throws the envelope away, anything you still need has to be re-attached. `add.fields` takes a comma-separated list — `op`, `table`, `lsn`, `source.ts_ms` — and adds each as a top-level field prefixed with `__` by default. `add.headers` does the same thing into Kafka message headers instead, which keeps the value clean when the value schema is under contract. ## Deletes, tombstones, and the classic bug A delete produces two records: the change event with `before` populated and `after` null, followed by a genuine tombstone — same key, null value — so that log compaction can eventually reclaim the key. The SMT cannot unwrap a null `after`, so it has to be told what to do. With the default settings it drops both the delete event and the tombstone. The result is the single most common report from teams new to this transform: *rows deleted in the source stay forever in the sink table*. Nothing errors; the events simply never arrive. The alternatives are (a) let the delete through as a record with a null value, which a delete-aware sink such as the JDBC sink with deletes enabled will translate into a `DELETE` statement, or (b) rewrite the delete into a normal-looking row built from the `before` image plus a `__deleted` flag, which is what you want when the target is an append-only or soft-delete model and you would rather mark than remove. In Debezium 2.x these are governed by `drop.tombstones` and `delete.handling.mode` (`drop` | `rewrite` | `none`); later 2.x releases consolidated them into a single delete/tombstone handling property and deprecated the older pair, so check the property names for the version you actually run. ## Where the transform runs, and what that implies The SMT executes inside the connector's task, per record, before the converter serialises it. That has two consequences worth stating in an interview. First, it changes the *schema* the converter registers — after flattening, the subject for that topic holds the row schema, not the envelope schema, so turning the SMT on or off mid-life is a breaking change for every consumer of that topic. Second, per-record transforms cannot join, aggregate or reorder; if you need the `before` image to compute a diff, flattening has already destroyed the information by the time a downstream job sees it. ## When not to flatten Keep the envelope when consumers need the transition rather than the state: audit trails, change-log tables, CDC-driven SCD Type 2 loads that compare old and new values, or any consumer that must distinguish an update from an insert on business grounds. Keep it too when different consumers of the same topic disagree — flattening is a topic-wide decision, and the safer pattern is to publish the envelope and let each consumer project it, or to move to an outbox table so the published shape is a designed event rather than a row image at all. ## Operational checklist Decide the delete behaviour explicitly rather than inheriting the default. Re-add `op` and a source timestamp, because a sink that cannot tell a snapshot row from a live insert cannot order or debug anything. Verify the flattened schema against your registry's compatibility rule before deploying, since the change of shape is not backwards compatible with the envelope. And test a delete end to end in staging: it is the path that silently does nothing.
- After flattening, how would a consumer still tell a snapshot row apart from a live insert?Only if you re-add the operation field. Snapshot reads carry `op: r` while streamed inserts carry `op: c`; once the envelope is gone that distinction is gone with it. Adding `op` through `add.fields` or `add.headers` preserves it, and adding the source timestamp lets a consumer reason about how far behind the row is.
- Why is enabling this transform on an existing topic a breaking change?The value schema changes from the envelope schema to the row schema. Every consumer deserialising the envelope fails, and with a schema registry the new schema is a different, incompatible shape under the same subject. Treat it as a new topic or a coordinated cutover, not a config tweak.
- Your sink needs both the old and new value of a column. Can this SMT give you that?No. It keeps only the `after` image, so the `before` values are discarded inside the connector before anything downstream sees them. Leave the envelope intact and compute the diff downstream, or publish a purpose-built event that carries both values explicitly.
saying these in an interview costs you the question
- Claims the SMT keeps the before image alongside the after image
- Assumes deletes propagate to the sink with default settings
- Thinks the transform runs on the consumer rather than in the connector task
- Says turning flattening on later is a transparent, non-breaking change
- Confuses the flattened null-value record with the outbox router's output