What is a Single Message Transform (SMT) in Kafka Connect, and how do you configure a chain of them on a connector?
answer
- transforms=<alias list>, runs left-to-right
- transforms.<alias>.type = FQ class
- per-record, stateless
- $Key vs $Value inner classes
- implements Transformation<R>
basics
~20 sAn SMT is a small function that modifies each record as it flows through a Kafka Connect connector. You list transforms by alias in the transforms config, then configure each alias with transforms.<alias>.type and its properties. They run in the order listed.
solid answer
~40 sA Single Message Transform (SMT) is a lightweight, stateless function applied to each individual record inside a Kafka Connect connector pipeline. For a source connector SMTs run after the connector produces records but before they hit Kafka; for a sink connector they run after reading from Kafka but before delivery to the sink. You declare a chain with the `transforms` property, listing comma-separated aliases (e.g. `transforms=route,addTs`). Each alias is then configured with `transforms.route.type=org.apache.kafka.connect.transforms.RegexRouter` plus that SMT's own properties like `transforms.route.regex` and `transforms.route.replacement`. The chain executes left-to-right in the listed order, each SMT receiving the previous one's output. SMTs implement the `org.apache.kafka.connect.transforms.Transformation` interface and are meant for simple per-record edits — not joins, aggregations, or anything needing state across records.
go deeper
Know SMTs edit one record at a time and that transforms= lists aliases configured by transforms.<alias>.type.
Explain source-vs-sink ordering relative to the converter, the $Key/$Value variants, and chain execution order.
Discuss schema propagation, statelessness limits, and when to reach for Kafka Streams instead.
Frame SMTs as a thin, deterministic hot-path concern; set org guidance on what belongs in SMTs vs stream processing vs upstream connectors.
## What problem SMTs solve Kafka Connect moves data between Kafka and external systems using **source connectors** (external system → Kafka) and **sink connectors** (Kafka → external system). Often the record needs a small adjustment in transit: rename a field, drop a field, change the destination topic, add a timestamp. Rewriting the connector for each tweak would be wasteful. A **Single Message Transform (SMT)** is a reusable, configuration-driven function that operates on **one record at a time** to make exactly these small edits. ## Where in the pipeline they run - **Source connector:** connector task produces a `SourceRecord` → SMT chain → converter serializes → record written to Kafka. - **Sink connector:** record read from Kafka → converter deserializes → SMT chain → `SinkRecord` handed to the sink task. So SMTs always sit between the connector and the converter/Kafka boundary. ## Configuring a chain Three pieces: 1. `transforms` — a comma-separated list of **aliases** you invent, e.g. `transforms=insertTs,route`. 2. `transforms.<alias>.type` — the fully-qualified class name of the SMT implementation. 3. `transforms.<alias>.<property>` — that SMT's own settings. Example: ``` transforms=insertTs,route transforms.insertTs.type=org.apache.kafka.connect.transforms.InsertField$Value transforms.insertTs.timestamp.field=ingestedAt transforms.route.type=org.apache.kafka.connect.transforms.RegexRouter transforms.route.regex=(.*) transforms.route.replacement=prod_$1 ``` The chain runs **in the order the aliases appear** in `transforms`. Here every record first gets an `ingestedAt` field inserted, then its topic is renamed with a `prod_` prefix. ## Key properties of SMTs - **Per-record and stateless:** they see one record and cannot aggregate or join across records. - **Key vs Value variants:** many SMTs ship as two inner classes, `...$Key` and `...$Value`, choosing whether to operate on the record key or value. - **Schema-aware:** SMTs work on both schemaful (e.g. Avro `Struct`) and schemaless (plain `Map`) records, transforming the schema as well as the data when present. - **Implement `Transformation<R>`:** the contract has `apply(R record)`, `configure(Map)`, `config()`, and `close()`. ## What SMTs are NOT for Anything requiring external lookups, buffering, windowing, or cross-record state. For that, use Kafka Streams or ksqlDB. SMTs are intentionally simple to keep them cheap and predictable on the hot path.
- Do SMTs run before or after the converter on a source connector?Before. On a source connector the SMT chain transforms the SourceRecord, then the converter serializes it to bytes for Kafka. On a sink connector the order is reversed: converter deserializes first, then the SMT chain runs.
- Why can't you use an SMT to join two records or compute a running total?SMTs are stateless and operate on a single record at a time — the apply() method only sees the current record. Cross-record operations need a stream processor like Kafka Streams or ksqlDB.
saying these in an interview costs you the question
- Saying SMTs can aggregate or join records — they are strictly per-record and stateless.
- Thinking the chain order doesn't matter — it executes strictly in the order aliases are listed.
- Confusing where SMTs run relative to the converter (it differs between source and sink).