Explain JsonConverter and the schemas.enable setting. What changes in the on-wire payload when it is true vs false?
answer
- true => {schema, payload} envelope, self-describing but bulky
- false => bare value, schemaless, type loss
- default true
- per key/value: *.schemas.enable
- JsonConverter != JsonSchemaConverter (registry)
basics
~10 sJsonConverter serializes Connect data as JSON. With schemas.enable=true it wraps the payload as {"schema":...,"payload":...} so the schema travels inline. With false it emits just the raw JSON value with no schema.
solid answer
~40 sJsonConverter (org.apache.kafka.connect.json.JsonConverter) serializes Connect records to JSON bytes without a schema registry. The key setting is schemas.enable (default true). When true, each message is an envelope: {"schema": {...Connect schema as JSON...}, "payload": {...actual data...}} — the schema is embedded in every record, which is self-describing but bulky and repetitive. When false, the converter emits only the bare value (e.g. {"id":1,"name":"x"}) with no schema, so the consumer/connector gets a schemaless Struct or a Map and loses type fidelity (e.g. it cannot distinguish int32 from int64, or know about decimals/timestamps). You set it independently for key and value: key.converter.schemas.enable and value.converter.schemas.enable. A frequent mistake is a source writing schemas.enable=true envelopes while a downstream consumer expects bare JSON, or a sink connector that needs a schema being fed schemas.enable=false data.
code
json · 11 lines{
"schema": {
"type": "struct",
"fields": [
{ "field": "id", "type": "int32", "optional": false },
{ "field": "name", "type": "string", "optional": true }
],
"optional": false
},
"payload": { "id": 1, "name": "x" }
}go deeper
Know JsonConverter outputs JSON and schemas.enable toggles whether a schema is included.
Describe the {schema,payload} envelope vs bare value, the default (true), and per-key/value config.
Explain type-fidelity loss, size/throughput tradeoffs, and the mismatch failure modes between producers and sinks.
Weigh JsonConverter against registry-backed formats org-wide and design a migration that preserves type semantics.
## JsonConverter `org.apache.kafka.connect.json.JsonConverter` converts Connect data to/from JSON-encoded bytes **without** needing a schema registry. It is popular because the output is human-readable and self-contained. ## schemas.enable This boolean (default `true`) controls whether the Connect schema is embedded in each message. ### schemas.enable = true (the envelope) Every message is a two-field JSON object: ``` { "schema": { "type": "struct", "fields": [ {"field":"id","type":"int32"}, ... ] }, "payload": { "id": 1, "name": "x" } } ``` Benefits: the message is **self-describing** — any consumer can reconstruct the exact Connect `Schema` (correct types, optionality, defaults, logical types like Decimal/Timestamp). Downside: the schema is repeated in *every* record, dramatically inflating message size and offering no enforced compatibility checking. ### schemas.enable = false (bare value) The message is just the data: ``` { "id": 1, "name": "x" } ``` No schema travels. On deserialization Connect produces a **schemaless** value (effectively a `Map`/`Struct` with `schema == null`). You lose type precision — JSON numbers are ambiguous, so int32 vs int64 vs float64 cannot be recovered, and logical types (decimal, date, timestamp) are gone. Many sink connectors that require a schema (e.g. JDBC sink building DDL) will fail or behave oddly against schemaless data. ## Key vs value, independently Because converters are set per key/value, so is this property: ``` key.converter=org.apache.kafka.connect.json.JsonConverter key.converter.schemas.enable=false value.converter=org.apache.kafka.connect.json.JsonConverter value.converter.schemas.enable=true ``` ## Common failure modes - **Mismatch:** a producer writes bare JSON but a sink runs with `schemas.enable=true`, so the converter tries to read a `schema`/`payload` envelope and fails (or treats the whole object as the schema). - **Bloat:** `schemas.enable=true` on high-volume topics multiplies storage and network because the schema is repeated per record. For that reason teams often move to Avro/Protobuf with a registry, where the schema is stored once and referenced by ID. - **Type loss:** sinks that depend on type information misbehave under `schemas.enable=false`. ## Relationship to schema registry JsonConverter does **not** use Confluent Schema Registry. (Confluent's `JsonSchemaConverter` is a different class that *does* use the registry and the JSON Schema format — do not confuse the two.)
- Why might you avoid schemas.enable=true on a high-throughput topic?Because the full Connect schema is embedded in every single record, massively inflating message size and network/storage cost. A registry-backed format (Avro/Protobuf) stores the schema once and references it by ID.
- What type-fidelity problems appear with schemas.enable=false?JSON numbers are ambiguous, so int8/16/32/64 and float types cannot be distinguished on read, and logical types like Decimal/Date/Timestamp are lost. Schema-dependent sinks (e.g. JDBC sink) can break.
saying these in an interview costs you the question
- Saying JsonConverter uses Schema Registry (it does not; JsonSchemaConverter is the registry one)
- Claiming schemas.enable defaults to false
- Thinking the envelope is {"key","value"} rather than {"schema","payload"}
- Asserting bare JSON preserves int32 vs int64 type distinctions