skip to content

Explain JsonConverter and the schemas.enable setting. What changes in the on-wire payload when it is true vs false?

level: middleimportance: must knowfreq 75%

answer

  1. true => {schema, payload} envelope, self-describing but bulky
  2. false => bare value, schemaless, type loss
  3. default true
  4. per key/value: *.schemas.enable
  5. JsonConverter != JsonSchemaConverter (registry)

basics

~10 s

JsonConverter serializes Connect data as JSON. With schemas.enable=true it wraps the payload as {"schema":...,"payload":...} so the schema travels inline. With false it emits just the raw JSON value with no schema.

solid answer

~40 s

JsonConverter (org.apache.kafka.connect.json.JsonConverter) serializes Connect records to JSON bytes without a schema registry. The key setting is schemas.enable (default true). When true, each message is an envelope: {"schema": {...Connect schema as JSON...}, "payload": {...actual data...}} — the schema is embedded in every record, which is self-describing but bulky and repetitive. When false, the converter emits only the bare value (e.g. {"id":1,"name":"x"}) with no schema, so the consumer/connector gets a schemaless Struct or a Map and loses type fidelity (e.g. it cannot distinguish int32 from int64, or know about decimals/timestamps). You set it independently for key and value: key.converter.schemas.enable and value.converter.schemas.enable. A frequent mistake is a source writing schemas.enable=true envelopes while a downstream consumer expects bare JSON, or a sink connector that needs a schema being fed schemas.enable=false data.

code

json · 11 lines
json
{
  "schema": {
    "type": "struct",
    "fields": [
      { "field": "id", "type": "int32", "optional": false },
      { "field": "name", "type": "string", "optional": true }
    ],
    "optional": false
  },
  "payload": { "id": 1, "name": "x" }
}

go deeper

for a junior

Know JsonConverter outputs JSON and schemas.enable toggles whether a schema is included.

for a middle

Describe the {schema,payload} envelope vs bare value, the default (true), and per-key/value config.

for a senior

Explain type-fidelity loss, size/throughput tradeoffs, and the mismatch failure modes between producers and sinks.

for a principal

Weigh JsonConverter against registry-backed formats org-wide and design a migration that preserves type semantics.

## JsonConverter `org.apache.kafka.connect.json.JsonConverter` converts Connect data to/from JSON-encoded bytes **without** needing a schema registry. It is popular because the output is human-readable and self-contained. ## schemas.enable This boolean (default `true`) controls whether the Connect schema is embedded in each message. ### schemas.enable = true (the envelope) Every message is a two-field JSON object: ``` { "schema": { "type": "struct", "fields": [ {"field":"id","type":"int32"}, ... ] }, "payload": { "id": 1, "name": "x" } } ``` Benefits: the message is **self-describing** — any consumer can reconstruct the exact Connect `Schema` (correct types, optionality, defaults, logical types like Decimal/Timestamp). Downside: the schema is repeated in *every* record, dramatically inflating message size and offering no enforced compatibility checking. ### schemas.enable = false (bare value) The message is just the data: ``` { "id": 1, "name": "x" } ``` No schema travels. On deserialization Connect produces a **schemaless** value (effectively a `Map`/`Struct` with `schema == null`). You lose type precision — JSON numbers are ambiguous, so int32 vs int64 vs float64 cannot be recovered, and logical types (decimal, date, timestamp) are gone. Many sink connectors that require a schema (e.g. JDBC sink building DDL) will fail or behave oddly against schemaless data. ## Key vs value, independently Because converters are set per key/value, so is this property: ``` key.converter=org.apache.kafka.connect.json.JsonConverter key.converter.schemas.enable=false value.converter=org.apache.kafka.connect.json.JsonConverter value.converter.schemas.enable=true ``` ## Common failure modes - **Mismatch:** a producer writes bare JSON but a sink runs with `schemas.enable=true`, so the converter tries to read a `schema`/`payload` envelope and fails (or treats the whole object as the schema). - **Bloat:** `schemas.enable=true` on high-volume topics multiplies storage and network because the schema is repeated per record. For that reason teams often move to Avro/Protobuf with a registry, where the schema is stored once and referenced by ID. - **Type loss:** sinks that depend on type information misbehave under `schemas.enable=false`. ## Relationship to schema registry JsonConverter does **not** use Confluent Schema Registry. (Confluent's `JsonSchemaConverter` is a different class that *does* use the registry and the JSON Schema format — do not confuse the two.)

  • Why might you avoid schemas.enable=true on a high-throughput topic?
    Because the full Connect schema is embedded in every single record, massively inflating message size and network/storage cost. A registry-backed format (Avro/Protobuf) stores the schema once and references it by ID.
  • What type-fidelity problems appear with schemas.enable=false?
    JSON numbers are ambiguous, so int8/16/32/64 and float types cannot be distinguished on read, and logical types like Decimal/Date/Timestamp are lost. Schema-dependent sinks (e.g. JDBC sink) can break.

saying these in an interview costs you the question

  • Saying JsonConverter uses Schema Registry (it does not; JsonSchemaConverter is the registry one)
  • Claiming schemas.enable defaults to false
  • Thinking the envelope is {"key","value"} rather than {"schema","payload"}
  • Asserting bare JSON preserves int32 vs int64 type distinctions

context