Describe Connect's internal data model — Schema and Struct — and how converters relate to it.
answer
- org.apache.kafka.connect.data: Schema + Struct
- primitives + STRUCT/ARRAY/MAP + logical types
- SchemaBuilder, Struct.put/get
- converter is the only wire-format-aware part
- schema == null => schemaless (Maps)
basics
~10 sConnect represents data internally with a Schema (the type/shape) and a Struct (the values), independent of any wire format. Converters translate between this Schema/Struct model and the raw bytes on the topic.
solid answer
~50 sKafka Connect has a format-agnostic in-memory data model in org.apache.kafka.connect.data. A Schema describes the type and structure: a primitive type (INT8/16/32/64, FLOAT32/64, BOOLEAN, STRING, BYTES) or a complex type (STRUCT, ARRAY, MAP), plus optionality, default values, a name/version, and logical types built on primitives (Decimal, Date, Time, Timestamp). A Struct is a typed record holding values that conform to a STRUCT schema; you access fields via get("field"). Connect records (SourceRecord/SinkRecord) carry keySchema+key and valueSchema+value. Converters are the only place that knows about wire formats: AvroConverter maps Connect Schema to an Avro schema, JsonConverter to a JSON envelope, etc. This decoupling is why the same connector and SMTs work across Avro, JSON, and Protobuf — they operate on Schema/Struct, and only the converter changes. Data can also be schemaless (schema == null), in which case values are plain Java objects/Maps.
code
java · 11 linesSchema schema = SchemaBuilder.struct().name("User")
.field("id", Schema.INT32_SCHEMA)
.field("name", Schema.OPTIONAL_STRING_SCHEMA)
.build();
Struct value = new Struct(schema)
.put("id", 1)
.put("name", "alice");
value.validate();
byte[] bytes = converter.fromConnectData("users", schema, value);go deeper
Know Connect has Schema (shape) and Struct (values) used internally, separate from the wire format.
Enumerate primitive/complex/logical types and explain converters translate Schema/Struct to bytes.
Explain format-neutrality enabling converter swapping, logical-type encoding, and schemaless degradation.
Discuss designing connectors/SMTs around the model and the implications of logical-type fidelity across formats org-wide.
## Why an internal model exists Kafka Connect must work with many wire formats (Avro, JSON, Protobuf, strings, raw bytes) and many connectors. To avoid every connector understanding every format, Connect defines a single **format-neutral data model** in the `org.apache.kafka.connect.data` package. Connectors and transforms speak this model; only converters touch the wire format. ## Schema A `Schema` describes the *type and shape* of a value. Components: - **Primitive types:** `INT8`, `INT16`, `INT32`, `INT64`, `FLOAT32`, `FLOAT64`, `BOOLEAN`, `STRING`, `BYTES`. - **Complex types:** `STRUCT` (named fields), `ARRAY` (element schema), `MAP` (key + value schema). - **Optionality:** a schema can be optional (`optional()`), meaning the value may be null. - **Default value:** used when a field is absent. - **Name & version:** a logical name and integer version, important for evolution and for logical types. - **Logical types:** semantic types layered on primitives via the schema name — `Decimal` (on BYTES), `Date`/`Time`/`Timestamp` (on INT32/INT64). The converter is responsible for encoding these correctly per format. Schemas are built with `SchemaBuilder`, e.g. `SchemaBuilder.struct().field("id", Schema.INT32_SCHEMA).build()`. ## Struct A `Struct` is an instance conforming to a STRUCT `Schema`. You set/get fields by name: `new Struct(schema).put("id", 1)`, and read with `struct.getInt32("id")` or `struct.get("id")`. `Struct.validate()` checks the values against the schema (e.g. non-optional fields present, correct types). ## Records `SourceRecord` and `SinkRecord` each carry a key schema + key value and a value schema + value value, plus topic/partition/offset metadata and headers. SMTs receive and return these records. ## Where converters fit The converter is the **only** component aware of the byte representation: - `fromConnectData(topic, schema, value)` turns a Connect `Schema` + value into `byte[]` (used by source connectors before producing, and conceptually for serialization). - `toConnectData(topic, bytes)` turns `byte[]` back into a `SchemaAndValue` (used by sink connectors after consuming). Because of this separation, you can swap `value.converter` from JSON to Avro without changing the connector or the SMT chain — they still see the same `Schema`/`Struct`. ## Schemaless data Not all data has a schema. With `schema == null`, values are plain Java objects (`Map`, `List`, `String`, numbers). JsonConverter with `schemas.enable=false` produces schemaless data; many SMTs and sinks behave differently (or fail) when the schema is null. ## Why this matters in interviews Understanding that Schema/Struct is wire-format-independent explains: (1) why converters are swappable, (2) why logical types can be lost across formats, and (3) why schemaless mode degrades downstream sinks.
- Name a few Connect logical types and the primitives they build on.Decimal (on BYTES), Date and Time (on INT32), Timestamp (on INT64). They are identified by the schema's name, and the converter encodes them appropriately per wire format.
- Why can you swap value.converter from JSON to Avro without changing the connector?Because the connector and SMTs operate on the format-neutral Schema/Struct model; only the converter knows the byte format, so changing it does not change the typed model the connector sees.
saying these in an interview costs you the question
- Saying connectors work directly with Avro/JSON bytes (they work with Schema/Struct)
- Claiming Connect has no internal type system
- Confusing Connect Schema with an Avro/JSON-Schema document — it is its own model
- Forgetting that schemaless (schema == null) data is valid and common