skip to content

Serializers, Deserializers and Built-in Serdes

Serializer and deserializer pairs, the built-in serdes, and the fact that the broker itself only ever sees byte arrays. Interviewers start here before asking anything about schemas.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is a Serializer in Kafka, and how do producers know which one to use for keys and values?

level: juniorimportance: must knowfreq 78%

answer

  1. object -> byte[]
  2. key.serializer / value.serializer
  3. broker stores opaque bytes
  4. Serializer<T>.serialize(topic, data)
  5. key and value independent

basics

~20 s

A Serializer turns a Java/Kotlin object into a byte[] so Kafka can store and send it. Producers pick one via the key.serializer and value.serializer config properties — one for the key, one for the value.

solid answer

~40 s

Kafka brokers only store and transfer raw bytes, so a producer must convert each record's key and value into a byte[] before sending. That conversion is done by a Serializer — a class implementing org.apache.kafka.common.serialization.Serializer<T>, whose core method is byte[] serialize(String topic, T data). You tell the producer which classes to use with the mandatory configs key.serializer and value.serializer (each set to a fully-qualified class name, e.g. org.apache.kafka.common.serialization.StringSerializer). The producer instantiates them via reflection and calls serialize() on every record. Keys and values are independent, so you commonly mix serializers — e.g. a String key with a JSON or Avro value. If a serializer is missing or wrong, you get a ConfigException at startup or a SerializationException at send time.

go deeper

for a junior

Know that producers turn objects into bytes via a Serializer, set with key.serializer/value.serializer.

for a middle

Know the Serializer<T> interface signature and that key/value are independent, plus the runtime errors.

for a senior

Discuss reflection-based instantiation, generic-vs-config mismatch failures, and serialization happening before partitioning.

for a principal

Frame serialization as the contract boundary between app types and the broker's byte model, and its implications for schema evolution and partitioning.

## The problem serializers solve A Kafka broker is intentionally dumb about your data: it stores each message as an opaque sequence of bytes (a `byte[]`) and never interprets it. But your application code works with typed objects — a `String` username, a `Long` id, an `Order` POJO. Something has to translate between the two worlds. On the way out (producer), that translator is a **Serializer**; on the way in (consumer), it's a **Deserializer**. ## The Serializer interface It lives in `org.apache.kafka.common.serialization.Serializer<T>`. The method that matters is: ``` byte[] serialize(String topic, T data) ``` Given the topic name and a typed object, it returns the raw bytes. (There is also an overload that receives the record `Headers`, plus optional `configure(...)` and `close()` lifecycle methods.) The broker stores exactly those bytes. ## How the producer chooses one — key.serializer / value.serializer A `KafkaProducer` is configured with two **mandatory** properties: - `key.serializer` — class used to serialize the record key - `value.serializer` — class used to serialize the record value Each is the **fully-qualified class name** of a `Serializer` implementation, e.g. `org.apache.kafka.common.serialization.StringSerializer`. The producer loads the class by reflection at construction time. These are **separate** settings because a record's key and value are independent and frequently have different types — a `String` key with a JSON value is common. ## Why typed generics matter `KafkaProducer<K, V>` is generic. If you declare `KafkaProducer<String, byte[]>` but configure a `LongSerializer`, the mismatch is only caught at runtime when `serialize()` casts and fails, throwing a `SerializationException`. The generic type and the configured serializer class must agree. ## What happens at send time For each `ProducerRecord`, the producer calls the key serializer on the key and the value serializer on the value, then appends the resulting byte arrays to a batch destined for a partition. Serialization happens **before** partitioning (the default partitioner can hash the serialized key bytes). ## Edge cases - Missing `key.serializer`/`value.serializer` → `ConfigException` when the producer is created. - A `null` key or value is passed through as `null` bytes, not run through the serializer for the actual content in most built-in serializers (they return `null` for `null` input). - Serialization is synchronous and on the calling thread; an expensive serializer adds latency to `send()`.

  • Why are there two separate config keys instead of one?
    Because a record's key and value are independent and often differ in type — e.g. a String key for partitioning with a JSON or Avro value as the payload, so each needs its own serializer.
  • What error do you get if you forget to set value.serializer?
    A ConfigException is thrown when the KafkaProducer is constructed, because key.serializer and value.serializer are mandatory configs with no default.

saying these in an interview costs you the question

  • Saying the broker deserializes or validates message content — it only stores opaque bytes.
  • Claiming one serializer config covers both key and value.
  • Confusing Serializer (producer side) with Deserializer (consumer side).

context

open as a page

Which serializers ship built-in with the Kafka clients library, and what wire format do they produce?

level: middleimportance: must knowfreq 70%

basics

~10 s

Kafka bundles StringSerializer, ByteArraySerializer, IntegerSerializer, LongSerializer, DoubleSerializer, ShortSerializer, FloatSerializer, and UUIDSerializer (plus ByteBuffer/Bytes/Void). Numbers use fixed-width big-endian bytes; String/UUID use a configurable charset (UTF-8 default).

open as a page

What happens when a serializer is given a null value, and what is a tombstone?

level: middleimportance: must knowfreq 58%

basics

~20 s

Built-in serializers return null bytes for null input — they don't crash. A record with a non-null key but a null value is a tombstone: on a log-compacted topic it signals 'delete this key,' and compaction eventually removes it.

open as a page

When would you use ByteArraySerializer versus StringSerializer or a typed serializer, and what are the trade-offs?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Use ByteArraySerializer when your data is already bytes or you serialize it yourself (e.g. Avro/Protobuf done manually). Use StringSerializer for text. Use typed serializers (Long/Integer/UUID) when keys/values are those primitives. ByteArray is most flexible but least type-safe.

open as a page

How would you implement a custom Serializer/Deserializer, and what production concerns must it handle?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Implement Serializer<T>/Deserializer<T>, putting your encoding in serialize()/deserialize(). Handle null in and out, make it thread-safe and stateless, honor the isKey flag in configure(), version your format for evolution, and fail with SerializationException — never crash the whole consumer.

open as a page

In Kafka Streams, what is a Serde and how does the Serdes factory class fit in?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A Serde<T> bundles a Serializer<T> and a Deserializer<T> into one object, because Streams both reads and writes data. The Serdes factory class provides ready-made ones, e.g. Serdes.String(), Serdes.Long(), and you set defaults with default.key.serde / default.value.serde.

open as a page