What are serializers and deserializers in the Kafka Java client, and what happens when a deserializer hits a bad (poison) record?
answer
- Kafka = bytes only
- Serializer obj->byte[], Deserializer byte[]->obj
- deser runs inside poll()
- poison pill = bad record stalls partition forever
- fix: ErrorHandlingDeserializer / DLT / seek past
basics
~20 sKafka moves only bytes. A Serializer<T> turns your key/value object into byte[] before producing; a Deserializer<T> turns byte[] back into an object when consuming. A bad record that the deserializer can't parse throws inside poll(), and by default the consumer keeps failing on that same offset — a 'poison pill' that blocks progress until you handle it.
solid answer
~40 sSerializers (Serializer<T>) and deserializers (Deserializer<T>) bridge your domain types and Kafka's byte[] wire format. The producer applies key.serializer/value.serializer; the consumer applies key.deserializer/value.deserializer. Built-ins exist for String, byte[], Integer, Long, etc., plus JSON/Avro/Protobuf serdes (often with Schema Registry). The poison-pill problem: deserialization runs inside poll(), so a record that throws (e.g. SerializationException — JSON that won't parse) makes poll() itself fail. Since you can't commit past a record you never decoded, the consumer retries the same offset forever and the partition stalls. Mitigations: use ErrorHandlingDeserializer (Spring Kafka) or a delegating deserializer that catches the exception, returns null/sentinel, and routes the raw bytes to a dead-letter topic; or set deserializer leniency. The key insight is that deserialization errors are a consume-side liveness hazard, not just a parsing detail.
code
properties · 4 lines# Spring Kafka: wrap real deserializers so poison pills don't stall the partition
spring.kafka.consumer.value-deserializer=org.springframework.kafka.support.serializer.ErrorHandlingDeserializer
spring.kafka.consumer.properties.spring.deserializer.value.delegate.class=org.apache.kafka.common.serialization.StringDeserializer
# pair with a DefaultErrorHandler + DeadLetterPublishingRecoverer to route bad records to a DLTgo deeper
Know serializers turn objects into bytes for producing and deserializers turn bytes back into objects for consuming.
Explain that deserialization runs in poll() and a bad record can stall the partition (poison pill).
Describe DLT / ErrorHandlingDeserializer / seek-past strategies and Schema Registry for prevention.
Set org policy for malformed-data handling (skip vs DLT vs halt), schema governance, and idempotent reprocessing of dead-lettered records.
## Why serdes exist Kafka brokers store and transfer only `byte[]` — they're agnostic to your data's meaning. The Java client therefore needs to convert between your application types and bytes: - **Producer**: `Serializer<T>.serialize(topic, T) -> byte[]`, configured via `key.serializer` and `value.serializer`. - **Consumer**: `Deserializer<T>.deserialize(topic, byte[]) -> T`, configured via `key.deserializer` and `value.deserializer`. The client ships built-ins in `org.apache.kafka.common.serialization`: `StringSerializer`/`StringDeserializer`, `ByteArray...`, `Integer...`, `Long...`, `Double...`, `UUID...`, etc. For structured data you use JSON, Avro, or Protobuf serdes, frequently paired with a **Schema Registry** for schema evolution. A `Serde<T>` (used by Kafka Streams) just bundles a matching serializer + deserializer pair. ## Generics must match The `<K,V>` on the producer/consumer must agree with the configured serdes; a mismatch is a `ClassCastException` at send/poll time, not at construction. ## The poison-pill problem Deserialization happens **inside `poll()`**, before records are handed to your code. If a record's bytes can't be deserialized — malformed JSON, wrong schema, a record produced with a different serializer — the deserializer throws (typically `org.apache.kafka.common.errors.SerializationException`). That exception propagates out of `poll()`. Why it's dangerous: the consumer's committed offset is still *before* that record. You never successfully read it, so you can't advance past it. The naive loop catches the error, loops, calls `poll()` again, gets the **same** record, and throws again — forever. One bad message ('poison pill') **stalls the entire partition**, blocking every later (good) record behind it. This is a liveness/availability bug, not a mere data-quality nuisance. ## Mitigations 1. **ErrorHandlingDeserializer (Spring Kafka)**: wraps your real deserializer; on failure it doesn't throw out of poll() but yields a record carrying the deserialization exception, which an error handler can route to a **dead-letter topic (DLT)** and let the consumer commit past it. 2. **Delegating/safe deserializer**: catch the exception inside `deserialize()`, return `null` or a sentinel, and side-channel the raw bytes + offset to a DLT for later inspection; your processing code skips nulls and commits. 3. **Manual handling**: catch the exception, log the failing topic-partition-offset, seek past it (`consumer.seek(tp, offset+1)`) deliberately, and commit — only safe if skipping is acceptable. 4. **Schema Registry + compatibility rules** prevent many poison pills upstream by rejecting incompatible producers. ## Practical guidance - Decide your policy up front: skip, DLT, or halt. Silent skipping can drop data; halting can take down a consumer. - DLT is the common production pattern: isolate the bad record, keep the partition flowing, alert, and reprocess later. - Validate at produce time too, so fewer bad records exist to consume.
- Why does one unparseable record block an entire partition?Deserialization happens inside poll(), so the bad record throws before your code sees it; the committed offset is still before it, so you can never advance past it. The loop re-polls the same offset and re-throws indefinitely, stalling that partition and all good records behind it.
- How do you handle poison pills in production?Wrap the deserializer (e.g. Spring's ErrorHandlingDeserializer) or use a delegating one that catches the failure, routes the raw bytes to a dead-letter topic with the topic/partition/offset, and lets the consumer commit past the bad record so the partition keeps flowing. Add a Schema Registry to prevent many bad records upstream.
- Where in the consumer lifecycle does deserialization run, and why does that matter?Inside poll(), before records reach your application loop. That's why a deserialization failure crashes poll() itself rather than being something your for-loop can simply skip — you must intervene at the deserializer or error-handler level.
saying these in an interview costs you the question
- Thinking Kafka stores typed objects, not bytes
- Believing a bad record is automatically skipped
- Claiming deserialization happens in your loop where you can try/catch per record
- Ignoring the partition-stall (liveness) impact of a poison pill