What are Avro logical types, and how do decimal, timestamp-millis, and uuid work over Kafka?
answer
- logicalType annotates a primitive base type
- decimal = bytes/fixed + precision + scale (unscaled int)
- timestamp-millis = long epoch ms UTC → Instant
- uuid = string → UUID
- no converter → you get the raw primitive
basics
~20 sLogical types annotate a primitive Avro type with semantic meaning. decimal is bytes/fixed with precision+scale, timestamp-millis is a long of epoch millis, and uuid is a string. Both sides must share the schema to decode them correctly.
solid answer
~50 sAvro **logical types** layer semantic meaning on top of primitive types via a `logicalType` attribute, so the wire encoding stays a known primitive while libraries map it to a richer language type. Key ones: **decimal** is stored as `bytes` (or `fixed`) carrying `precision` and `scale`, encoding the unscaled two's-complement integer — used for exact money math without float error. **timestamp-millis** (and `timestamp-micros`) is a `long` of milliseconds since the Unix epoch (UTC); **date** is an int of days; **time-millis** is an int. **uuid** is a `string` constrained to UUID format. On Kafka, the writer schema (carrying these annotations) is registered, and the deserializer maps decimal→`BigDecimal`, timestamp-millis→`Instant`/`long`, uuid→`UUID`, provided conversions are enabled (Avro `SpecificData`/`GenericData` with logical-type conversions, or `avro.use.logical.type.converters` in newer plugins). If a consumer ignores the logicalType, it just sees the raw primitive (a long or bytes).
go deeper
Know logical types add meaning (timestamp, decimal, uuid) on top of plain Avro primitives.
State the base type and target type for decimal, timestamp-millis, and uuid.
Explain unscaled-integer decimal encoding, converter requirements, and the silent-degradation bug.
Govern decimal precision/scale as contract, align Avro and Connect logical-type mappings across the platform.
Avro has a small set of **primitive types**: null, boolean, int, long, float, double, bytes, string. These can't natively express "this long is a timestamp" or "these bytes are an exact decimal." **Logical types** solve that: they are an annotation (`"logicalType": "..."`) attached to a primitive (the **base type**), telling Avro libraries to map the raw bytes to a richer in-memory type while keeping the wire encoding as that primitive. **decimal** — base type `bytes` or `fixed`. It carries two attributes: `precision` (total significant digits) and `scale` (digits after the decimal point). The value is encoded as the **unscaled integer** in two's-complement big-endian bytes; e.g. `12.34` with scale 2 is the integer `1234`. This gives **exact** arithmetic (critical for money), avoiding binary floating-point rounding. ```json {"name":"amount","type":{"type":"bytes","logicalType":"decimal","precision":10,"scale":2}} ``` Maps to `java.math.BigDecimal`. **timestamp-millis** — base type `long`, value = milliseconds since `1970-01-01T00:00:00Z` (UTC). `timestamp-micros` is the microsecond variant. Maps to `java.time.Instant` (or a long). **date** = int days since epoch → `LocalDate`; **time-millis** = int ms since midnight → `LocalTime`. Avro also has `local-timestamp-millis` (wall-clock, not UTC) since 1.10. **uuid** — base type `string`, value must be a canonical UUID string. Maps to `java.util.UUID`. **How conversion happens:** logical types only become rich types if the Avro **conversion** is registered. With generated SpecificRecords, the Avro Maven/Gradle plugin can emit code that uses `BigDecimal`/`Instant`/`UUID` when configured (e.g. `enableDecimalLogicalType=true`, and logical-type converters enabled). With GenericRecord you add conversions to `GenericData`. If conversions aren't enabled, you just receive the **base primitive** — a `ByteBuffer`, a `long`, or a `String`. **Kafka specifics & edge cases:** - The `logicalType` lives in the **schema**, which is registered and resolved by ID — so both producer and consumer see the same annotation. A consumer that doesn't enable converters silently degrades to the raw primitive, which is a common "why is my decimal a ByteBuffer?" bug. - An **unknown** logicalType is, by spec, ignored — Avro falls back to the base type rather than failing. So a typo in `logicalType` won't error; you just lose the rich mapping. - **decimal** evolution: changing `scale`/`precision` is *not* a transparent change — the unscaled-integer encoding depends on scale, so altering it can corrupt interpretation. Treat precision/scale as part of the contract. - **Connect** has its own logical-type mapping (`org.apache.kafka.connect.data.Decimal/Timestamp`) layered over Avro — mismatches here cause classic Connect type errors. **Summary:** logical types = primitive on the wire + semantic annotation in the schema + a library conversion to a rich type. Money → decimal(bytes, precision, scale)→BigDecimal; time → timestamp-millis(long)→Instant; identity → uuid(string)→UUID.
- How is an Avro decimal physically encoded on the wire?As the unscaled integer value in two's-complement big-endian bytes (base type bytes or fixed), with precision and scale declared in the schema. 12.34 scale 2 = integer 1234.
- What happens if a consumer doesn't enable logical-type conversions?It receives the underlying base primitive — a ByteBuffer for decimal, a long for timestamp-millis, a String for uuid — instead of BigDecimal/Instant/UUID.
saying these in an interview costs you the question
- Saying decimal is stored as a double/float (defeats the exactness purpose)
- Claiming timestamp-millis stores a formatted date string rather than an epoch long
- Assuming changing a decimal's scale is a safe, transparent schema change