skip to content

What are Avro logical types, and how do decimal, timestamp-millis, and uuid work over Kafka?

level: seniorimportance: should knowfreq 40%

answer

  1. logicalType annotates a primitive base type
  2. decimal = bytes/fixed + precision + scale (unscaled int)
  3. timestamp-millis = long epoch ms UTC → Instant
  4. uuid = string → UUID
  5. no converter → you get the raw primitive

basics

~20 s

Logical types annotate a primitive Avro type with semantic meaning. decimal is bytes/fixed with precision+scale, timestamp-millis is a long of epoch millis, and uuid is a string. Both sides must share the schema to decode them correctly.

solid answer

~50 s

Avro **logical types** layer semantic meaning on top of primitive types via a `logicalType` attribute, so the wire encoding stays a known primitive while libraries map it to a richer language type. Key ones: **decimal** is stored as `bytes` (or `fixed`) carrying `precision` and `scale`, encoding the unscaled two's-complement integer — used for exact money math without float error. **timestamp-millis** (and `timestamp-micros`) is a `long` of milliseconds since the Unix epoch (UTC); **date** is an int of days; **time-millis** is an int. **uuid** is a `string` constrained to UUID format. On Kafka, the writer schema (carrying these annotations) is registered, and the deserializer maps decimal→`BigDecimal`, timestamp-millis→`Instant`/`long`, uuid→`UUID`, provided conversions are enabled (Avro `SpecificData`/`GenericData` with logical-type conversions, or `avro.use.logical.type.converters` in newer plugins). If a consumer ignores the logicalType, it just sees the raw primitive (a long or bytes).

go deeper

for a junior

Know logical types add meaning (timestamp, decimal, uuid) on top of plain Avro primitives.

for a middle

State the base type and target type for decimal, timestamp-millis, and uuid.

for a senior

Explain unscaled-integer decimal encoding, converter requirements, and the silent-degradation bug.

for a principal

Govern decimal precision/scale as contract, align Avro and Connect logical-type mappings across the platform.

Avro has a small set of **primitive types**: null, boolean, int, long, float, double, bytes, string. These can't natively express "this long is a timestamp" or "these bytes are an exact decimal." **Logical types** solve that: they are an annotation (`"logicalType": "..."`) attached to a primitive (the **base type**), telling Avro libraries to map the raw bytes to a richer in-memory type while keeping the wire encoding as that primitive. **decimal** — base type `bytes` or `fixed`. It carries two attributes: `precision` (total significant digits) and `scale` (digits after the decimal point). The value is encoded as the **unscaled integer** in two's-complement big-endian bytes; e.g. `12.34` with scale 2 is the integer `1234`. This gives **exact** arithmetic (critical for money), avoiding binary floating-point rounding. ```json {"name":"amount","type":{"type":"bytes","logicalType":"decimal","precision":10,"scale":2}} ``` Maps to `java.math.BigDecimal`. **timestamp-millis** — base type `long`, value = milliseconds since `1970-01-01T00:00:00Z` (UTC). `timestamp-micros` is the microsecond variant. Maps to `java.time.Instant` (or a long). **date** = int days since epoch → `LocalDate`; **time-millis** = int ms since midnight → `LocalTime`. Avro also has `local-timestamp-millis` (wall-clock, not UTC) since 1.10. **uuid** — base type `string`, value must be a canonical UUID string. Maps to `java.util.UUID`. **How conversion happens:** logical types only become rich types if the Avro **conversion** is registered. With generated SpecificRecords, the Avro Maven/Gradle plugin can emit code that uses `BigDecimal`/`Instant`/`UUID` when configured (e.g. `enableDecimalLogicalType=true`, and logical-type converters enabled). With GenericRecord you add conversions to `GenericData`. If conversions aren't enabled, you just receive the **base primitive** — a `ByteBuffer`, a `long`, or a `String`. **Kafka specifics & edge cases:** - The `logicalType` lives in the **schema**, which is registered and resolved by ID — so both producer and consumer see the same annotation. A consumer that doesn't enable converters silently degrades to the raw primitive, which is a common "why is my decimal a ByteBuffer?" bug. - An **unknown** logicalType is, by spec, ignored — Avro falls back to the base type rather than failing. So a typo in `logicalType` won't error; you just lose the rich mapping. - **decimal** evolution: changing `scale`/`precision` is *not* a transparent change — the unscaled-integer encoding depends on scale, so altering it can corrupt interpretation. Treat precision/scale as part of the contract. - **Connect** has its own logical-type mapping (`org.apache.kafka.connect.data.Decimal/Timestamp`) layered over Avro — mismatches here cause classic Connect type errors. **Summary:** logical types = primitive on the wire + semantic annotation in the schema + a library conversion to a rich type. Money → decimal(bytes, precision, scale)→BigDecimal; time → timestamp-millis(long)→Instant; identity → uuid(string)→UUID.

  • How is an Avro decimal physically encoded on the wire?
    As the unscaled integer value in two's-complement big-endian bytes (base type bytes or fixed), with precision and scale declared in the schema. 12.34 scale 2 = integer 1234.
  • What happens if a consumer doesn't enable logical-type conversions?
    It receives the underlying base primitive — a ByteBuffer for decimal, a long for timestamp-millis, a String for uuid — instead of BigDecimal/Instant/UUID.

saying these in an interview costs you the question

  • Saying decimal is stored as a double/float (defeats the exactness purpose)
  • Claiming timestamp-millis stores a formatted date string rather than an epoch long
  • Assuming changing a decimal's scale is a safe, transparent schema change

context