skip to content

Binary Formats (ProtoBuf, CBOR)

ProtoBuf and CBOR encode the same @Serializable classes compactly, with field numbers controlled by annotation. Their experimental status and the need for stable field numbers are the caveats to state.

part ofKotlinoverview, primer and where to startread it →
on this pageshow

questions

5

How does @ProtoNumber work in kotlinx.serialization ProtoBuf, and what happens if you omit it?

level: middleimportance: must knowfreq 55%

answer

  1. Tag = field number, not name, on the wire
  2. Omit → auto 1,2,3 in declaration order
  3. Reorder/insert/delete shifts implicit numbers → silent break
  4. Annotate explicitly, never reuse a number, only append
  5. Gaps in numbers are allowed

basics

~10 s

@ProtoNumber sets the field's tag number on the wire. ProtoBuf stores fields by number, not name. If you omit it, numbers are assigned automatically starting at 1 in declaration order.

solid answer

~40 s

In Protocol Buffers each field is identified on the wire by an integer **field number** (tag), not by its name. In kotlinx.serialization you set it with `@ProtoNumber(n)` on the property. These numbers are the contract between encoder and decoder: matching numbers must line up, and they must be stable across versions for compatibility. If you omit `@ProtoNumber`, kotlinx assigns numbers automatically as 1, 2, 3… in property declaration order. That auto-assignment is dangerous for schema evolution: reordering, inserting, or removing a property shifts the implicit numbers and silently breaks previously written data. So for any persisted or cross-service ProtoBuf, you annotate every field explicitly with a fixed `@ProtoNumber`, never reuse a retired number, and only append new ones. `@ProtoNumber` is `@ExperimentalSerializationApi`.

code

kotlin · 8 lines
kotlin
@Serializable
data class Order(
    @ProtoNumber(1) val id: Long,
    @ProtoNumber(2) val total: Int,
    @ProtoNumber(3) val note: String? = null,
)
// Adding a new field later: give it @ProtoNumber(4), keep it nullable.
// Never renumber id, total, or note.

go deeper

for a junior

Knows @ProtoNumber sets a tag and that ProtoBuf is number-based.

for a middle

Explains auto-assignment in declaration order and why reorder/insert breaks compatibility; applies explicit annotation.

for a senior

States the full evolution discipline (never renumber/reuse, append-only, nullable additions) and ties it to silent corruption.

for a principal

Establishes team conventions/lint for numbering, reserved ranges, and review gates so the wire contract stays stable across services.

## Field numbers are the identity ProtoBuf does not put field *names* on the wire. Each field is encoded with a small integer **tag** = `(field_number << 3) | wire_type`. The **field number** is what both sides agree on. The decoder reads tag 1, knows that means the first agreed field, and maps it back. ## @ProtoNumber `@ProtoNumber(n)` is the kotlinx annotation (in package `kotlinx.serialization.protobuf`) that pins a property to wire number `n`: ```kotlin import kotlinx.serialization.Serializable import kotlinx.serialization.protobuf.ProtoNumber @Serializable data class User( @ProtoNumber(1) val id: Long, @ProtoNumber(2) val name: String, @ProtoNumber(5) val email: String? = null, // gaps are fine ) ``` Numbers need not be contiguous; gaps are allowed and normal (you leave room for future fields). ## What omitting it does If you do **not** annotate, kotlinx auto-assigns `1, 2, 3, …` in **declaration order**. That works for a quick local round-trip, but ties the wire format to source order: - Reorder two properties → their numbers swap → old bytes decode into the wrong fields. - Insert a property in the middle → everything after it shifts by one. - Delete a property → all following numbers shift down. Because the decoder trusts the number, these are *silent* corruptions, not exceptions, when types happen to be compatible. ## Rules for safe evolution - **Annotate every field explicitly** in any schema that is persisted or shared. - **Never change** a field's number once data exists. - **Never reuse** a retired number for a different meaning. - **Only append** new fields with fresh numbers, and make them nullable / give defaults so old readers can skip them. ## Status `@ProtoNumber` is part of the experimental ProtoBuf API (`@ExperimentalSerializationApi`).

  • Are field numbers required to be contiguous?
    No. Gaps are allowed and even encouraged so you can reserve ranges for future fields without renumbering existing ones.
  • Why is auto-assignment risky in production?
    It binds the wire numbers to property declaration order, so any reorder/insert/delete silently changes the encoding and breaks already-stored or in-flight data.

Field numbers are like seat numbers on a ticket: the usher seats you by number, not by your name — change the numbering and everyone ends up in the wrong seat.

saying these in an interview costs you the question

  • Thinks ProtoBuf matches fields by property name
  • Says auto-assigned numbers are safe to reorder
  • Proposes reusing a deleted field's number
  • Believes numbers must be 1..N contiguous
  • Doesn't make appended fields nullable/defaulted

context

open as a page

What are the binary serialization formats in kotlinx.serialization (ProtoBuf, CBOR), and why would you choose them over JSON?

level: juniorimportance: should knowfreq 45%

basics

~20 s

ProtoBuf and CBOR encode your data into compact bytes instead of text like JSON. You use them when you want smaller, faster messages, for example over a network. Mark the class @Serializable and call encodeToByteArray.

open as a page

Show how to round-trip a value through ProtoBuf, and name common pitfalls (nullability, defaults, byte handling) when using binary formats.

level: middleimportance: should knowfreq 38%

basics

~10 s

Encode with ProtoBuf.encodeToByteArray and decode with decodeFromByteArray of the same type. Common mistakes: treating bytes as a String, mismatched field numbers, and assuming missing optional fields will fail instead of using defaults.

open as a page

What does the @ExperimentalSerializationApi status of ProtoBuf/CBOR mean in practice, and how do you handle it in production code?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Experimental means the API can change between versions and the compiler warns unless you opt in. In production you opt in deliberately, pin the library version, and wrap usage so a future change touches one place.

open as a page

Compare ProtoBuf and CBOR in kotlinx.serialization: wire format, size, and when to choose each.

level: seniorimportance: should knowfreq 40%

basics

~20 s

ProtoBuf stores only field numbers and values, so it is very compact but both sides need the same schema. CBOR also stores the keys, so it is a bit bigger but more self-describing, like a binary JSON.

open as a page