skip to content

What are the binary serialization formats in kotlinx.serialization (ProtoBuf, CBOR), and why would you choose them over JSON?

level: juniorimportance: should knowfreq 45%

answer

  1. Same @Serializable model, different format object
  2. encodeToByteArray / decodeFromByteArray (not String)
  3. ProtoBuf = tag numbers, no names; CBOR = self-describing
  4. Separate artifacts: -protobuf, -cbor
  5. Both still @ExperimentalSerializationApi

basics

~20 s

ProtoBuf and CBOR encode your data into compact bytes instead of text like JSON. You use them when you want smaller, faster messages, for example over a network. Mark the class @Serializable and call encodeToByteArray.

solid answer

~40 s

kotlinx.serialization ships pluggable formats. JSON is text; ProtoBuf and CBOR are binary, producing a ByteArray that is smaller and faster to parse. You annotate the class with @Serializable, then instead of Json.encodeToString you call ProtoBuf.encodeToByteArray(value) / ProtoBuf.decodeFromByteArray(bytes) (or the Cbor instance). The same @Serializable model works across formats — that is the whole point of the format-agnostic design. ProtoBuf is best when payload size and speed matter (RPC, mobile, storage) and you control both ends; CBOR is a self-describing binary format (an IETF standard, RFC 8949) often used in IoT/COSE. Both are in the kotlinx-serialization-protobuf / kotlinx-serialization-cbor artifacts, separate from -json, and both are still marked @ExperimentalSerializationApi.

code

kotlin · 6 lines
kotlin
@Serializable
data class Point(val x: Int, val y: Int)

val bytes = ProtoBuf.encodeToByteArray(Point(3, 4))
val p = ProtoBuf.decodeFromByteArray<Point>(bytes)
// JSON of the same value would be the text {"x":3,"y":4}

go deeper

for a junior

Knows binary formats produce bytes, are smaller than JSON, and reuse @Serializable.

for a middle

Names the correct functions and artifacts and distinguishes ProtoBuf (tags) from CBOR (self-describing).

for a senior

Explains the format-agnostic architecture, when each format fits, and the experimental status implications.

for a principal

Frames the tradeoff for a system: schema control, interop with other languages/COSE, and governance of an experimental API in production.

## The format-agnostic design kotlinx.serialization separates *what* is serializable (your `@Serializable` class) from *how* it is encoded (the **format**). You write the model once and pick a format object: `Json`, `ProtoBuf`, `Cbor`. Each format exposes encode/decode functions. ## JSON vs binary - **JSON** is human-readable text: `{"id":1,"name":"Ada"}`. Easy to debug, but verbose and slower to parse. - **Binary** formats emit raw bytes (a `ByteArray`). They are smaller and faster, but not human-readable. ## ProtoBuf Implements the **Protocol Buffers** wire format. Fields are identified by a **field number** (a tag), not by name, so names are not stored — that is a big part of the size saving. Best when you control both producer and consumer and want maximum compactness/speed (RPC, persistence, mobile). ## CBOR **Concise Binary Object Representation**, an IETF standard (RFC 8949). It is *self-describing*: field names/keys are encoded, more like a binary JSON. Common in IoT, WebAuthn/COSE. ## The API ```kotlin import kotlinx.serialization.Serializable import kotlinx.serialization.protobuf.ProtoBuf import kotlinx.serialization.cbor.Cbor import kotlinx.serialization.encodeToByteArray import kotlinx.serialization.decodeFromByteArray @Serializable data class User(val id: Int, val name: String) val bytes: ByteArray = ProtoBuf.encodeToByteArray(User(1, "Ada")) val back: User = ProtoBuf.decodeFromByteArray(bytes) val cborBytes = Cbor.encodeToByteArray(User(1, "Ada")) ``` Note the function names differ from JSON: binary formats use `encodeToByteArray` / `decodeFromByteArray` (there is no `encodeToString`, because the output is bytes). ## Artifacts & status - ProtoBuf lives in `org.jetbrains.kotlinx:kotlinx-serialization-protobuf`. - CBOR lives in `kotlinx-serialization-cbor`. - Both APIs are annotated `@ExperimentalSerializationApi`, so they may change between versions and require an opt-in.

  • Why is there no encodeToString for ProtoBuf?
    Because its output is raw bytes, not text. Binary formats expose encodeToByteArray/decodeFromByteArray; only text formats like Json have encodeToString.
  • Can the same @Serializable class be encoded as both JSON and ProtoBuf?
    Yes — that is the format-agnostic design. The generated serializer is reused; you just swap the format object.

JSON is a labelled spreadsheet you can read; ProtoBuf is the same data zipped into numbered columns only the program understands.

saying these in an interview costs you the question

  • Thinks ProtoBuf/CBOR need a different annotation than @Serializable
  • Calls encodeToString on ProtoBuf
  • Claims they are part of the JSON artifact / need no extra dependency
  • Says CBOR and ProtoBuf are identical wire formats
  • Unaware both are experimental

context