skip to content

Why are data formats like JSON or Protobuf preferred over Java native serialization, and what must you still watch for?

level: middleimportance: should knowfreq 55%

answer

  1. data vs object graph — no class identity in JSON/Protobuf
  2. consumer picks the type, not the attacker
  3. cross-language + schema evolution + readable
  4. danger returns with default/polymorphic typing
  5. bind to concrete types; allowlist subtypes if polymorphism needed

basics

~20 s

JSON and Protobuf describe plain data, not Java objects, so reading them doesn't automatically run code or rebuild arbitrary classes. They're also cross-language and more stable across versions. You still must avoid letting a library auto-create arbitrary types from the input.

solid answer

~50 s

Native Java serialization encodes a Java object graph including class identities and reconstruction logic, so reading it instantiates classes and runs magic methods — the root cause of gadget-chain RCE. JSON and Protobuf instead encode **pure data**: fields and values with no embedded behavior. A JSON or Protobuf parser maps that data into types **you** declare, so attacker-controlled bytes don't choose which classes get built. They're also language-agnostic, human-readable (JSON) or compact and schema-evolved (Protobuf), and avoid serialVersionUID brittleness. The caveat: safety isn't automatic. **Polymorphic deserialization** features — e.g. Jackson's default typing, or any 'deserialize to type named in the payload' setting — re-create the exact same vulnerability by letting the document pick the target class. So you must: disable default/polymorphic typing, bind to concrete known types, validate the resulting data, and keep parser libraries patched.

code

java · 8 lines
java
ObjectMapper mapper = new ObjectMapper();

// SAFE: bind to a concrete, known type — payload cannot pick the class.
User user = mapper.readValue(untrustedJson, User.class);

// DANGEROUS on untrusted input: default typing lets the JSON name the class
// (e.g. {"@class":"com.evil.Gadget", ...}) -> reintroduces RCE risk.
// mapper.activateDefaultTyping(LaissezFaireSubTypeValidator.instance); // do NOT do this

go deeper

for a junior

Knows JSON/Protobuf are common, cross-language data formats and are generally safer than Java serialization.

for a middle

Explains that these formats carry data not class identity, so the consumer chooses the type, and that polymorphic typing can reintroduce the risk.

for a senior

Details the data-vs-behavior distinction, schema evolution benefits, Jackson default-typing CVEs, and concrete mitigations (concrete binding, subtype allowlists, validation, size caps).

for a principal

Sets a serialization-format strategy across services (Protobuf/JSON contracts, banning native serialization and polymorphic typing on untrusted boundaries), and governs library patch policy and schema-evolution standards.

## The fundamental difference: data vs. objects Java native serialization encodes an **object graph**: not just the values, but the **class identity** of each object and instructions to invoke its reconstruction logic. Reading it therefore *instantiates classes and runs code* — the source of gadget-chain RCE. **JSON** (a text format of objects, arrays, strings, numbers, booleans, null) and **Protocol Buffers / Protobuf** (a compact binary format defined by a `.proto` schema) encode **pure data only**. There is no behavior and no Java class identity in the bytes. The bytes say 'name = "Ada", age = 36'; they do **not** say 'build a com.evil.Gadget and call its readObject'. ## Why that's safer Because the payload carries no class identity, the **consumer** decides what type to build it into. You write `mapper.readValue(json, User.class)` or use the generated Protobuf message class. Attacker-controlled bytes can fill in field values but cannot choose *which* class is instantiated. That removes the 'data drives which code runs' property at the core of deserialization attacks. The worst a malicious value usually does is be a wrong-but-typed value, which your normal input validation handles. ## Additional practical benefits - **Cross-language / interoperable:** JSON and Protobuf are read by virtually every language; native Java serialization is Java-only. - **Schema evolution:** Protobuf has explicit forward/backward-compatibility rules (numbered fields, optional/repeated); JSON tolerates extra/missing fields gracefully. Native serialization is brittle — a class change can break old data and forces careful `serialVersionUID` management. - **Inspectable:** JSON is human-readable; Protobuf has tooling. Easier to debug, log (carefully), and validate. - **Compact & fast (Protobuf):** smaller and quicker to parse than verbose formats when performance matters. ## The crucial caveat: polymorphic / type-from-payload deserialization Switching format is **not** automatically safe. Many libraries offer a feature where the **document itself names the concrete class** to instantiate, to support polymorphism: - Jackson's **default typing** / `@JsonTypeInfo` writes a type id (e.g. `@class` or `@type`) into the JSON, and on read it instantiates that class. If enabled with untrusted input, an attacker puts a gadget class name there — **the JSON-deserialization equivalents of the native attack** (historically a long list of Jackson CVEs). - The same risk exists in any 'instantiate the type named in the data' mechanism in other libraries/languages (YAML loaders that build objects, .NET `TypeNameHandling`, etc.). **Mitigations:** 1. **Bind to concrete, known types.** Never enable global default typing on untrusted input. If polymorphism is required, use a **closed allowlist** of subtypes (e.g. Jackson `@JsonSubTypes` / a validated `PolymorphicTypeValidator`). 2. **Validate the data after parsing** (ranges, required fields, sizes) — JSON parsing doesn't validate business rules. 3. **Cap input size and nesting** to avoid DoS (deeply nested JSON, huge numbers/arrays). 4. **Keep parser libraries current** — patched promptly for the recurring polymorphic-typing CVEs. ## Bottom line Prefer JSON/Protobuf because they transport **data, not behavior**, so the consumer — not the attacker — picks the target type. Just don't reintroduce the original flaw by enabling 'type named in the payload' features on untrusted input.

  • How can JSON deserialization still lead to RCE?
    Via polymorphic/default typing where the JSON document names the concrete class to instantiate (e.g. Jackson default typing, an @class field). An attacker supplies a gadget class name, re-creating the native-serialization attack.
  • What's the safest way to support polymorphism in JSON?
    Use a closed allowlist of permitted subtypes (e.g. Jackson @JsonSubTypes with a strict PolymorphicTypeValidator) rather than global default typing, so only known types can be instantiated.

saying these in an interview costs you the question

  • Assuming any JSON/Protobuf usage is automatically safe — polymorphic typing reintroduces RCE
  • Enabling Jackson default typing or @class type ids on untrusted input
  • Skipping post-parse validation because 'the parser handled it'
  • Thinking the benefit is only readability rather than the data-vs-behavior security property

context