Why are data formats like JSON or Protobuf preferred over Java native serialization, and what must you still watch for?
answer
- data vs object graph — no class identity in JSON/Protobuf
- consumer picks the type, not the attacker
- cross-language + schema evolution + readable
- danger returns with default/polymorphic typing
- bind to concrete types; allowlist subtypes if polymorphism needed
basics
~20 sJSON and Protobuf describe plain data, not Java objects, so reading them doesn't automatically run code or rebuild arbitrary classes. They're also cross-language and more stable across versions. You still must avoid letting a library auto-create arbitrary types from the input.
solid answer
~50 sNative Java serialization encodes a Java object graph including class identities and reconstruction logic, so reading it instantiates classes and runs magic methods — the root cause of gadget-chain RCE. JSON and Protobuf instead encode **pure data**: fields and values with no embedded behavior. A JSON or Protobuf parser maps that data into types **you** declare, so attacker-controlled bytes don't choose which classes get built. They're also language-agnostic, human-readable (JSON) or compact and schema-evolved (Protobuf), and avoid serialVersionUID brittleness. The caveat: safety isn't automatic. **Polymorphic deserialization** features — e.g. Jackson's default typing, or any 'deserialize to type named in the payload' setting — re-create the exact same vulnerability by letting the document pick the target class. So you must: disable default/polymorphic typing, bind to concrete known types, validate the resulting data, and keep parser libraries patched.
code
java · 8 linesObjectMapper mapper = new ObjectMapper();
// SAFE: bind to a concrete, known type — payload cannot pick the class.
User user = mapper.readValue(untrustedJson, User.class);
// DANGEROUS on untrusted input: default typing lets the JSON name the class
// (e.g. {"@class":"com.evil.Gadget", ...}) -> reintroduces RCE risk.
// mapper.activateDefaultTyping(LaissezFaireSubTypeValidator.instance); // do NOT do thisgo deeper
Knows JSON/Protobuf are common, cross-language data formats and are generally safer than Java serialization.
Explains that these formats carry data not class identity, so the consumer chooses the type, and that polymorphic typing can reintroduce the risk.
Details the data-vs-behavior distinction, schema evolution benefits, Jackson default-typing CVEs, and concrete mitigations (concrete binding, subtype allowlists, validation, size caps).
Sets a serialization-format strategy across services (Protobuf/JSON contracts, banning native serialization and polymorphic typing on untrusted boundaries), and governs library patch policy and schema-evolution standards.
## The fundamental difference: data vs. objects Java native serialization encodes an **object graph**: not just the values, but the **class identity** of each object and instructions to invoke its reconstruction logic. Reading it therefore *instantiates classes and runs code* — the source of gadget-chain RCE. **JSON** (a text format of objects, arrays, strings, numbers, booleans, null) and **Protocol Buffers / Protobuf** (a compact binary format defined by a `.proto` schema) encode **pure data only**. There is no behavior and no Java class identity in the bytes. The bytes say 'name = "Ada", age = 36'; they do **not** say 'build a com.evil.Gadget and call its readObject'. ## Why that's safer Because the payload carries no class identity, the **consumer** decides what type to build it into. You write `mapper.readValue(json, User.class)` or use the generated Protobuf message class. Attacker-controlled bytes can fill in field values but cannot choose *which* class is instantiated. That removes the 'data drives which code runs' property at the core of deserialization attacks. The worst a malicious value usually does is be a wrong-but-typed value, which your normal input validation handles. ## Additional practical benefits - **Cross-language / interoperable:** JSON and Protobuf are read by virtually every language; native Java serialization is Java-only. - **Schema evolution:** Protobuf has explicit forward/backward-compatibility rules (numbered fields, optional/repeated); JSON tolerates extra/missing fields gracefully. Native serialization is brittle — a class change can break old data and forces careful `serialVersionUID` management. - **Inspectable:** JSON is human-readable; Protobuf has tooling. Easier to debug, log (carefully), and validate. - **Compact & fast (Protobuf):** smaller and quicker to parse than verbose formats when performance matters. ## The crucial caveat: polymorphic / type-from-payload deserialization Switching format is **not** automatically safe. Many libraries offer a feature where the **document itself names the concrete class** to instantiate, to support polymorphism: - Jackson's **default typing** / `@JsonTypeInfo` writes a type id (e.g. `@class` or `@type`) into the JSON, and on read it instantiates that class. If enabled with untrusted input, an attacker puts a gadget class name there — **the JSON-deserialization equivalents of the native attack** (historically a long list of Jackson CVEs). - The same risk exists in any 'instantiate the type named in the data' mechanism in other libraries/languages (YAML loaders that build objects, .NET `TypeNameHandling`, etc.). **Mitigations:** 1. **Bind to concrete, known types.** Never enable global default typing on untrusted input. If polymorphism is required, use a **closed allowlist** of subtypes (e.g. Jackson `@JsonSubTypes` / a validated `PolymorphicTypeValidator`). 2. **Validate the data after parsing** (ranges, required fields, sizes) — JSON parsing doesn't validate business rules. 3. **Cap input size and nesting** to avoid DoS (deeply nested JSON, huge numbers/arrays). 4. **Keep parser libraries current** — patched promptly for the recurring polymorphic-typing CVEs. ## Bottom line Prefer JSON/Protobuf because they transport **data, not behavior**, so the consumer — not the attacker — picks the target type. Just don't reintroduce the original flaw by enabling 'type named in the payload' features on untrusted input.
- How can JSON deserialization still lead to RCE?Via polymorphic/default typing where the JSON document names the concrete class to instantiate (e.g. Jackson default typing, an @class field). An attacker supplies a gadget class name, re-creating the native-serialization attack.
- What's the safest way to support polymorphism in JSON?Use a closed allowlist of permitted subtypes (e.g. Jackson @JsonSubTypes with a strict PolymorphicTypeValidator) rather than global default typing, so only known types can be instantiated.
saying these in an interview costs you the question
- Assuming any JSON/Protobuf usage is automatically safe — polymorphic typing reintroduces RCE
- Enabling Jackson default typing or @class type ids on untrusted input
- Skipping post-parse validation because 'the parser handled it'
- Thinking the benefit is only readability rather than the data-vs-behavior security property