Why is deserializing untrusted data with Java's built-in serialization dangerous, and what mitigations exist?
answer
- Deserialization runs code (readObject/readResolve) via reflection
- Attacker picks classes from your classpath -> gadget chains
- RCE or DoS, regardless of data sensitivity
- ObjectInputFilter allowlist + depth/size limits (JEP 290)
- Best fix: don't deserialize untrusted data; use JSON/Protobuf
basics
~20 sReading objects from untrusted bytes lets an attacker control what classes get instantiated and which methods run during deserialization, which can lead to remote code execution. Don't deserialize untrusted input; prefer safe formats like JSON, or filter allowed classes.
solid answer
~40 sJava deserialization rebuilds objects from a byte stream by instantiating classes and invoking lifecycle hooks (readObject, readResolve, finalize) via reflection — before your code ever inspects the result. If the bytes are attacker-controlled, the attacker chooses which serializable classes on your classpath get constructed and chains their side effects ('gadget chains') to achieve denial of service or remote code execution. The data being non-sensitive doesn't matter; the act of deserializing is the vulnerability. Mitigations: never deserialize untrusted input; if you must, install an ObjectInputFilter (allowlist of expected classes, depth/size limits, JDK 9+), keep libraries patched, run with least privilege, and prefer data-only formats (JSON, Protobuf) with strict schemas and no automatic type binding. Architecturally, treat Java serialization as a JVM-internal, trusted-boundary tool only, and keep it off any external or cross-trust boundary.
code
java · 8 lines// Defense-in-depth if you MUST use java serialization (JDK 9+):
ObjectInputStream in = new ObjectInputStream(stream);
in.setObjectInputFilter(ObjectInputFilter.Config.createFilter(
"com.example.MyDto;java.util.*;" + // allowlist expected classes
"maxdepth=10;maxarray=1000;maxrefs=5000;" + // resource caps
"!*")); // reject everything else
MyDto dto = (MyDto) in.readObject(); // filtered before instantiation
// Better: avoid java serialization entirely for untrusted input.go deeper
Aware that deserializing untrusted data is dangerous and that JSON is generally safer than Java serialization.
Explains that deserialization runs code (readObject) before validation, so attacker-controlled bytes are risky.
Describes gadget chains, RCE vs DoS, and concrete mitigations like ObjectInputFilter allowlists, patching, and preferring data-only formats.
Sets organizational policy: ban Java serialization on untrusted boundaries, mandate versioned schema-based formats, define defense-in-depth (filters, least privilege, integrity) and a migration plan for legacy serialized stores.
## Recap: what deserialization does `ObjectInputStream.readObject()` reconstructs a Java object from bytes. Crucially it does this by **instantiating classes and invoking code via reflection** — the class's `readObject`, `readResolve`/`readExternal`, validation callbacks, and even `finalize` can run — **before** your application logic gets a chance to validate anything. The bytes essentially say 'create an instance of class X and populate it,' and the JVM obeys. ## Why untrusted input is the problem If the byte stream comes from somewhere an attacker can influence (an HTTP request, a queue message, a cookie, a cache), the attacker controls **which classes are instantiated** and with **what field values**. They don't need your class — any **serializable class on your classpath** is fair game, including ones in libraries. ## Gadget chains A **gadget** is a class whose deserialization side effects do something useful to an attacker (e.g. a `readObject` that calls a method on a field, which calls another method...). By nesting objects so that one gadget's side effect triggers the next, an attacker builds a **gadget chain** that ends in something dangerous — running a system command, opening a connection, or exhausting resources. Famous chains were found in widely used libraries (e.g. Commons-Collections), which is why the data's apparent harmlessness is irrelevant: the **mechanism** is the exploit. Outcomes range from **denial of service** (a deeply nested or huge stream that consumes CPU/memory — a 'billion laughs'-style bomb) to full **remote code execution (RCE)**. ## Mitigations (in priority order) 1. **Don't deserialize untrusted data at all.** This is the only fully reliable defense. Use a **data-only format** — JSON, Protobuf, Avro — parsed into known types with a strict schema and *no* polymorphic 'instantiate the class named in the payload' behavior (disable default typing in libraries like Jackson). 2. **ObjectInputFilter (JEP 290, JDK 9+; backported).** If you truly must use Java serialization, install a filter that **allowlists** the exact classes you expect and rejects everything else, and caps stream **depth, array size, and reference count** to block resource-exhaustion bombs. You can set it per-stream or process-wide (`jdk.serialFilter`). 3. **Keep dependencies patched** and minimize the classpath to shrink the gadget surface. 4. **Least privilege / isolation** — run the process with minimal OS permissions, network egress restrictions, and resource limits so a successful exploit is contained. 5. **Integrity protection** — if a stream must travel through an untrusted channel between two trusted endpoints, sign/MAC it so tampering is detected before deserialization (defense in depth, not a substitute for the above). ## Architectural stance At a system-design level, treat Java's built-in serialization as a **JVM-internal, fully trusted boundary** tool (e.g. between cooperating JVMs you fully control) and **never** expose it on an external or cross-trust boundary. New designs should default to a versioned, language-neutral, data-only format. The Java platform itself has been steadily de-emphasizing serialization for exactly these reasons. ## Relating back to the basics This is why `transient` matters for secrets (excluded data can't leak via the stream) and why the serialized form should be considered an **API and an attack surface**, not an implementation detail.
- If the serialized data contains no sensitive fields, is it safe to deserialize from an untrusted source?No. The vulnerability is the deserialization mechanism itself — it instantiates classes and runs their readObject/readResolve code before you can inspect anything. An attacker exploits the process (gadget chains) regardless of how innocuous the apparent data is.
- What does an ObjectInputFilter protect against, and what are its limits?It allowlists which classes may be instantiated and caps stream depth/array size/reference count, blocking unexpected gadget classes and resource-exhaustion bombs. Limits: it only helps if the allowlist is tight and correct, it doesn't make an allowed-but-vulnerable class safe, and it's weaker than simply not deserializing untrusted data.
Deserializing untrusted bytes is like running a program someone mailed you on a USB stick: you're not just reading data, you're executing whatever instructions it contains — using any tools (classes) already installed on your machine.
saying these in an interview costs you the question
- Believing it's safe because the payload 'is just data'
- Thinking validating the object after readObject() prevents the attack (code already ran)
- Relying on a denylist instead of an allowlist of classes
- Assuming a TLS channel makes the deserialized content trustworthy
- Treating Java serialization as appropriate for external/public APIs