Why can calling ObjectInputStream.readObject() on attacker-controlled bytes lead to remote code execution, even when the expected class looks harmless?
answer
- stream names the classes -> attacker picks classpath classes
- magic methods run: readObject/readResolve/readExternal
- gadget chain = wiring ordinary library classes together
- ysoserial / Commons Collections InvokerTransformer
- vulnerable just by having a library on the classpath
basics
~20 sreadObject() rebuilds objects from raw bytes and runs special hook methods (like readObject/readResolve) on each class while doing so. An attacker sends bytes describing other classes already on your classpath, and those hooks run code the attacker chose.
solid answer
~40 sJava native deserialization is not a simple data parse: ObjectInputStream reads a class descriptor from the stream and reconstructs an object graph, invoking each class's magic methods (readObject, readResolve, readExternal, even finalize) as it goes. The attacker doesn't need your class to be dangerous; they craft a stream that instantiates *any* Serializable class on your classpath. By chaining several such classes whose hook methods call each other, they build a 'gadget chain' that ends in something like Runtime.exec or a JNDI lookup. Libraries such as Apache Commons Collections historically supplied ready-made gadgets. The root cause is that the stream chooses which classes to instantiate and what fields to set, before you ever see the result. Defenses: avoid native serialization for untrusted input, or constrain it with a JEP 290 ObjectInputFilter allowlist.
code
java · 9 lines// DANGEROUS: bytes come from an attacker (HTTP body, cookie, queue...)
try (var in = new ObjectInputStream(request.getInputStream())) {
// readObject() reads class names FROM THE STREAM, instantiates them,
// and runs their readObject/readResolve hooks during reconstruction.
Object o = in.readObject(); // gadget chain can fire here -> RCE
User user = (User) o; // the cast happens AFTER the damage
}
// The cast to User does NOT protect you: the malicious classes were
// already instantiated and their hook methods already executed.go deeper
Knows readObject() rebuilds objects from bytes and that doing it on untrusted data is dangerous; can name 'insecure deserialization' as a risk.
Explains that the stream names the classes and that hook methods (readObject/readResolve) run during reconstruction, so RCE is possible even for unexpected classes.
Articulates gadget chains, names a real example (Commons Collections / ysoserial), explains why the cast and transient don't help, and prescribes avoiding native deser + allowlists.
Frames it as a trust-boundary/code-execution-surface problem; weighs migrating off native serialization vs. JEP 290 filters org-wide, dependency hygiene to shrink the gadget surface, and detection (filter logging, RASP).
## What 'serialization' means **Serialization** is turning a live in-memory object into a flat sequence of bytes you can store or send over a network. **Deserialization** is the reverse: taking those bytes and rebuilding an equivalent object in memory. Java has a built-in ("native") mechanism for this: a class declares it implements the marker interface `java.io.Serializable`, and then `ObjectOutputStream.writeObject(obj)` produces the bytes and `ObjectInputStream.readObject()` reconstructs the object. ## Why it is not a harmless data parse A naive mental model is: "deserialization just fills in the fields of the object I asked for." That is wrong on two counts. 1. **The byte stream itself names the classes.** The serialized stream contains *class descriptors* — the fully-qualified class names and field layouts of every object in the graph. When `readObject()` runs, it reads a name from the stream, looks that class up on the JVM's classpath, and instantiates it. So the *attacker's bytes*, not your code, decide which classes get created. If a class named `EvilGadget` is on your classpath, a crafted stream can ask for it even if your code only ever expected a `User`. 2. **Reconstruction runs code — the 'magic methods'.** Serializable classes may define hook methods that the deserializer calls automatically during reconstruction: - `private void readObject(ObjectInputStream in)` — custom read logic, run for that class. - `Object readResolve()` — lets a class substitute the object returned (used for singletons/enums). - `void readExternal(ObjectInput in)` — for `Externalizable` classes, full control over reading. - Even `finalize()` can be reached when crafted objects are garbage-collected. Because these methods execute automatically, deserializing untrusted bytes is effectively *running code from those classes*. ## Gadget chains No single class usually says "run this shell command." Instead attackers assemble a **gadget chain**: a sequence of ordinary library classes whose hook methods, fields, and side effects can be wired together so that finishing the deserialization triggers a dangerous operation. A classic example used `org.apache.commons.collections.functors.InvokerTransformer`, which can be configured (via its serialized fields) to call an arbitrary method by reflection — chained through a `LazyMap` and a proxy so that simply building the map invokes `Runtime.getRuntime().exec("...")`. The famous **ysoserial** tool packages dozens of such chains for common libraries (Commons Collections, Spring, Groovy, etc.). Key insight: **you can be vulnerable just by having a library on the classpath**, even if your own code never uses that library's serialization. The attacker only needs *gadgets* to exist; they supply the chain in the bytes. ## The trust boundary The danger appears wherever Java native deserialization is applied to data an attacker can influence: HTTP request bodies, cookies, RMI/JMX endpoints, JMS messages, caches, message queues, deserialized session state, etc. This is catalogued as **insecure deserialization** (an OWASP Top-10 category). ## Why `transient` and `serialVersionUID` do not save you - `transient` marks a field to be *skipped* during serialization (e.g. a password or a socket). It limits what data crosses the wire but does **not** stop gadget classes from being instantiated. - `serialVersionUID` is a version stamp used to detect class-incompatibility between writer and reader; it is about correctness, not security. An attacker controls it in their stream anyway. ## Defenses (overview) 1. **Don't natively deserialize untrusted input.** Prefer a data-only format (JSON, Protocol Buffers) parsed into known types. 2. **JEP 290 `ObjectInputFilter` allowlists** (Java 9+, backported to 8u121+): constrain *which classes* may be deserialized, before they are instantiated — covered in the look-ahead/JEP-290 questions. 3. **Look-ahead deserialization**: validate the class name in the stream header *before* the object is built. 4. Remove unneeded gadget-bearing libraries; keep dependencies patched. The core takeaway: native `readObject()` is a *code-execution surface*, not a parser, because the stream decides the classes and the classes' hook methods run during reconstruction.
- If my code casts the result to User immediately, why doesn't that prevent the attack?Because the ClassCastException (if any) happens *after* readObject() finished building the whole object graph and ran every hook method. The malicious code already executed during reconstruction; the cast is too late.
- Does removing Commons Collections make me safe?It removes one well-known gadget source, but other libraries (Spring, Groovy, BeanShell, JDK internals) supply gadgets too. Removing gadgets is defense-in-depth, not a complete fix — you still need to avoid deserializing untrusted data or apply a class allowlist.
It's like a mail-order kit where the package itself tells the factory which parts to assemble and in what order. You expected a toy car, but the shipping label can instead order the factory to build a working flamethrower from parts already in the warehouse.
saying these in an interview costs you the question
- Thinking deserialization only 'fills fields' and can't run code.
- Believing the cast to the expected type protects you (it runs too late).
- Assuming you're safe because *your* classes are harmless — gadgets come from libraries.
- Confusing serialVersionUID or transient with a security control.