skip to content

Data & Object Streams

DataInputStream and DataOutputStream read and write primitives in a portable binary format, while the Object streams serialize whole object graphs. Both are decorators over plain byte streams, which is the structural point worth making.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

How does Java object serialization with ObjectOutputStream/ObjectInputStream work, and what is required to make a class serializable?

level: middleimportance: must knowfreq 65%

answer

  1. Serializable = empty marker interface
  2. Recursive over the whole object graph
  3. Handle table → shared refs + cycles, written once
  4. readObject skips the constructor
  5. transient/static excluded; serialVersionUID controls versions

basics

~10 s

ObjectOutputStream.writeObject turns an object (and everything it references) into bytes; ObjectInputStream.readObject rebuilds it. The class must implement the Serializable marker interface, and every field that should be saved must itself be serializable.

solid answer

~50 s

Java serialization converts an object graph into a byte stream and back. You wrap a byte stream in ObjectOutputStream and call writeObject(obj); ObjectInputStream.readObject() reconstructs it. For a class to be serializable it must implement java.io.Serializable, a marker interface with no methods — it just flags the class as eligible. Serialization is recursive: every non-transient, non-static field is written, so all referenced objects must also be serializable or you get a NotSerializableException. The runtime walks the object graph, tracking already-seen objects by identity so shared and cyclic references are preserved and written only once. On read, deserialization does NOT call the class's constructor — it allocates the object and restores fields directly. Mark fields you don't want saved (caches, passwords, threads) as transient. Add a serialVersionUID to control version compatibility. It's powerful but has real security and evolution pitfalls.

code

java · 12 lines
java
class Account implements Serializable {
    private static final long serialVersionUID = 1L;
    private String owner;          // serialized
    private transient String pin;  // skipped -- becomes null on read
}

try (ObjectOutputStream out = new ObjectOutputStream(new FileOutputStream("acct.ser"))) {
    out.writeObject(account);
}
try (ObjectInputStream in = new ObjectInputStream(new FileInputStream("acct.ser"))) {
    Account a = (Account) in.readObject(); // pin == null
}

go deeper

for a junior

Knows you call writeObject/readObject, the class must implement Serializable, and transient skips a field.

for a middle

Explains the recursive object-graph walk, the requirement that referenced fields be serializable, transient/static exclusion, and the basic role of serialVersionUID.

for a senior

Covers the handle table preserving shared/cyclic references, constructor bypass on read, writeObject/readObject and Externalizable customization, versioning strategy, and the security implications of untrusted deserialization.

for a principal

Sets organizational policy against native serialization for untrusted or long-lived data, favors schema-based formats with explicit migration, knows gadget-chain attack surface and JEP 290 filtering, and guides safe migration of legacy serialized formats.

## What serialization is **Serialization** is turning a live in-memory object into a flat sequence of bytes that can be stored or sent, and **deserialization** is rebuilding the object from those bytes. Java has this built into the standard library via two filter streams: `ObjectOutputStream` (writes objects) and `ObjectInputStream` (reads them). ``` ObjectOutputStream out = new ObjectOutputStream(new FileOutputStream("obj.ser")); out.writeObject(myObject); ObjectInputStream in = new ObjectInputStream(new FileInputStream("obj.ser")); MyType x = (MyType) in.readObject(); ``` These stream classes are built **on top of byte streams** (just like Data streams), and internally they also use the Data-stream primitive-writing machinery to encode field values. ## The Serializable marker interface For an object to be written, its class must implement `java.io.Serializable`. This is a **marker interface** — it declares no methods at all. Its only purpose is to signal to the runtime "instances of this class are allowed to be serialized." If you try to `writeObject` an instance whose class is not `Serializable`, you get a `NotSerializableException` at runtime. ## It is recursive over the object graph An object usually references other objects (a `Person` has an `Address`, a `List<Order>`, etc.). Serialization is **recursive**: when you serialize an object, the runtime also serializes every object reachable through its non-`static`, non-`transient` fields. This whole web is the **object graph**. Two consequences: 1. **Every reachable object must also be serializable**, or the whole operation fails with `NotSerializableException` pointing at the offending field. 2. The runtime tracks objects it has already written **by identity** using a handle table. If the same object is referenced from two places, it is written **once** and the second reference becomes a back-pointer. This means **shared references and cycles** (A points to B, B points back to A) are handled correctly and won't loop forever. ## Deserialization bypasses the constructor This surprises many people: when `readObject` rebuilds an object, it does **not** run the class's normal constructor. It allocates a blank instance and restores the field values directly from the stream. (The constructor of the first *non-serializable* superclass does run, to initialize the inherited part.) This is why invariants you normally enforce in a constructor can be violated by a crafted stream — a security concern. ## transient and static fields - A field marked **`transient`** is skipped — its value is not written, and on read it gets the default (`null`, `0`, `false`). Use this for things that shouldn't or can't be serialized: caches, open file handles, threads, secrets like passwords. - **`static`** fields belong to the class, not the instance, so they are never part of an instance's serialized form. ## serialVersionUID and versioning Each serializable class has a `serialVersionUID` — a `long` version stamp. If you don't declare one, the compiler/runtime computes one from the class's structure. When you deserialize, the stream's UID must match the current class's UID, or you get an `InvalidClassException`. Because the auto-computed value changes whenever you tweak the class, the recommended practice is to **declare it explicitly**: ``` private static final long serialVersionUID = 1L; ``` This gives you control: you decide when a change is compatible (keep the UID) versus breaking (bump it). ## Customization hooks Classes can customize the process by implementing private `writeObject(ObjectOutputStream)` / `readObject(ObjectInputStream)` methods, or by implementing `Externalizable` for fully manual control. The `readResolve`/`writeReplace` hooks let you substitute objects (e.g. to preserve singletons). ## Why people increasingly avoid it Native Java serialization is convenient but has well-known problems: it is a **security risk** (deserializing untrusted data can trigger remote code execution via gadget chains), it is **brittle across versions**, it is **Java-only** (not interoperable), and it can be verbose. For new systems, JSON, Protocol Buffers, or other schema-based formats are usually preferred. Serialization is still important to *understand* because it appears in legacy code, RMI, caching layers, and many interview questions.

  • Why mark a field transient?
    To exclude it from serialization — for data that can't or shouldn't be persisted (caches, threads, file handles, secrets). On deserialization a transient field gets its default value, so you may need to recompute it in a custom readObject.
  • What does serialVersionUID do and why declare it explicitly?
    It is a version stamp compared at deserialization; a mismatch throws InvalidClassException. Auto-computed values change on any structural edit, so declaring it explicitly gives you control over which class changes are treated as compatible.

saying these in an interview costs you the question

  • Saying Serializable has methods you must implement — it is empty.
  • Claiming deserialization calls the no-arg constructor (it generally does not for serializable classes).
  • Forgetting that referenced fields must also be serializable.
  • Thinking static fields are serialized with the instance.
  • Ignoring the security risk of deserializing untrusted input.

context

open as a page

What are DataInputStream and DataOutputStream, and why would you use them instead of a plain InputStream/OutputStream?

level: juniorimportance: should knowfreq 45%

basics

~20 s

They are filter streams that read and write Java primitive types (int, double, boolean, etc.) and strings as raw bytes. You use them so you don't have to manually convert each primitive to and from bytes yourself.

open as a page

When would you choose DataOutputStream over ObjectOutputStream, and how do the two relate?

level: middleimportance: should knowfreq 42%

basics

~20 s

Use Data streams for a small, fixed set of primitives and strings where you control the exact byte layout. Use Object streams when you need to save whole objects and their references automatically. Both are filter streams built on byte streams.

open as a page

How does object serialization handle shared references, cyclic references, and transient fields in an object graph?

level: seniorimportance: should knowfreq 50%

basics

~10 s

Serialization remembers each object it has already written, so an object referenced from two places is saved once and cycles don't loop forever. Fields marked transient are not saved and come back as defaults.

open as a page

What are the security and schema-evolution risks of native Java serialization, and how do you mitigate them?

level: principalimportance: should knowfreq 40%

basics

~20 s

Deserializing untrusted bytes can let an attacker run code on your machine, and changing a class can break old saved data. Avoid deserializing untrusted input, use input filters, and prefer schema-based formats like JSON or Protocol Buffers.

open as a page