skip to content

What are DataInputStream and DataOutputStream, and why would you use them instead of a plain InputStream/OutputStream?

level: juniorimportance: should knowfreq 45%

answer

  1. Filter/decorator wrapping a byte stream
  2. writeInt/readInt etc. for primitives + writeUTF for strings
  3. Big-endian, fixed layout, portable
  4. Not self-describing — read in same order/type
  5. Wrap a Buffered stream for speed

basics

~20 s

They are filter streams that read and write Java primitive types (int, double, boolean, etc.) and strings as raw bytes. You use them so you don't have to manually convert each primitive to and from bytes yourself.

solid answer

~40 s

DataOutputStream and DataInputStream are filter streams (you wrap them around another byte stream) that read/write Java primitives — int, long, double, float, boolean, char, byte, short — plus UTF strings, in a fixed, portable binary format. Instead of manually splitting an int into 4 bytes, you call writeInt(x) and later readInt(). They write in big-endian (network) byte order, so a file written on one JVM reads correctly on another regardless of the host's native endianness. You must read values back in exactly the same order and type you wrote them; there is no self-describing schema. Typical use: a compact custom binary file or wire format where you control both ends. For text, prefer character streams; for arbitrary object graphs, prefer object streams or a real serialization library.

code

java · 13 lines
java
try (DataOutputStream out =
         new DataOutputStream(new BufferedOutputStream(new FileOutputStream("data.bin")))) {
    out.writeInt(42);
    out.writeDouble(3.14);
    out.writeUTF("hello");
}

try (DataInputStream in =
         new DataInputStream(new BufferedInputStream(new FileInputStream("data.bin")))) {
    int i = in.readInt();      // 42  -- same order/type as written
    double d = in.readDouble(); // 3.14
    String s = in.readUTF();    // "hello"
}

go deeper

for a junior

Knows Data streams read/write primitives so you don't hand-convert bytes, and that you wrap them around a FileOutputStream/FileInputStream.

for a middle

Explains the fixed big-endian layout, the read-in-same-order requirement, writeUTF's length prefix, and adds buffering for performance.

for a senior

Discusses portability guarantees, the lack of a self-describing schema as a design trade-off, the 64KB writeUTF limit, and when to choose Data streams vs object streams vs a schema format.

for a principal

Frames Data streams within the decorator-based java.io design, weighs them against schema-evolving formats (protobuf/Avro) for long-lived persisted data, and sets team conventions for binary formats and versioning.

## The problem they solve In Java, the lowest-level I/O abstraction is the **byte stream**: `InputStream` (read bytes) and `OutputStream` (write bytes). A byte is a value from 0–255. But programs work with richer values — an `int` (32 bits), a `long` (64 bits), a `double`, a `boolean`, text. To store an `int` like `1000` in a file, you must somehow turn it into a sequence of bytes, and to read it back you must reassemble those bytes into an `int`. Doing that by hand (bit-shifting, masking) is tedious and error-prone. ## What Data streams are `DataOutputStream` and `DataInputStream` are **filter streams** (also called decorator streams). A filter stream does not talk to a file or socket directly; it **wraps** another stream and adds behavior. You construct them like this: ``` DataOutputStream out = new DataOutputStream(new FileOutputStream("data.bin")); ``` The `FileOutputStream` does the actual writing of bytes to disk; the `DataOutputStream` adds convenient methods that turn primitives into bytes for you: `writeInt`, `writeLong`, `writeDouble`, `writeFloat`, `writeBoolean`, `writeChar`, `writeByte`, `writeShort`, and `writeUTF` (for strings). On the reading side, `DataInputStream` mirrors them: `readInt`, `readLong`, `readDouble`, `readUTF`, etc. ## Key properties 1. **Fixed binary layout.** `writeInt` always emits exactly 4 bytes; `writeLong` 8 bytes; `writeDouble` 8 bytes. The format is defined by the `DataInput`/`DataOutput` interfaces, not left to the platform. 2. **Big-endian (network) byte order.** The most-significant byte is written first. This is the same order used in network protocols. Because it is fixed, a file written on a little-endian machine can be read correctly on a big-endian machine — the format is **portable** across platforms. 3. **No schema / not self-describing.** The bytes contain only the values, not their types or names. The reader must know the exact order and types that were written. If you write an int then a double, you must read an int then a double. Reading them in the wrong order, or calling `readLong` where `writeInt` was used, silently produces garbage. 4. **`writeUTF`/`readUTF` use a modified UTF-8.** The string is prefixed with a 2-byte unsigned length, so the reader knows where it ends. (This 2-byte length limits a single `writeUTF` string to 65535 encoded bytes.) ## When to use them - A **compact custom binary file format** where you own both the writer and the reader. - A **simple wire protocol** over a socket where both ends agree on the field order. - Anywhere you need primitives stored portably without dragging in a serialization framework. ## When not to - For human-readable **text**, use character streams (`Reader`/`Writer`) instead — Data streams are binary. - For **arbitrary object graphs**, Object streams or a real serialization library (Protocol Buffers, JSON, etc.) are better because Data streams have no notion of objects. - When the format will evolve and you need forward/backward compatibility, a schema-based format beats hand-ordered Data-stream fields. ## A note on buffering Data streams call the underlying stream once per primitive, which can be slow for many small writes. Wrap a `BufferedOutputStream`/`BufferedInputStream` in between for performance: ``` new DataOutputStream(new BufferedOutputStream(new FileOutputStream("data.bin"))) ```

  • Why is fixed big-endian byte order important?
    It makes the format portable: a file written on one machine/JVM reads identically on another regardless of the host's native endianness, because the layout never depends on the platform.
  • What happens if you write an int but read it back as a long?
    readLong consumes 8 bytes instead of 4, so it pulls in bytes that belonged to the next value and returns a meaningless number; everything after that is misaligned. There is no error — just silent corruption.

saying these in an interview costs you the question

  • Thinking Data streams store type information so any reader can decode them — they do not; the reader must know the layout.
  • Confusing them with character streams (Reader/Writer) for text.
  • Assuming byte order depends on the host platform — it is always big-endian.
  • Believing writeUTF can store arbitrarily long strings (it is capped near 64KB encoded).

context