skip to content

I/O Streams

The java.io hierarchy: byte versus character streams, the decorator layering that adds buffering and typing, and the closing, flushing and encoding rules. Interviewers use it as the textbook example of the Decorator pattern.

part ofJavaoverview, primer and where to startread it →
on this pageshow

explore

questions

18

What is the difference between byte streams and character streams in java.io, and when do you use each?

level: juniorimportance: must knowfreq 70%

answer

  1. InputStream/OutputStream = bytes; Reader/Writer = chars
  2. byte = 8 bits raw; char = 16-bit UTF-16, charset applied
  3. binary -> byte streams; text -> character streams
  4. parallel names: FileInputStream vs FileReader
  5. wrong choice on UTF-8 text = mojibake

basics

~20 s

Byte streams (InputStream/OutputStream) move raw bytes and suit binary data like images. Character streams (Reader/Writer) move text characters and handle the conversion between bytes and characters. Use byte streams for binary, character streams for text.

solid answer

~40 s

java.io has two parallel families. Byte streams are rooted at InputStream and OutputStream; they read and write data 8 bits at a time as raw bytes, with no idea what the bytes mean. Use them for binary content: images, audio, serialized objects, or when copying data verbatim. Character streams are rooted at Reader and Writer; they read and write text as 16-bit Java chars, automatically applying a character encoding (charset) to translate between bytes on disk and chars in memory. Use them for text files, so accented letters and non-Latin scripts decode correctly. The two families have parallel class names (FileInputStream vs FileReader, BufferedInputStream vs BufferedReader). Choosing wrong corrupts text: reading UTF-8 text byte-by-byte and treating each byte as a character breaks any multi-byte character.

go deeper

for a junior

Knows byte streams = InputStream/OutputStream for binary, character streams = Reader/Writer for text, and can pick the right one.

for a middle

Explains that character streams apply a charset to decode/encode, and that misusing byte streams on text causes garbled multi-byte characters.

for a senior

Articulates the encoding boundary precisely, knows the parallel class taxonomy, and always specifies the charset rather than the platform default.

for a principal

Frames the two hierarchies as an encoding-boundary design, can reason about UTF-16 surrogate pairs, supplementary characters, and where NIO (Charset/CharsetDecoder, channels) supersedes java.io for performance or control.

## The problem these solve A computer file or network connection is fundamentally a sequence of **bytes** — a byte is 8 bits, a number from 0 to 255. Some data is naturally binary (a JPEG image, an MP3, a compiled program): the bytes *are* the data. Other data is **text** — human-readable characters like `A`, `é`, or `中`. But text must also be stored as bytes, so there has to be a rule mapping characters to bytes and back. That rule is a **character encoding** (also called a **charset**), e.g. UTF-8, ISO-8859-1, or UTF-16. Java's `java.io` package gives you **two parallel stream hierarchies** to handle these two cases. ## Byte streams: raw bytes - **Roots:** `InputStream` (reading) and `OutputStream` (writing). - They deal in **bytes** — `int read()` returns one byte (0–255, or -1 at end of stream); `read(byte[])` fills a buffer. - They make **no interpretation** of the data. A byte is a byte. - **Use for:** images, audio, video, PDFs, ZIP files, serialized objects, encrypted blobs, or any time you want to copy data exactly without changing it. - **Example subclasses:** `FileInputStream`, `FileOutputStream`, `BufferedInputStream`, `ByteArrayInputStream`. A **stream** here just means "a sequence you read or write one piece at a time, in order" — not the Java 8 `java.util.stream.Stream`, which is unrelated. ## Character streams: text characters - **Roots:** `Reader` (reading) and `Writer` (writing). - They deal in **characters** — a Java `char` is 16 bits (a UTF-16 code unit). `int read()` returns one char value (or -1 at end). - They **automatically apply a charset** to convert between the underlying bytes and `char`s. So if a file is UTF-8 and you read it with a `Reader`, multi-byte sequences are decoded into the correct characters. - **Use for:** any text — config files, CSVs, logs, source code, JSON — especially with non-ASCII content. - **Example subclasses:** `FileReader`, `FileWriter`, `BufferedReader`, `StringReader`. ## Why two families instead of one If you only had byte streams and read a UTF-8 file byte-by-byte, you would mis-handle any character that takes more than one byte: `é` is two bytes in UTF-8, `中` is three. Treating each byte as a character splits them into garbage. Character streams exist precisely to do correct **decoding** (bytes → chars on read) and **encoding** (chars → bytes on write). ## The naming pattern The two hierarchies mirror each other, which makes them easy to learn: | Byte stream | Character stream | |---|---| | `InputStream` | `Reader` | | `OutputStream` | `Writer` | | `FileInputStream` | `FileReader` | | `BufferedInputStream` | `BufferedReader` | | `ByteArrayInputStream` | `CharArrayReader` | ## How to choose 1. **Is the data binary?** Use byte streams. 2. **Is the data text?** Use character streams, and **specify the charset explicitly** (e.g. `StandardCharsets.UTF_8`) rather than relying on the platform default — see the bridge question. Getting this wrong is one of the classic sources of "mojibake" (garbled text) bugs.

  • Why can't you just always use character streams to be safe?
    Character streams apply a charset, so running binary data through them corrupts it — bytes that are not valid in the charset get replaced or remapped. Binary data must go through byte streams to be preserved exactly.
  • How many bytes does a Java char represent?
    A char is a 16-bit UTF-16 code unit. Most characters fit in one char, but characters beyond the Basic Multilingual Plane (e.g. many emoji) need a surrogate pair of two chars.

saying these in an interview costs you the question

  • Saying 'just use FileInputStream for everything' — it corrupts multi-byte text
  • Thinking a Java char is always one byte (it is a 16-bit UTF-16 code unit)
  • Confusing java.io streams with java.util.stream.Stream
  • Believing character streams skip encoding — they always apply a charset

context

open as a page

Why must I/O streams be closed, and what happens if you forget to close one?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Streams hold operating-system resources like file handles or sockets. If you never close them, those resources leak and can run out, and buffered data may never be written. Always close them when you are done.

open as a page

What does wrapping an InputStream in a BufferedInputStream actually do, and why does it speed things up?

level: juniorimportance: must knowfreq 70%

basics

~20 s

BufferedInputStream wraps another stream and reads a big chunk of bytes at once into an in-memory array. Your reads come from that array instead of hitting the disk or network each time, so there are far fewer slow system calls.

open as a page

What do InputStreamReader and OutputStreamWriter do, and why are they called bridges?

level: middleimportance: must knowfreq 60%

basics

~20 s

They connect the byte world to the character world. InputStreamReader wraps a byte InputStream and decodes its bytes into characters using a charset. OutputStreamWriter wraps a byte OutputStream and encodes characters into bytes. They are the bridge between the two hierarchies.

open as a page

Why should you always specify a charset when converting between bytes and text in Java, and what goes wrong if you don't?

level: middleimportance: must knowfreq 70%

basics

~20 s

Bytes are not text until you choose a character encoding (a charset) to interpret them. If you don't specify one, older Java uses the machine's default, so the same code produces different or corrupted text on different machines. Always pass an explicit charset, normally UTF-8.

open as a page

What does flush() do on an output stream, and when do you need to call it explicitly?

level: middleimportance: must knowfreq 62%

basics

~20 s

flush() pushes data that is sitting in an in-memory buffer out to the underlying destination (file, socket, OS). You need it when you want buffered bytes to appear before you close the stream, such as on a long-lived socket.

open as a page

How does Java object serialization with ObjectOutputStream/ObjectInputStream work, and what is required to make a class serializable?

level: middleimportance: must knowfreq 65%

basics

~10 s

ObjectOutputStream.writeObject turns an object (and everything it references) into bytes; ObjectInputStream.readObject rebuilds it. The class must implement the Serializable marker interface, and every field that should be saved must itself be serializable.

open as a page

What are DataInputStream and DataOutputStream, and why would you use them instead of a plain InputStream/OutputStream?

level: juniorimportance: should knowfreq 45%

basics

~20 s

They are filter streams that read and write Java primitive types (int, double, boolean, etc.) and strings as raw bytes. You use them so you don't have to manually convert each primitive to and from bytes yourself.

open as a page

When would you choose DataOutputStream over ObjectOutputStream, and how do the two relate?

level: middleimportance: should knowfreq 42%

basics

~20 s

Use Data streams for a small, fixed set of primitives and strings where you control the exact byte layout. Use Object streams when you need to save whole objects and their references automatically. Both are filter streams built on byte streams.

open as a page

Why is java.io considered a textbook example of the Decorator design pattern? Explain how stream wrapping illustrates it.

level: middleimportance: should knowfreq 62%

basics

~20 s

Each stream class wraps another stream of the same type and adds one behavior (buffering, data conversion, etc.). Because the wrapper has the same interface as what it wraps, you can stack wrappers in any order to combine features without subclassing every combination.

open as a page

Concretely, what goes wrong when you read a UTF-8 text file through a byte stream and treat each byte as a character?

level: seniorimportance: should knowfreq 45%

basics

~10 s

In UTF-8 some characters take more than one byte. If you treat each byte as a character, those multi-byte characters get split into several wrong characters, so accented and non-Latin text comes out garbled.

open as a page

How does try-with-resources manage multiple streams, and what is a suppressed exception?

level: seniorimportance: should knowfreq 48%

basics

~20 s

You can declare several resources in one try-with-resources; Java closes them automatically in reverse order. If the body throws and a close() also throws, Java keeps the body's exception as the main one and attaches the close() exception as a 'suppressed' exception so neither is lost.

open as a page

How does object serialization handle shared references, cyclic references, and transient fields in an object graph?

level: seniorimportance: should knowfreq 50%

basics

~10 s

Serialization remembers each object it has already written, so an object referenced from two places is saved once and cycles don't loop forever. Fields marked transient are not saved and come back as defaults.

open as a page

Does the order in which you wrap decorator streams matter? Give an example involving buffering, compression, and flushing.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Yes. Each layer sees the bytes the layer below produces, so wrapping order decides whether you buffer raw or compressed data. Put the buffer next to the slow source (the file/socket) so it batches real I/O, and remember to flush the whole chain before relying on the output.

open as a page

What are the security and schema-evolution risks of native Java serialization, and how do you mitigate them?

level: principalimportance: should knowfreq 40%

basics

~20 s

Deserializing untrusted bytes can let an attacker run code on your machine, and changing a class can break old saved data. Avoid deserializing untrusted input, use input filters, and prefer schema-based formats like JSON or Protocol Buffers.

open as a page

How does the java.io stream design relate to the decorator pattern, and when would you move from java.io character streams to NIO?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

java.io streams are built with the decorator pattern: you wrap a base stream with layers that each add a feature, like buffering or character decoding. NIO is a newer, channel-and-buffer based API you reach for when you need better performance, non-blocking I/O, or fine-grained charset control.

open as a page

Design a write path that must survive a power loss. Where do flush(), close(), and fsync fit, and what guarantees does each give?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Writing data passes through layers: your JVM buffer, then the operating system's cache, then the physical disk. flush() moves data from the JVM to the OS, close() flushes and releases the handle, but only an OS sync (fsync) forces the OS to write to the disk so it survives a power loss.

open as a page

When would you reach for java.io decorator streams with buffering versus NIO channels and ByteBuffers? What does each model optimize?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The java.io decorator streams (with BufferedInputStream/Reader) are simple, blocking, byte/character-at-a-time APIs that are great for straightforward sequential file and text work. NIO channels with ByteBuffers are lower-level and support non-blocking, selectable, and bulk transfers, which scale better for many concurrent connections.

open as a page