skip to content

How do you process a large text file line by line in a memory-efficient way using Files.lines, and what must you be careful about?

level: middleimportance: should knowfreq 55%

answer

  1. Files.lines -> lazy Stream<String>, constant memory
  2. MUST close: try-with-resources (open file handle)
  3. UTF-8 default; IOException -> UncheckedIOException
  4. vs readAllLines (loads all) and BufferedReader loop
  5. Windows file-lock if not closed

basics

~20 s

Use Files.lines(path), which gives a Stream<String> that reads lines lazily one at a time instead of loading the whole file. Wrap it in try-with-resources so the underlying file gets closed, because a stream from a file holds an open file handle.

solid answer

~40 s

Files.lines(Path) returns a Stream<String> backed by a BufferedReader, so lines are read lazily as you consume them — memory stays roughly constant regardless of file size, unlike readAllLines which materializes everything. You can then filter, map, count, etc. The crucial caveat: this stream wraps an open file, so it must be closed; Stream implements AutoCloseable, so put it in try-with-resources, otherwise you leak a file handle (and on Windows the file may stay locked). Also note it decodes as UTF-8 by default, and any IOException during iteration is wrapped in an UncheckedIOException. For terminal aggregations like counting matching lines it is concise and fast; for simple side-effect loops a plain BufferedReader is equally fine.

code

java · 8 lines
java
import java.nio.file.*;
import java.util.stream.Stream;

// Count ERROR lines in a huge log without loading it all
try (Stream<String> lines = Files.lines(Path.of("app.log"))) {
    long errors = lines.filter(l -> l.contains("ERROR")).count();
    System.out.println(errors);
} // <- stream closed here, releasing the file handle

go deeper

for a junior

Knows Files.lines streams a file line by line instead of loading it all.

for a middle

Uses it inside try-with-resources and explains the lazy, constant-memory behavior vs readAllLines.

for a senior

Explains the file-handle leak / Windows lock if unclosed, the UncheckedIOException wrapping, and when a BufferedReader loop is the better tool.

for a principal

Guides when streaming vs whole-file is appropriate, the parallel-stream non-benefit on I/O, and codifies resource-closing discipline in reviews/lint rules.

## The goal You have a big file — say a 5 GB log — and you want to, e.g., count lines containing "ERROR". You must **not** load it all into memory. The answer is *lazy streaming*: read and process one line at a time. ## What Files.lines does ```java try (Stream<String> lines = Files.lines(Path.of("app.log"))) { long errors = lines.filter(l -> l.contains("ERROR")).count(); } ``` `Files.lines(Path)` returns a **`Stream<String>`**. A *Stream* in Java is a lazy pipeline of elements: nothing is read until a *terminal operation* (like `count()`, `forEach()`, `collect()`) pulls elements through. Internally it is backed by a `BufferedReader`, reading one line per element on demand. So at any moment only one line (plus an 8 KB buffer) is in memory — constant memory regardless of file size. Contrast `Files.readAllLines`, which builds a `List<String>` of *every* line up front. ## The non-obvious danger: closing the file A normal collection stream (e.g. `list.stream()`) needs no closing. But `Files.lines` holds an **open file handle** the whole time the stream is alive. `Stream` implements `AutoCloseable`, and `Files.lines` is documented to require closing. If you forget, you **leak a file descriptor** — and on Windows the OS keeps the file *locked*, so later deletes/renames fail. Therefore always use **try-with-resources**: ```java try (Stream<String> s = Files.lines(p)) { ... } // s.close() releases the handle ``` This is the single most common mistake with `Files.lines`. ## Encoding and exceptions It decodes with **UTF-8** by default (an overload accepts an explicit `Charset`). Because `Stream` methods can't throw checked exceptions, an `IOException` that occurs *while iterating* is wrapped and rethrown as an **`UncheckedIOException`** (a `RuntimeException`). Also, if the file contains bytes that aren't valid UTF-8, decoding fails the same way. ## When to use which - **`Files.lines`**: large files + a functional aggregation (filter/map/count/collect). Concise. - **`BufferedReader` + while loop**: large files + a plain imperative loop with side effects; equally memory-efficient, no stream overhead, and checked exceptions stay checked. - **`Files.readAllLines`**: only small files (you want random access to all lines). ## Parallelism caveat `Files.lines(...).parallel()` rarely helps for line processing because the source is sequential I/O; don't reach for it by default.

  • Why must you close a Files.lines stream but not a list.stream()?
    Files.lines is backed by an open file; the stream holds the file handle until closed. A collection stream has no underlying resource.
  • What exception type surfaces if an I/O error happens mid-iteration?
    UncheckedIOException, because Stream operations cannot throw checked exceptions; the original IOException is its cause.

saying these in an interview costs you the question

  • Not closing the Files.lines stream (file-descriptor / lock leak)
  • Using it on the assumption it loads everything like readAllLines
  • Expecting checked IOException mid-stream instead of UncheckedIOException
  • Reaching for .parallel() expecting a speedup on sequential I/O

context