How do you process a large text file line by line in a memory-efficient way using Files.lines, and what must you be careful about?
answer
- Files.lines -> lazy Stream<String>, constant memory
- MUST close: try-with-resources (open file handle)
- UTF-8 default; IOException -> UncheckedIOException
- vs readAllLines (loads all) and BufferedReader loop
- Windows file-lock if not closed
basics
~20 sUse Files.lines(path), which gives a Stream<String> that reads lines lazily one at a time instead of loading the whole file. Wrap it in try-with-resources so the underlying file gets closed, because a stream from a file holds an open file handle.
solid answer
~40 sFiles.lines(Path) returns a Stream<String> backed by a BufferedReader, so lines are read lazily as you consume them — memory stays roughly constant regardless of file size, unlike readAllLines which materializes everything. You can then filter, map, count, etc. The crucial caveat: this stream wraps an open file, so it must be closed; Stream implements AutoCloseable, so put it in try-with-resources, otherwise you leak a file handle (and on Windows the file may stay locked). Also note it decodes as UTF-8 by default, and any IOException during iteration is wrapped in an UncheckedIOException. For terminal aggregations like counting matching lines it is concise and fast; for simple side-effect loops a plain BufferedReader is equally fine.
code
java · 8 linesimport java.nio.file.*;
import java.util.stream.Stream;
// Count ERROR lines in a huge log without loading it all
try (Stream<String> lines = Files.lines(Path.of("app.log"))) {
long errors = lines.filter(l -> l.contains("ERROR")).count();
System.out.println(errors);
} // <- stream closed here, releasing the file handlego deeper
Knows Files.lines streams a file line by line instead of loading it all.
Uses it inside try-with-resources and explains the lazy, constant-memory behavior vs readAllLines.
Explains the file-handle leak / Windows lock if unclosed, the UncheckedIOException wrapping, and when a BufferedReader loop is the better tool.
Guides when streaming vs whole-file is appropriate, the parallel-stream non-benefit on I/O, and codifies resource-closing discipline in reviews/lint rules.
## The goal You have a big file — say a 5 GB log — and you want to, e.g., count lines containing "ERROR". You must **not** load it all into memory. The answer is *lazy streaming*: read and process one line at a time. ## What Files.lines does ```java try (Stream<String> lines = Files.lines(Path.of("app.log"))) { long errors = lines.filter(l -> l.contains("ERROR")).count(); } ``` `Files.lines(Path)` returns a **`Stream<String>`**. A *Stream* in Java is a lazy pipeline of elements: nothing is read until a *terminal operation* (like `count()`, `forEach()`, `collect()`) pulls elements through. Internally it is backed by a `BufferedReader`, reading one line per element on demand. So at any moment only one line (plus an 8 KB buffer) is in memory — constant memory regardless of file size. Contrast `Files.readAllLines`, which builds a `List<String>` of *every* line up front. ## The non-obvious danger: closing the file A normal collection stream (e.g. `list.stream()`) needs no closing. But `Files.lines` holds an **open file handle** the whole time the stream is alive. `Stream` implements `AutoCloseable`, and `Files.lines` is documented to require closing. If you forget, you **leak a file descriptor** — and on Windows the OS keeps the file *locked*, so later deletes/renames fail. Therefore always use **try-with-resources**: ```java try (Stream<String> s = Files.lines(p)) { ... } // s.close() releases the handle ``` This is the single most common mistake with `Files.lines`. ## Encoding and exceptions It decodes with **UTF-8** by default (an overload accepts an explicit `Charset`). Because `Stream` methods can't throw checked exceptions, an `IOException` that occurs *while iterating* is wrapped and rethrown as an **`UncheckedIOException`** (a `RuntimeException`). Also, if the file contains bytes that aren't valid UTF-8, decoding fails the same way. ## When to use which - **`Files.lines`**: large files + a functional aggregation (filter/map/count/collect). Concise. - **`BufferedReader` + while loop**: large files + a plain imperative loop with side effects; equally memory-efficient, no stream overhead, and checked exceptions stay checked. - **`Files.readAllLines`**: only small files (you want random access to all lines). ## Parallelism caveat `Files.lines(...).parallel()` rarely helps for line processing because the source is sequential I/O; don't reach for it by default.
- Why must you close a Files.lines stream but not a list.stream()?Files.lines is backed by an open file; the stream holds the file handle until closed. A collection stream has no underlying resource.
- What exception type surfaces if an I/O error happens mid-iteration?UncheckedIOException, because Stream operations cannot throw checked exceptions; the original IOException is its cause.
saying these in an interview costs you the question
- Not closing the Files.lines stream (file-descriptor / lock leak)
- Using it on the assumption it loads everything like readAllLines
- Expecting checked IOException mid-stream instead of UncheckedIOException
- Reaching for .parallel() expecting a speedup on sequential I/O