skip to content

A teammate uses File.readLines() to process a 5 GB log file and the process runs out of memory. How would you fix it with the Kotlin stdlib, and why does that fix work?

level: middleimportance: must knowfreq 65%

answer

  1. readLines = whole file in a List (OOM risk)
  2. useLines gives a lazy Sequence, auto-closed
  3. Don't leak the Sequence outside the lambda
  4. forEachLine = per-line side-effect, auto-closed
  5. Streaming = O(one line) memory

basics

~20 s

readLines loads every line into memory at once. Replace it with useLines or forEachLine, which read the file line by line and only keep one line in memory at a time, so a huge file no longer blows up the heap.

solid answer

~50 s

`readLines()` returns a `List<String>` containing the whole file, so a 5 GB file needs ~5 GB of heap — hence the OOM. The streaming alternatives keep memory bounded. `File.useLines(charset) { lines: Sequence<String> -> ... }` gives you a lazy `Sequence<String>`; lines are read on demand and the reader is closed automatically when the lambda returns. Because it's a `Sequence`, operators like `filter`, `map`, `count` run lazily without materializing the file. `File.forEachLine { line -> ... }` is the simplest form: it invokes the lambda once per line and closes the stream itself. The catch with `useLines`: the `Sequence` is only valid inside the lambda — you must consume it there (e.g. return a count or terminal result), not leak it out, because the underlying reader is closed on exit. For both, only one line plus buffer is resident, so memory is O(longest line), not O(file).

code

kotlin · 10 lines
kotlin
import java.io.File

// BAD: materializes 5 GB
// val errors = File("app.log").readLines().count { it.contains("ERROR") }

// GOOD: bounded memory, reader auto-closed
val errors = File("app.log").useLines { lines ->
    lines.count { it.contains("ERROR") }
}
println("errors=$errors")

go deeper

for a junior

Recognizes that loading a 5 GB file into a List is the problem and that a line-by-line approach helps.

for a middle

Names useLines/forEachLine, explains the lazy Sequence and automatic close, picks the right one.

for a senior

Articulates the O(one line) memory model, the leaked-Sequence pitfall, and when to drop to bufferedReader().use for control.

for a principal

Weighs throughput/back-pressure, charset cost, and whether streaming line-by-line vs. chunked binary reads fits the workload at scale.

## Why readLines OOMs `File.readLines()` returns a `List<String>` — an *eager* collection holding **every** line simultaneously. For a 5 GB file that means roughly 5 GB of `String` objects on the heap at once, which exceeds typical JVM heap limits and throws `OutOfMemoryError`. ## The streaming fix The stdlib offers two memory-bounded helpers that read **one line at a time**: ### useLines — lazy Sequence ```kotlin import java.io.File val errorCount = File("app.log").useLines { lines -> lines.count { it.contains("ERROR") } } ``` - `useLines(charset = Charsets.UTF_8, block: (Sequence<String>) -> T): T` opens a `BufferedReader`, hands your lambda a **lazy** `Sequence<String>`, and **closes the reader automatically** when `block` returns (it's built on `use`). - A `Sequence` is lazy: intermediate operators (`filter`, `map`) don't build new lists; lines flow through one at a time until a terminal operator (`count`, `sumOf`, `forEach`, `toList`). - **Critical rule:** the `Sequence` is only valid **inside** the lambda. Returning it (e.g. `useLines { it }`) gives you a Sequence whose underlying reader is already closed — iterating it later throws. Always finish your work and return a *terminal* value. ### forEachLine — simplest per-line callback ```kotlin File("app.log").forEachLine { line -> if (line.contains("ERROR")) println(line) } ``` - `forEachLine(charset, action: (String) -> Unit)` calls `action` once per line and closes the stream for you. No Sequence to mishandle. Use it for pure side-effects. ### bufferedReader().use — full control ```kotlin File("app.log").bufferedReader().use { reader -> reader.lineSequence().filter { it.isNotBlank() }.forEach(::println) } ``` `bufferedReader()` returns a `BufferedReader`; `.use { }` guarantees `close()` even on exception. `lineSequence()` lazily yields lines. ## Memory model comparison | API | Memory | Stream closed by | Result type | |-----|--------|------------------|-------------| | `readLines()` | O(file) | itself (eager) | `List<String>` | | `useLines { }` | O(one line) | itself, on lambda exit | your terminal `T` | | `forEachLine { }` | O(one line) | itself | `Unit` | | `bufferedReader().use { }` | O(one line) | `use` | your `T` | The streaming options keep only the current line (plus an 8 KB-ish buffer) resident, so memory is O(longest line) regardless of file size.

  • Why can't you return the Sequence from useLines and iterate it later?
    useLines closes the underlying reader when the lambda returns. The leaked Sequence is backed by a closed stream, so iterating it later throws an exception.
  • When would you pick forEachLine over useLines?
    When you only need a side-effect per line (printing, writing elsewhere) and don't need Sequence operators or a computed return value.

readLines pours the whole reservoir into a bucket; useLines runs water through a pipe one cup at a time so the bucket never overflows.

saying these in an interview costs you the question

  • Not knowing readLines holds the whole file in memory
  • Returning the useLines Sequence out of the lambda and iterating it afterward
  • Thinking useLines is eager like toList
  • Manually closing the reader inside useLines (it's already managed)
  • Confusing useLines (lazy Sequence) with readLines (eager List)

context