A Ruby job calls File.readlines on a 5 GB access log and gets killed for memory; why, and how do you rewrite it?
answer
- one Array holding every line
- a String object per line
- File.foreach reads through a buffer
- no block: Enumerator, chain lazy
- don't collect the lines again
basics
~20 sFile.readlines reads the whole file into an Array with one String per line, so a 5 GB log needs more than 5 GB of heap. File.foreach, or each_line on an open File, yields one line at a time and keeps memory flat.
solid answer
~40 s`File.readlines` (like `File.read`) is eager: it reads the entire file and returns an `Array` holding one `String` per line, so memory is the file size **plus** per-object overhead for tens of millions of strings, and the process is killed long before it does any work. The fix is to stream: `File.foreach(path, chomp: true) { |line| ... }` reads through the IO buffer, yields each line, and closes the file when done, so only the current line and your aggregates stay live. Without a block, `File.foreach` returns an `Enumerator`, so you can write `File.foreach(path).lazy.select { ... }.first(10)` to stop reading early, or `each_slice(5_000)` to batch database inserts. Streaming only helps if the loop keeps aggregates rather than pushing every line into another array.
code
ruby · 19 lines# Eager: builds an Array of ~25 million Strings before the loop starts
# File.readlines("access.log").each { |line| ... }
# Streaming: one line at a time, file closed afterwards
status_counts = Hash.new(0)
File.foreach("access.log", chomp: true) do |line|
status_counts[line.split[8]] += 1
end
# Without a block: an Enumerator; lazy stops reading after 10 matches
first_errors = File.foreach("access.log")
.lazy
.select { it.include?(" 500 ") }
.first(10)
# Batching without loading the file
File.foreach("access.log").each_slice(5_000) do |batch|
puts batch.size
endgo deeper
Recall that File.readlines and File.read load everything, while File.foreach and each_line read one line at a time.
Explain the per-line String overhead, that File.foreach without a block returns an Enumerator, and how lazy and each_slice keep a chain streaming.
Diagnose the OOM-killed job, rewrite it to stream with bounded aggregates, cap pathological line lengths, and verify flat memory over a full run.
Set the rule that any input without a guaranteed size bound is streamed, and review data-processing code for hidden accumulation rather than just the read call.
## Eager versus streaming reads Ruby offers two families of whole-file reads, and the difference is **when the data enters memory**. **Eager** — the entire file is read before your code sees any of it: - `File.read(path)` returns one `String` holding every byte. - `File.readlines(path)` returns an `Array` with one `String` per line. - `file.readlines` on an open `File` does the same from the current position. **Streaming** — one line is read at a time through the IO object's internal buffer: - `File.foreach(path) { |line| ... }` opens, yields each line, and closes. - `file.each_line { |line| ... }` does the same on a `File` you opened, returning the file. - `file.gets` in a `while` loop returns one line per call and `nil` at end of file. ## Why readlines on 5 GB fails A 5 GB access log with ~200-byte lines has about 25 million lines. `File.readlines` must hold all of them at once, and: 1. Every line becomes a separate `String` object with its own header and heap buffer, so the total is well above 5 GB. 2. The `Array` holding 25 million references needs its own large contiguous allocation. 3. Nothing can be freed until the job drops the array, so the garbage collector cannot help. In a container with a memory limit, the kernel's out-of-memory killer ends the process — often with no Ruby exception at all, just a killed worker. ## The streaming rewrite ```ruby status_counts = Hash.new(0) File.foreach("access.log", chomp: true) do |line| status_counts[line.split[8]] += 1 # status field in combined log format end ``` Memory now holds one line, the buffer, and a `Hash` with a few dozen keys, whatever the file size. `File.foreach` wraps the loop in an ensure step, so the file is closed even if a line blows up the parser. ## Streaming from a handle you opened Sometimes you need the handle itself — to skip a header, or to resume from where the previous run stopped. The same rules apply with `IO#each_line`: ```ruby File.open("access.log") do |f| f.seek(last_offset) f.each_line(chomp: true) { |line| handle(line) } save_offset(f.pos) end ``` `each_line` streams from the current position and returns the file; after the loop, `f.pos` is the byte offset to resume from on the next run, and the block form of `File.open` closes the handle however the loop ends. Writing `f.readlines.each` in the same place would make it eager again. ## Enumerators make streaming composable Called **without a block**, `File.foreach` returns an `Enumerator`, and the file is opened only when something iterates it. That lets you use `Enumerable` methods without loading the file: | Goal | Streaming form | |---|---| | First 10 server errors | `File.foreach(path).lazy.select { it.include?(" 500 ") }.first(10)` | | Batched inserts | `File.foreach(path).each_slice(5_000) { \|batch\| insert(batch) }` | | Line numbers | `File.foreach(path).with_index(1) { \|line, n\| ... }` | | Count matches | `File.foreach(path).count { it.include?("POST") }` | `lazy` matters for chains: a plain `select` on the enumerator would build a full array of matches before `first` ran. With `lazy`, `first(10)` stops the read after the tenth match. ## Traps that bring the memory back - **Calling `lazy` on `readlines`.** `File.readlines(path).lazy` is too late — the array already exists. - **Accumulating lines.** `errors << line` for every error line recreates the problem on a smaller scale; keep counts, keys or a bounded sample. - **Huge lines.** A "line" is everything up to the separator. A file with no newlines is one 5 GB line. Pass a limit — `File.foreach(path, 1_048_576)` — so each yielded chunk stays bounded, or read fixed-size blocks with `IO#read(bytes)`. - **Keeping the whole file for a second pass.** Stream twice rather than caching everything, or record only the offsets you need. ## When eager is fine `File.read` and `File.readlines` are the right choice for small, bounded files — a config file, a template, a fixture. The judgement is about the **upper bound** on the input: if nobody can promise one, stream it.
- Does File.read have the same problem, and when is it acceptable?Yes: `File.read` returns the whole file as one `String`, so a 5 GB file needs a 5 GB string. It is fine for small inputs with a known upper bound, such as a config file or template. For anything whose size you cannot bound, stream with `File.foreach` or `each_line`.
- The job now uses File.foreach but memory still grows steadily; what do you look for?Something is retaining lines: an array that collects every match, a `Hash` keyed by full lines, a cache or memo, or objects built per line and stored. Keep aggregates, keys or a bounded sample instead of the lines themselves, and watch the process's resident memory over the run to confirm it stays flat.
- What happens with File.foreach if the file contains no newline characters at all?A line is everything up to the separator, so the whole file arrives as one giant `String` and streaming buys nothing. Pass a limit, as in `File.foreach(path, 1_048_576)`, so each yielded chunk is capped, or read fixed-size blocks with `IO#read(bytes)` in a loop.
saying these in an interview costs you the question
- File.readlines is lazy and only reads lines as you iterate the array.
- File.foreach loads the file into memory first and then yields lines.
- Calling lazy on File.readlines(path) fixes the memory problem.
- A 5 GB file needs about 5 GB of memory with readlines, no more.
- Core Ruby cannot stream a file; a gem is needed for line iteration.