In PHP, how would you stream a 2 GB access log line by line with generators, and what pulls it all back into memory?
answer
- memory_limit defaults to 128M
- file() builds an array of every line
- fopen plus fgets inside a generator
- chain generator stages, aggregate as you go
- iterator_to_array and count() undo it
basics
~20 sYield each fgets() line from a generator that opened the log with fopen(), chain filter stages as further generators, and aggregate counters in the final foreach. file(), iterator_to_array() or storing every line reloads the whole file.
solid answer
~40 s`file()` or `file_get_contents()` on a 2 GB log builds the whole content in memory and dies against `memory_limit` (128M by default) with an *Allowed memory size exhausted* fatal error. Instead, a generator opens the file with `fopen()`, loops `while (($line = fgets($h)) !== false)` and yields each line, closing the handle in a `finally`. Filtering and parsing become further generators that take an `iterable` and yield what passes, so only one line is in flight at a time. The consumer aggregates, for example counting 5xx responses per path, rather than storing lines. Laziness is lost the moment something materialises the stream: `iterator_to_array()`, appending every line to an array, sorting, or `yield from file($path)`. `count()` on a `Generator` does not count at all; since PHP 8.0 it throws `TypeError`.
code
php · 26 lines<?php
declare(strict_types=1);
function lines(string $path): Generator
{
$h = fopen($path, 'r');
if ($h === false) {
throw new RuntimeException("Cannot open $path");
}
try {
while (($line = fgets($h)) !== false) {
yield rtrim($line, "\n");
}
} finally {
fclose($h);
}
}
function serverErrors(iterable $lines): Generator
{
foreach ($lines as $n => $line) {
if (preg_match('/" 5\d\d /', $line) === 1) {
yield $n => $line;
}
}
}go deeper
Recall that file() loads every line at once and that a generator around fopen() and fgets() yields one line at a time.
Explain how chained generator stages keep one line in flight, why the sink must aggregate, and which calls such as iterator_to_array() or count() break the pipeline.
Show the production details: closing the handle in finally for early exits, proving flat memory with memory_get_peak_usage(), and refusing to paper over the design by raising memory_limit.
Judge when a streaming PHP job is the right tool and when the question, such as an exact median, needs storage with its own sort or index instead.
## Why the obvious code fails The shortest way to read a file in PHP is `file($path)`, which returns an array with one element per line, or `file_get_contents($path)`, which returns one string. Both put the **entire** file into the process's memory. PHP caps that memory per script with the `memory_limit` directive, which defaults to `128M` both in the engine and in the shipped `php.ini-production` and `php.ini-development`. A 2 GB access log exceeds that by an order of magnitude, so the script stops with a fatal error of the form *Allowed memory size of N bytes exhausted (tried to allocate M bytes)*. Raising `memory_limit` only moves the cliff: the next, larger log fails again, and a web worker holding gigabytes starves its neighbours. ## The streaming shape The fix is to hold **one line at a time**. A generator makes that shape reusable: 1. **Source stage.** A generator opens the file with `fopen($path, 'r')`, loops on `fgets()`, which returns the next line or `false` at end of file, and yields each line. The `fclose()` goes in a `finally` block. 2. **Transform stages.** Each further step is a generator that accepts an `iterable`, loops over it with `foreach`, and yields what it keeps: lines matching a status code, parsed arrays, normalised paths. 3. **Sink.** The final `foreach` consumes the chain and **aggregates**: counters, sums, the current maximum, or rows written straight to another stream. When the sink pulls one value, each stage runs just far enough to produce it. At any moment the process holds one line per stage plus whatever the aggregate needs. Memory stays roughly constant whether the log is 2 MB or 2 GB, which `memory_get_peak_usage()` confirms. ## Why the finally block matters If the consumer stops early, with `break`, a `return`, or an exception, the source generator is left paused inside its loop. When the last reference to it goes away, PHP destroys it and runs any pending `finally` block, so the file handle is closed promptly. Without `try`/`finally`, the handle stays open until the generator object is destroyed, or until the request ends, which matters in long-running workers. ## What silently brings the file back into memory Laziness survives only while every stage passes values along one at a time. These patterns undo it: | Pattern | What happens | |---|---| | `file($path)` or `yield from file($path)` | the whole file is read into an array before the first line is yielded | | `iterator_to_array($lines)` | every yielded value is copied into one array | | appending each line to `$all[]` in the sink | the consumer rebuilds the array the generator avoided | | `sort()` or grouping that needs every line | sorting needs the full set; aggregate into counters instead | | `array_map()` / `array_filter()` on the stream | they accept only arrays, so someone first converts the generator | | `count($generator)` | since PHP 8.0, `TypeError`: `count()` accepts `Countable|array`, and `Generator` is neither | The rule of thumb: every step between the file and the sink must be a `foreach` that yields, and the sink must keep a **bounded** summary rather than the lines themselves. ## Designing the aggregate Most questions asked of an access log reduce to bounded state: - **Counts per status code** grow with the number of distinct codes, not the number of lines. - **Top N slow paths** needs a structure capped at N entries, updated per line. - **Error lines for later review** can be written to another file with `fwrite()` as they stream past, instead of being held in memory. When a question genuinely needs the whole data set, such as an exact median over every response time, that is a signal to use a tool with its own storage (a database or a sort on disk), not to raise `memory_limit`. ## Keys as a bonus Because a generator yields keys too, the source stage's automatic integer keys give each line its zero-based line number. A filter stage that re-yields `$n => $line` keeps the original line numbers attached, so the sink can report *line 18204377: 503 on /checkout* without having counted lines itself.
- The consumer breaks out of the loop after the first 100 error lines. Is the log file's handle leaked?Not if the source generator closes it in `finally`. When the abandoned generator's last reference goes away, PHP destroys it and runs the pending `finally` block, so `fclose()` executes. Without `try`/`finally`, the handle stays open until the generator object is destroyed or the request ends, which matters in long-running CLI workers.
- Why not simply raise memory_limit to 4G for this job?It moves the failure to the next, larger log and lets one process hold gigabytes that other workers on the host need. A streaming pipeline has flat memory whatever the file size, so it needs no tuning as logs grow; the limit stays a guard against genuine leaks.
saying these in an interview costs you the question
- file() reads lines lazily, so it is fine for multi-gigabyte logs.
- Raising memory_limit is the proper fix for reading huge files.
- count() on a generator returns how many values it will yield.
- iterator_to_array() keeps a generator's memory benefit.
- Wrapping file() in a generator with yield from makes it lazy.