skip to content

In PHP, how do you process a multi-gigabyte log file line by line without hitting memory_limit, and which loop mistakes still break it?

level: seniorimportance: must knowfreq 55%

answer

  1. one line in memory at a time
  2. file() builds an array of every line
  3. fgets() returns false when nothing is left
  4. while (!feof()) runs one extra time
  5. SplFileObject DROP_NEW_LINE, READ_AHEAD, SKIP_EMPTY

basics

~20 s

Open the file with fopen() and loop while fgets() does not return false, or iterate an SplFileObject, so only one line is in memory. file() and file_get_contents() load everything, and collecting results in an array grows memory again.

solid answer

~40 s

`file()` and `file_get_contents()` load the whole file, so a multi-gigabyte log exceeds `memory_limit` (128M by default) long before it is parsed. The streaming version is `fopen($path, 'r')` and `while (($line = fgets($h)) !== false)`, which holds one line at a time; `SplFileObject` gives the same thing as a `foreach` with the flags `DROP_NEW_LINE | READ_AHEAD | SKIP_EMPTY`. The common mistakes are `while (!feof($h))`, which runs one extra iteration with `false` because `feof()` turns true only after a read has hit the end; appending every parsed row to an array, which rebuilds the memory problem; and a file with no newlines, where one `fgets()` call without a length reads everything. I would check the fix with `memory_get_peak_usage()`, not by raising `memory_limit` to `-1`.

go deeper

for a junior

Recall that fgets() in a loop holds one line at a time, returns false at the end, and that file() and file_get_contents() load the whole file.

for a middle

Explain why while (!feof()) runs one extra iteration, why the check is !== false, and how SplFileObject's flags change what each iteration yields.

for a senior

Show you measure peak memory, guard against huge lines with a length or fread() chunks, and keep aggregates bounded so the import survives a tenfold bigger file.

for a principal

Frame it as a pipeline decision: which imports must stream end to end, where results are flushed, and why memory_limit stays a safety net rather than a tuning knob.

## Why the obvious code fails A PHP script runs under **`memory_limit`**, 128M in the built-in default and in both shipped php.ini files. Anything that turns the file into PHP values at once breaks on a large log: - `file_get_contents($path)` makes one string as big as the file. - `file($path)` first reads the whole file, then builds an **array with one string per line**; each string carries its own header and each element an array slot, so the footprint can be much larger than the file itself. - `explode("\n", file_get_contents($path))` briefly holds both the string and the array. When the limit is crossed the script stops with the fatal error *Allowed memory size of N bytes exhausted*. Raising `memory_limit` to `-1` just moves the failure to the machine's RAM. ## The streaming loop ```php <?php declare(strict_types=1); $h = fopen('/var/log/app/access.log', 'r'); if ($h === false) { throw new RuntimeException('cannot open log'); } $errors = 0; while (($line = fgets($h)) !== false) { if (str_contains($line, ' 500 ')) { $errors++; } } fclose($h); echo $errors, PHP_EOL, memory_get_peak_usage(), PHP_EOL; ``` Key points: 1. **`fgets($h)` returns the next line including its newline, or `false`** when there is nothing more to read (or on error). 2. The condition compares with `!== false`, because a line containing only `"0"` is falsy. 3. Only the current line and the running aggregate (`$errors`) are kept, so peak memory stays roughly constant however large the file is. ## Mistakes that still break it | Mistake | What goes wrong | |---|---| | `while (!feof($h)) { $line = fgets($h); … }` | `feof()` becomes true only after a read has hit the end. After the last line it is still false, so the body runs once more with `$line === false`. | | Collecting every parsed row in an array | memory grows with the file again; aggregate as you go or write results out in batches | | A file with one enormous line (minified JSON, a binary blob) | `fgets()` without a length reads until a newline, so it reads everything; pass a length (`fgets($h, 8192)` returns at most 8191 bytes) or use `fread()` in chunks | | Not checking `fopen()` | on failure `$h` is `false`, and since PHP 8.0 passing it to `fgets()` or `feof()` throws a `TypeError` instead of warning and looping | | Loading the file with `file()` "because it is convenient" | the whole file becomes an array before the first line is processed | ## SplFileObject: the same stream as an object `SplFileObject` wraps a file handle and is iterable, one line per iteration: ```php <?php declare(strict_types=1); $file = new SplFileObject('/var/log/app/access.log', 'r'); $file->setFlags(SplFileObject::DROP_NEW_LINE | SplFileObject::READ_AHEAD | SplFileObject::SKIP_EMPTY); foreach ($file as $lineNo => $line) { // $line has no trailing newline; empty lines are skipped } ``` - The constructor throws a `RuntimeException` if the file cannot be opened, instead of returning `false`. - `DROP_NEW_LINE` strips the line ending; `SKIP_EMPTY` skips empty lines and, per the manual, needs `READ_AHEAD` to work as expected. - `$file->seek($n)` moves to line `$n`, but it has to read through the lines before it; it is not a byte offset. ## Related tools on the same handle - `fread($h, $n)` reads fixed-size chunks when line structure does not matter. - `fseek($h, $offset, SEEK_END)` and `ftell($h)` jump near the end of a log to process only its tail. `fseek()` returns `0` on success and `-1` on failure. - Wrapping the loop in a function that `yield`s lines turns it into a reusable lazy iterator; that technique belongs with generators, but the memory behaviour comes from `fgets()` underneath. ## Where the time goes Memory is fixed by the loop; time is usually spent in what you do with each line. A `preg_match()` per line on a multi-gigabyte log dominates the cost, so cheap pre-filters such as `str_contains()` before the regex pay off. For progress reporting on long runs, `ftell($h)` gives the byte offset reached, which you can compare against the file size to print a percentage without counting lines first. ## How to prove it works Measure instead of guessing: log `memory_get_peak_usage()` at the end of a run on a small file and on a large one. With a streaming loop the two numbers are close; with `file()` the second grows with the file.

  • How do you process only the last 10 MB of a large, growing log in PHP?
    Open it with `'r'`, call `fseek($h, -10 * 1024 * 1024, SEEK_END)`, discard the first `fgets()` result because it is probably a partial line, then loop on `fgets()` as usual. `fseek()` returns `-1` if the offset is before the start, so fall back to `rewind()` for small files. `ftell()` tells you the offset you stopped at, which you can store to resume the next run.
  • Why can peak memory still climb in a correct fgets() loop?
    Because something else in the loop keeps data: an array of parsed rows, a growing string, a cache keyed by user, or an ORM that remembers every entity it saw. The loop itself holds one line; anything you accumulate is on you. Aggregate into counters, flush batches to a database or output file, and check `memory_get_peak_usage()` on inputs of two sizes.
  • What does SplFileObject::SKIP_EMPTY need to work as expected?
    The manual says it requires `READ_AHEAD`. In practice you combine `DROP_NEW_LINE | READ_AHEAD | SKIP_EMPTY`: without dropping the newline, a blank line is still `"\n"` rather than empty, and without read-ahead the iterator does not skip it reliably.

Reading with fgets() is like reading a long scroll through a window one line wide: the window never grows, however long the scroll is. file() unrolls the whole scroll across the floor before you read the first line.

saying these in an interview costs you the question

  • file() is fine for big files because it returns lines one at a time.
  • while (!feof($h)) is the correct way to loop over fgets().
  • Raising memory_limit to -1 is a proper fix for large imports.
  • fgets() with no length stops at 1024 bytes, so huge lines are safe.
  • SplFileObject::seek() jumps straight to a byte offset.