skip to content

How would you use Laravel's LazyCollection::make() to count visits per country in a 4 GB access log without exhausting memory?

level: middleimportance: should knowfreq 38%

answer

  1. a closure that yields, not a Generator
  2. fgets inside a while loop
  3. File::lines($path) does it for you
  4. countBy keeps one counter per group
  5. try/finally closes the handle

basics

~20 s

Pass LazyCollection::make() a closure that opens the file and yields one line at a time, then chain map() to extract the country and countBy() to tally; only one line plus one counter per country is held in memory.

solid answer

~30 s

Give `LazyCollection::make()` a **generator function**: open the file with `fopen`, loop with `fgets`, `yield` each line, and close the handle in a `finally`. Chain `->map(fn ($line) => explode(',', $line)[2] ?? null)` to extract the country, `->filter()` to drop blanks, then `->countBy()`. On a lazy collection `countBy()` is streaming: it keeps one integer per country and yields the totals after the last line, so memory is one line plus the counter array rather than 4 GB. `File::lines($path)` returns the same kind of line-by-line `LazyCollection` without writing the loop. Avoid `groupBy()` here — on a lazy collection it collects every line first.

code

php · 11 lines
php
<?php

use Illuminate\Support\Facades\File;

$visitsByCountry = File::lines(storage_path('logs/access.log'))
    ->map(fn (string $line) => explode(',', $line)[2] ?? null)
    ->filter()
    ->countBy()           // streaming: one counter per country
    ->sortDesc()          // sorts only the small totals array
    ->take(10)
    ->all();

go deeper

for a junior

Know that LazyCollection::make() takes a closure that yields lines, and that File::lines() returns a lazy collection over a file.

for a middle

Build the map, filter and countBy pipeline, explain why countBy stays small, and close the handle in a finally block.

for a senior

Spot the calls that silently materialise or re-read the file, and bound memory for writes with lazy chunk() batches.

for a principal

Judge when a streaming PHP job is right versus log-pipeline tooling or database-side aggregation, based on volume and how often the report runs.

## The problem A 4 GB access log will not fit in a PHP process with a typical `memory_limit`, and `file()` or `file_get_contents()` try to load all of it. Plain loops work, but you lose the collection vocabulary. Laravel's **`LazyCollection`** gives you both: a streaming source with `map`, `filter` and aggregation methods on top. ## Step 1: a generator function as the source `LazyCollection::make()` accepts a **closure that yields**. The closure is invoked each time the collection is enumerated, and each `yield` hands one line to the pipeline: ```php <?php use Illuminate\Support\LazyCollection; $lines = LazyCollection::make(function () { $handle = fopen(storage_path('logs/access.log'), 'r'); try { while (($line = fgets($handle)) !== false) { yield $line; } } finally { fclose($handle); } }); ``` Points interviewers look for: - pass the **function**, not the result of calling it — the constructor rejects a `Generator` object with an `InvalidArgumentException`; - read with **`fgets`** (one line) rather than `file()` (the whole file); - put `fclose` in a **`finally`** block: when a later step such as `take()` or `first()` stops early, the generator is abandoned mid-loop, and code after the loop never runs, while a `finally` still does when the generator is destroyed. Laravel also has a ready-made version: **`File::lines($path)`** on the `File` facade returns a `LazyCollection` that reads the file with `SplFileObject` and drops the newline characters. ## Step 2: a streaming pipeline Assume each line looks like `2026-09-29T10:00:00Z,203.0.113.5,DE,/pricing`: ```php <?php $visitsByCountry = $lines ->map(fn (string $line) => explode(',', trim($line))[2] ?? null) ->filter() ->countBy() ->sortDesc() ->all(); // ['US' => 812344, 'DE' => 301877, ...] ``` What each step holds in memory: | Step | Lazy behaviour | Memory it keeps | |---|---|---| | `map()` | yields one mapped value per line | one value | | `filter()` | yields only truthy values | one value | | `countBy()` | counts as it goes, yields the totals at the end | one integer per distinct country | | `sortDesc()` | collects its input, then sorts | the counts array — small here | | `all()` | enumerates into an array | the final counts | The whole run holds one line at a time plus a couple of hundred counters. Sorting is fine **after** `countBy()` because its input is tiny; sorting the raw lines would pull 4 GB into memory. ## Traps that undo the saving 1. **`groupBy('country')->map->count()`** — `groupBy()` on a lazy collection delegates to an eager `Collection`, so every line is materialised before grouping. Use `countBy()` or `reduce()` instead. 2. **Calling `->all()` or `->collect()` early** — converts the stream into an array on the spot. 3. **Enumerating twice** — `$lines->count()` followed by a second pipeline re-reads the whole file, because the generator function runs again. 4. **Accumulating in a closure** — a `map()` callback that appends every line to an outer array rebuilds the problem by hand. ## Batching writes When the goal is to store results instead of counting them, `chunk(1000)` on the lazy collection yields groups of 1,000 lines as small lazy collections, so each batch can go to one insert statement while memory stays bounded by the chunk size. ## Proving the memory claim Interviewers sometimes ask how you would show that the pipeline really streams. A quick check: 1. run the job against a small fixture and a large file; 2. log `memory_get_peak_usage(true)` at the end of each run; 3. compare — a streaming pipeline shows roughly the same peak for both, while an accidental `groupBy()` or `collect()` makes the peak grow with file size. Put a `tapEach()` counter in the chain to log progress every million lines without a second pass over the file. ## Why this is a Laravel question, not only a PHP one Generators are PHP's mechanism; the interview angle is knowing which **Laravel methods** keep streaming and which silently collect, that `make()` wants the function, and that `File::lines()` exists.

  • Why is groupBy('country') a poor choice on a Laravel LazyCollection over a huge log?
    `LazyCollection::groupBy()` passes through to the eager `Collection` implementation, which first collects every item into an array. For a 4 GB log that means holding every line in memory. `countBy()` keeps a running counter per group instead, and `reduce()` works when you need more than a count.
  • What happens to an fclose() placed after the loop in a LazyCollection source when you only take(10) lines?
    It never runs during the pipeline. `take(10)` stops pulling after ten lines, leaving the generator suspended inside the loop, and plain code after the loop is unreachable. PHP does run `finally` blocks when a suspended generator is destroyed, so put the `fclose()` in a `try`/`finally`.

saying these in an interview costs you the question

  • Reads the file with file() and then wraps the array in a LazyCollection.
  • Passes an already-called generator object to LazyCollection::make().
  • Uses groupBy() on the lazy stream to count per country.
  • Calls count() first to size a progress bar, re-reading the whole file.
  • Believes chunk() on a LazyCollection loads the file in chunks of bytes.