A nightly Eloquent job walks two million product rows and updates each one; when do you choose chunk(), chunkById(), lazy() or cursor(), and why?
answer
- avoid all() and get() on millions
- offset paging skips rows when filter changes
- chunkById pages with where id > last
- cursor: one query, no eager loading
- lazy() defaults to 1000 per page
basics
~20 sUse chunkById() or lazyById() when the loop updates rows it filters on, since offset-based chunk() and lazy() then skip rows. cursor() runs one query and holds one model but cannot eager load, and the driver may still buffer the result.
solid answer
~40 sAll four avoid hydrating two million models at once. `chunk()` and `lazy()` page with limit and offset: `chunk()` hands each page to a callback, `lazy()` flattens the pages into one lazy stream. Offset paging breaks when the loop changes the filtered column, because each page shifts and rows are skipped, so for an update job I use `chunkById()` or `lazyById()`, which ask for keys greater than the last one seen. `cursor()` runs a single query and hydrates one model at a time, but it cannot eager load relations and the driver may still buffer the raw rows, so the docs point to `lazy()` for very large sets. For a job that updates every product, `chunkById(500, ...)` with per-page work is my default.
go deeper
Recall that all() and get() load everything, and that chunk(), lazy() and cursor() exist to process large tables in pieces.
Explain offset paging versus key paging, and what cursor() trades for its single query.
Pick the reader by write pattern and relations: chunkById for mutating loops, lazy for eager-loaded reads, set-based updates where no per-row logic is needed.
Weigh long-running PHP batch jobs against moving the work into set-based SQL or a streaming pipeline outside the request stack.
## The problem the four methods solve `Product::all()` or `->get()` loads **every** matching row and hydrates **every** model before your loop starts. For two million products that exhausts PHP's memory limit. Eloquent offers four ways to walk a large result without holding it all at once. They differ in how many queries they run, how much they hold, and whether they survive the loop changing the data it walks. | Method | Queries | Held in PHP at once | Returns | |---|---|---|---| | `chunk(500, fn)` | one per page, `limit/offset` | one page of models | `bool`, pages go to the callback | | `chunkById(500, fn)` | one per page, `where id > last` | one page of models | `bool` | | `lazy(500)` | one per page, `limit/offset` | one page of models | a lazy stream of models | | `cursor()` | **one** | one model, plus the driver's buffer | a lazy stream of models | `lazyById()` is to `lazy()` what `chunkById()` is to `chunk()`, and `each()` / `eachById()` are per-model callback forms of the chunked readers. `lazy()`, `lazyById()`, `each()` and `eachById()` default to 1000 rows per page; `chunk()` takes the size as a required argument. ## Offset paging and the skipped-rows bug `chunk()` and `lazy()` page with `limit` and `offset`, ordering by the primary key when the query has no order. That is correct as long as the set being paged does not change. It breaks when the loop modifies the column it filters on: 1. The job selects `where('needs_reindex', true)` and sets `needs_reindex = false` on each product. 2. Page 1 takes rows 1-500 and clears their flag. 3. Page 2 asks for offset 500 of a set that is now 500 rows smaller, so it skips the next 500 rows entirely. `chunkById()` and `lazyById()` avoid this by remembering the last key they saw and asking for `where id > last` on the next page, so a shrinking set cannot shift the window. Because they add their own `where`, group your own `orWhere` conditions in a closure so the key condition is not swallowed by an `or`. ## What `cursor()` trades `cursor()` runs **one** query and hydrates one model at a time as you iterate, so the PHP-side model count stays at one. Two costs come with it: - It **cannot eager load** relationships, because there is never a set of parents to load them for; touching a relation inside the loop runs a query per row. - The PDO driver may still buffer the raw result set; the Laravel docs warn that with a very large result this eventually runs out of memory and recommend `lazy()` instead. It also holds one connection and one open result for the whole walk, which matters for a long nightly job. ## Choosing for a nightly electronics job - **Updating each product as you go**, especially a column in the filter: `chunkById()` or `lazyById()`. - **Read-only export or report where the set is stable**: `lazy()` for a flat loop with eager loading per page, or `chunk()` when you want the page as a unit, for example to send a batch to a search index. - **A moderate read-only walk with no relations**: `cursor()` is simplest and issues a single query. - **A single set-based change**: none of them; a query-level `update()` is one statement. ## Memory beyond the reader Even the right reader leaks if the loop holds on to models: appending every product to an array, or keeping a growing log of changes, defeats paging. Keep per-page work self-contained and let each page go out of scope. ## Variants worth naming - `chunkByIdDesc()` and `lazyByIdDesc()` walk the key downwards, useful when the newest products matter most and the job may be stopped early. - `chunkById()` and `lazyById()` accept a `column` argument, and an `alias` for joined queries, when the paging key is not the model's primary key or must be qualified. - `each(fn, 1000)` and `eachById(fn, 1000)` hide the page entirely and call the closure per model; returning `false` from it stops the walk. - Every chunked reader respects an existing `limit` and `offset` on the query, so `->limit(10000)->chunkById(...)` processes at most ten thousand rows. ```php Product::where('needs_reindex', true) ->chunkById(500, function ($products) { foreach ($products as $product) { $product->update(['needs_reindex' => false]); } }); ```
- Why must you group `orWhere` conditions in a closure before calling `chunkById()`?`chunkById()` appends its own `where id > ?` condition. Written ungrouped, `where('a', 1)->orWhere('b', 1)` becomes `a = 1 or b = 1 and id > ?`, so the key condition binds only to the second branch and pages repeat rows. Wrapping your conditions in `where(function ($q) { ... })` keeps the key condition applied to the whole filter.
- Why does `cursor()` not help when the loop reads `$product->category->name`?`cursor()` hydrates one model at a time, so there is never a batch of parents to eager load relations for, and it cannot eager load them. Each access lazy-loads the category with its own query. `lazy()` pages through the table, so relations can be eager loaded once per page.
saying these in an interview costs you the question
- chunk() is safe even when the loop updates the filtered column
- cursor() never runs out of memory because it holds one model
- cursor() can eager load relations like get() can
- lazy() runs a single query like cursor()
- Product::all() is fine for a nightly job because it runs off-peak