With Laravel collections, how do pluck(), keyBy() and groupBy() differ when summarising survey answers by department?
answer
- one value, one item, many items
- pluck('score', 'respondent_id')
- keyBy: last duplicate wins
- groupBy: collection of collections
- modelKeys() returns a plain array
basics
~10 spluck() extracts one field per item, keyBy() indexes whole items by a field with one item per key, and groupBy() builds a collection of sub-collections, one per distinct value, so duplicates are kept.
solid answer
~30 s`pluck('score')` returns just that field from every item, re-indexed; `pluck('score', 'respondent_id')` keys those values by another field, and dot notation such as `department.name` reaches into nested data. `keyBy('respondent_id')` keeps each whole item under the field's value, one item per key — if two items share a key, the later one silently replaces the earlier. `groupBy('department')` never drops anything: each key maps to a collection of every matching item, which you then aggregate with `map(fn ($g) => $g->avg('score'))`. For a two-way split use `partition()`, and on an Eloquent collection `modelKeys()` gives a plain array of primary keys.
code
php · 12 lines<?php
$answers = collect([
['respondent_id' => 17, 'department' => 'Sales', 'score' => 4],
['respondent_id' => 23, 'department' => 'IT', 'score' => 2],
['respondent_id' => 31, 'department' => 'Sales', 'score' => 5],
]);
$answers->pluck('score'); // [4, 2, 5]
$answers->pluck('score', 'respondent_id'); // [17 => 4, 23 => 2, 31 => 5]
$answers->keyBy('department')->count(); // 2 - Sales row 17 was overwritten
$answers->groupBy('department')->map->count(); // ['Sales' => 2, 'IT' => 1]go deeper
Know that pluck() takes one field, keyBy() indexes whole items, and groupBy() returns a collection of collections.
Explain the duplicate behaviour of keyBy() and pluck() with a key, groupBy's multi-level and preserveKeys options, and how to aggregate each group.
Spot silent data loss from keyBy() on non-unique fields in code review, and judge when aggregation belongs in SQL instead of in PHP after loading every row.
Set guidance for when reports aggregate in the database versus in collections, based on data volume, reuse of loaded models and query cost.
## Three different shapes Take a collection of survey answers, each with `respondent_id`, `department` and `score`. The three methods answer three different questions: | Method | Question it answers | Result shape | Duplicates | |---|---|---|---| | `pluck('score')` | "just the scores" | `[4, 5, 2, …]`, keys `0..n-1` | all kept | | `pluck('score', 'respondent_id')` | "score per respondent" | `[17 => 4, 23 => 5, …]` | later key wins | | `keyBy('respondent_id')` | "look an answer up by respondent" | `[17 => answer, 23 => answer]` | later item wins | | `groupBy('department')` | "all answers per department" | `['Sales' => Collection, 'IT' => Collection]` | all kept | ## pluck() `pluck($value, $key = null)` delegates to `Arr::pluck`, which reads each item with `data_get`. That means: - it works on arrays **and** objects, including Eloquent models and their loaded relations; - **dot notation** walks nested data: `pluck('respondent.email')`; - without a key argument, values are appended, so the result is always a list; - with a key argument, the result is associative; a repeated key keeps the **last** value. On an Eloquent collection, `pluck()` returns a **base** `Illuminate\Support\Collection`, because the values are no longer models. ## keyBy() `keyBy($keyBy)` accepts a field name or a closure and stores the **whole item** under the computed key. It is the right tool for fast lookups — `$byRespondent->get(17)` instead of searching — but it assumes the key is **unique**. With a non-unique field such as `department`, every department ends up with only its last answer, and nothing warns you. Enum keys are converted to their value, and objects with `__toString` are cast to strings. ## groupBy() `groupBy($groupBy, $preserveKeys = false)` builds a collection whose values are **collections**. Useful details: 1. The key can be a field name, a dotted path or a closure; a closure may return an **array of keys** to put one item in several groups. 2. Passing an **array** of keys groups on several levels: `groupBy(['department', 'question_id'])`. 3. Inner groups are **re-indexed** unless you pass `true` as the second argument. 4. Boolean keys become `0`/`1` and a `null` key becomes the empty string. Because each group is a collection, aggregation chains straight on: ```php <?php $summary = $answers ->groupBy('department') ->map(fn ($group) => [ 'responses' => $group->count(), 'average' => round($group->avg('score'), 1), ]); // ['Sales' => ['responses' => 12, 'average' => 4.2], 'IT' => [...]] ``` ## partition() and reduce() for the rest - `partition(fn ($a) => $a['score'] >= 4)` returns a collection of **two** collections — passing and failing — that you destructure: `[$promoters, $others] = $answers->partition(...)`. - `reduce(fn ($carry, $answer) => $carry + $answer['score'], 0)` folds everything into one value; the callback receives the carry first, then the item, then the key. Without the second argument the initial carry is `null`. - `countBy('department')` is the shortcut when all you need is a count per value. ## modelKeys() on Eloquent collections `Illuminate\Database\Eloquent\Collection::modelKeys()` returns a **plain PHP array** of each model's primary key via `getKey()`. It differs from `pluck('id')` in two ways: it respects a custom primary key name, and it returns an array rather than a collection — convenient for `whereIn()` or `sync()`. ## PHP-side or database-side? Everything above runs **after** the rows are loaded into memory. For a few hundred answers that is fine and keeps the report logic readable. For hundreds of thousands, loading every row to call `groupBy('department')->map->avg('score')` costs memory and transfer that a `GROUP BY` in the query would avoid. A reasonable split: - aggregate in the **database** when the report only needs the totals; - aggregate in a **collection** when the same loaded models are also rendered, or when the grouping logic is awkward in SQL (a closure returning several group keys, for example). ## Choosing in an interview - Need **values only** → `pluck`. - Need **one item per unique key** → `keyBy`, after confirming the key really is unique. - Need **every item per category** → `groupBy`, then aggregate each group. - Need **yes/no buckets** → `partition`. A strong answer names the duplicate behaviour of `keyBy`, because silently losing answers is the classic production bug with it.
- How would you group Laravel survey answers by department and then by question in one call?Pass an array to `groupBy`: `$answers->groupBy(['department', 'question_id'])`. The framework groups by the first key, then maps each group through `groupBy` with the remaining keys, giving a collection of collections of collections.
- On an Eloquent collection, why might you prefer modelKeys() to pluck('id')?`modelKeys()` calls `getKey()` on each model, so it follows the model's configured primary key rather than a hard-coded `id`, and it returns a plain PHP array ready for `whereIn()` or `sync()`. `pluck('id')` returns a base collection and assumes the column name.
saying these in an interview costs you the question
- Uses keyBy('department') to build per-department totals.
- Thinks groupBy() returns arrays of IDs rather than collections of items.
- Believes pluck() cannot read nested relations with dot notation.
- Assumes keyBy() throws an exception on a duplicate key.
- Says modelKeys() returns an Eloquent collection of models.