skip to content

With Laravel collections, how do pluck(), keyBy() and groupBy() differ when summarising survey answers by department?

level: middleimportance: must knowfreq 58%

answer

  1. one value, one item, many items
  2. pluck('score', 'respondent_id')
  3. keyBy: last duplicate wins
  4. groupBy: collection of collections
  5. modelKeys() returns a plain array

basics

~10 s

pluck() extracts one field per item, keyBy() indexes whole items by a field with one item per key, and groupBy() builds a collection of sub-collections, one per distinct value, so duplicates are kept.

solid answer

~30 s

`pluck('score')` returns just that field from every item, re-indexed; `pluck('score', 'respondent_id')` keys those values by another field, and dot notation such as `department.name` reaches into nested data. `keyBy('respondent_id')` keeps each whole item under the field's value, one item per key — if two items share a key, the later one silently replaces the earlier. `groupBy('department')` never drops anything: each key maps to a collection of every matching item, which you then aggregate with `map(fn ($g) => $g->avg('score'))`. For a two-way split use `partition()`, and on an Eloquent collection `modelKeys()` gives a plain array of primary keys.

code

php · 12 lines
php
<?php

$answers = collect([
    ['respondent_id' => 17, 'department' => 'Sales', 'score' => 4],
    ['respondent_id' => 23, 'department' => 'IT',    'score' => 2],
    ['respondent_id' => 31, 'department' => 'Sales', 'score' => 5],
]);

$answers->pluck('score');                 // [4, 2, 5]
$answers->pluck('score', 'respondent_id'); // [17 => 4, 23 => 2, 31 => 5]
$answers->keyBy('department')->count();    // 2 - Sales row 17 was overwritten
$answers->groupBy('department')->map->count(); // ['Sales' => 2, 'IT' => 1]

go deeper

for a junior

Know that pluck() takes one field, keyBy() indexes whole items, and groupBy() returns a collection of collections.

for a middle

Explain the duplicate behaviour of keyBy() and pluck() with a key, groupBy's multi-level and preserveKeys options, and how to aggregate each group.

for a senior

Spot silent data loss from keyBy() on non-unique fields in code review, and judge when aggregation belongs in SQL instead of in PHP after loading every row.

for a principal

Set guidance for when reports aggregate in the database versus in collections, based on data volume, reuse of loaded models and query cost.

## Three different shapes Take a collection of survey answers, each with `respondent_id`, `department` and `score`. The three methods answer three different questions: | Method | Question it answers | Result shape | Duplicates | |---|---|---|---| | `pluck('score')` | "just the scores" | `[4, 5, 2, …]`, keys `0..n-1` | all kept | | `pluck('score', 'respondent_id')` | "score per respondent" | `[17 => 4, 23 => 5, …]` | later key wins | | `keyBy('respondent_id')` | "look an answer up by respondent" | `[17 => answer, 23 => answer]` | later item wins | | `groupBy('department')` | "all answers per department" | `['Sales' => Collection, 'IT' => Collection]` | all kept | ## pluck() `pluck($value, $key = null)` delegates to `Arr::pluck`, which reads each item with `data_get`. That means: - it works on arrays **and** objects, including Eloquent models and their loaded relations; - **dot notation** walks nested data: `pluck('respondent.email')`; - without a key argument, values are appended, so the result is always a list; - with a key argument, the result is associative; a repeated key keeps the **last** value. On an Eloquent collection, `pluck()` returns a **base** `Illuminate\Support\Collection`, because the values are no longer models. ## keyBy() `keyBy($keyBy)` accepts a field name or a closure and stores the **whole item** under the computed key. It is the right tool for fast lookups — `$byRespondent->get(17)` instead of searching — but it assumes the key is **unique**. With a non-unique field such as `department`, every department ends up with only its last answer, and nothing warns you. Enum keys are converted to their value, and objects with `__toString` are cast to strings. ## groupBy() `groupBy($groupBy, $preserveKeys = false)` builds a collection whose values are **collections**. Useful details: 1. The key can be a field name, a dotted path or a closure; a closure may return an **array of keys** to put one item in several groups. 2. Passing an **array** of keys groups on several levels: `groupBy(['department', 'question_id'])`. 3. Inner groups are **re-indexed** unless you pass `true` as the second argument. 4. Boolean keys become `0`/`1` and a `null` key becomes the empty string. Because each group is a collection, aggregation chains straight on: ```php <?php $summary = $answers ->groupBy('department') ->map(fn ($group) => [ 'responses' => $group->count(), 'average' => round($group->avg('score'), 1), ]); // ['Sales' => ['responses' => 12, 'average' => 4.2], 'IT' => [...]] ``` ## partition() and reduce() for the rest - `partition(fn ($a) => $a['score'] >= 4)` returns a collection of **two** collections — passing and failing — that you destructure: `[$promoters, $others] = $answers->partition(...)`. - `reduce(fn ($carry, $answer) => $carry + $answer['score'], 0)` folds everything into one value; the callback receives the carry first, then the item, then the key. Without the second argument the initial carry is `null`. - `countBy('department')` is the shortcut when all you need is a count per value. ## modelKeys() on Eloquent collections `Illuminate\Database\Eloquent\Collection::modelKeys()` returns a **plain PHP array** of each model's primary key via `getKey()`. It differs from `pluck('id')` in two ways: it respects a custom primary key name, and it returns an array rather than a collection — convenient for `whereIn()` or `sync()`. ## PHP-side or database-side? Everything above runs **after** the rows are loaded into memory. For a few hundred answers that is fine and keeps the report logic readable. For hundreds of thousands, loading every row to call `groupBy('department')->map->avg('score')` costs memory and transfer that a `GROUP BY` in the query would avoid. A reasonable split: - aggregate in the **database** when the report only needs the totals; - aggregate in a **collection** when the same loaded models are also rendered, or when the grouping logic is awkward in SQL (a closure returning several group keys, for example). ## Choosing in an interview - Need **values only** → `pluck`. - Need **one item per unique key** → `keyBy`, after confirming the key really is unique. - Need **every item per category** → `groupBy`, then aggregate each group. - Need **yes/no buckets** → `partition`. A strong answer names the duplicate behaviour of `keyBy`, because silently losing answers is the classic production bug with it.

  • How would you group Laravel survey answers by department and then by question in one call?
    Pass an array to `groupBy`: `$answers->groupBy(['department', 'question_id'])`. The framework groups by the first key, then maps each group through `groupBy` with the remaining keys, giving a collection of collections of collections.
  • On an Eloquent collection, why might you prefer modelKeys() to pluck('id')?
    `modelKeys()` calls `getKey()` on each model, so it follows the model's configured primary key rather than a hard-coded `id`, and it returns a plain PHP array ready for `whereIn()` or `sync()`. `pluck('id')` returns a base collection and assumes the column name.

saying these in an interview costs you the question

  • Uses keyBy('department') to build per-department totals.
  • Thinks groupBy() returns arrays of IDs rather than collections of items.
  • Believes pluck() cannot read nested relations with dot notation.
  • Assumes keyBy() throws an exception on a duplicate key.
  • Says modelKeys() returns an Eloquent collection of models.