skip to content

What does Collectors.groupingBy do, and what is the shape of the result?

level: juniorimportance: must knowfreq 78%

answer

  1. Classifier function -> key
  2. Result is Map<K, List<T>>
  3. Stream equivalent of SQL GROUP BY
  4. Default HashMap + ArrayList values
  5. Null key throws NPE

basics

~20 s

groupingBy takes a function that maps each element to a key. It returns a Map where each key points to a list of all the elements that produced that key, like sorting items into labeled buckets.

solid answer

~40 s

Collectors.groupingBy(classifier) is a stream collector that buckets elements by a key. You pass a classifier function; for each element it computes a key, then collects every element with the same key into a List under that key. The result is a Map<K, List<T>>. For example, grouping a list of people by their city gives Map<String, List<Person>>. By default the map is a HashMap and the values are ArrayLists. It is the stream equivalent of a SQL GROUP BY. The single-argument form always produces lists; if you want counts, sums, or another structure per group, you add a second argument (a downstream collector). groupingBy never returns null keys gracefully — a null key from the classifier throws NullPointerException in the standard implementation.

code

java · 13 lines
java
import java.util.*;
import java.util.stream.*;

record Person(String name, String city) {}

List<Person> people = List.of(
    new Person("Alice", "Berlin"),
    new Person("Bob",   "Berlin"),
    new Person("Carol", "Oslo"));

Map<String, List<Person>> byCity =
    people.stream().collect(Collectors.groupingBy(Person::city));
// {Berlin=[Person[Alice], Person[Bob]], Oslo=[Person[Carol]]}

go deeper

for a junior

Can state that groupingBy takes a key-extractor and returns Map<K,List<T>>, and give a simple example like grouping by first letter.

for a middle

Knows the defaults (HashMap, ArrayList), the SQL GROUP BY analogy, and the null-key NPE pitfall.

for a senior

Explains it as the one-arg special case of a more general API, and immediately reaches for downstream collectors when a list-per-group is not what is wanted.

for a principal

Frames groupingBy in terms of the Collector contract (supplier/accumulator/combiner/finisher) and reasons about its parallel-friendliness and memory profile for large cardinality keys.

## What problem this solves You often have a flat collection and want to split it into named groups: orders by customer, words by first letter, employees by department. Doing this by hand means creating a `Map`, looping, calling `computeIfAbsent`, and appending. `Collectors.groupingBy` does exactly that in one declarative call inside a stream. ## Terms, defined from scratch - **Stream**: a pipeline that processes a sequence of elements one by one (e.g. `list.stream()`). - **Collector**: an object that tells a stream's terminal `collect(...)` operation *how to accumulate* elements into a final result (a `List`, a `Map`, a count, etc.). `Collectors` is a factory class full of ready-made collectors. - **Classifier function**: a function you supply, of type `Function<T, K>`, that takes one element and returns the *key* (the group label) it belongs to. ## The single-argument form ```java Map<K, List<T>> result = stream.collect(Collectors.groupingBy(classifier)); ``` For each element `t`, it computes `k = classifier.apply(t)`, then adds `t` to the list stored under key `k`. Elements that yield the same key land in the same list. Order of elements *within* each list follows encounter order of the stream. Concretely, grouping people by city: ```java Map<String, List<Person>> byCity = people.stream().collect(Collectors.groupingBy(Person::city)); ``` If Alice and Bob live in Berlin and Carol lives in Oslo, you get `{"Berlin":[Alice,Bob], "Oslo":[Carol]}`. ## What the defaults are - The returned `Map` is a `HashMap` (no ordering guarantee). - Each group's container is an `ArrayList`. - There is **no guarantee** the map is mutable beyond being a standard HashMap; treat it as a normal map. ## Relationship to SQL Think of it as `SELECT key, <aggregate> FROM table GROUP BY key`. The single-arg form is like `GROUP BY` that hands you every row in each group (the list). Adding a *downstream collector* (a second argument) is like applying an aggregate (`COUNT(*)`, `SUM(x)`) to each group. ## Edge cases to know - A **null classifier result** throws `NullPointerException` in the standard collector — there is no group for `null`. (Some custom collectors tolerate it, but the JDK one does not.) - Empty stream → empty map (not null). - It is an *eager* terminal operation: the whole stream is consumed and the map is built before `collect` returns. From this base, every other variant is just: supply a different map factory and/or a different downstream collector.

  • What happens if the classifier returns null for some element?
    The standard JDK groupingBy collector throws NullPointerException; there is no automatic 'null' bucket. You must filter out or map nulls to a sentinel key beforehand.
  • What concrete Map and List types do you get by default?
    A HashMap whose values are ArrayLists. Neither ordering nor a specific mutable type is guaranteed beyond standard HashMap/ArrayList behavior.

Sorting a deck of cards into piles by suit: the classifier 'which suit?' decides the pile, and each pile is the list of cards that share that suit.

saying these in an interview costs you the question

  • Saying groupingBy returns a List instead of a Map
  • Assuming the result map is sorted or insertion-ordered (it is a HashMap)
  • Claiming a null classifier key just creates a 'null' group (it throws NPE)
  • Confusing the classifier (element -> key) with a downstream collector

context