skip to content

How do you produce a multi-level grouping such as Map<Country, Map<City, Long>>, and how do nested downstream collectors compose?

level: seniorimportance: should knowfreq 55%

answer

  1. Downstream can be another groupingBy -> nested maps
  2. Outside-in: outer classifier = top key, innermost collector = leaf value
  3. 3-arg form: classifier, mapFactory (TreeMap/LinkedHashMap), downstream
  4. collectingAndThen to post-process (e.g. unmodifiableList)
  5. mapping/summingInt/toSet mix at any level

basics

~10 s

Use groupingBy whose downstream is another groupingBy. The outer classifier picks the first key, the inner groupingBy groups within each outer group, and an innermost collector like counting() produces the leaf value.

solid answer

~40 s

Because a downstream collector is itself a Collector, you can nest groupingBy inside groupingBy to build multi-level maps. The outer groupingBy(Person::country, ...) partitions by country; its downstream groupingBy(Person::city, counting()) then partitions each country's people by city and counts them, yielding Map<String, Map<String, Long>>. You can keep nesting and mix in mapping, summingInt, or toSet at any level. Two practical refinements: use the three-argument groupingBy(classifier, mapFactory, downstream) to control the map type — e.g. a TreeMap for sorted keys or LinkedHashMap for insertion order — and use collectingAndThen to post-process a group result, such as making each inner list immutable. The composition is what makes Collectors powerful: each layer only knows it produces a value for its key, so arbitrarily deep aggregations read as nested factory calls instead of nested loops.

code

java · 20 lines
java
import java.util.*;
import java.util.stream.*;

record Person(String country, String city) {}

List<Person> people = List.of(
    new Person("DE", "Berlin"), new Person("DE", "Berlin"),
    new Person("DE", "Munich"), new Person("NO", "Oslo"));

// Map<country, Map<city, count>>, country keys sorted
Map<String, Map<String, Long>> nested = people.stream().collect(
    Collectors.groupingBy(Person::country, TreeMap::new,
        Collectors.groupingBy(Person::city, Collectors.counting())));
// {DE={Berlin=2, Munich=1}, NO={Oslo=1}}

// each group made immutable via collectingAndThen
Map<String, List<Person>> byCity = people.stream().collect(
    Collectors.groupingBy(Person::city,
        Collectors.collectingAndThen(Collectors.toList(),
            Collections::unmodifiableList)));

go deeper

for a junior

Recognizes that a groupingBy can contain another collector but may not yet build a map-of-maps unaided.

for a middle

Can write a two-level groupingBy with counting and read the resulting Map<K, Map<K2, Long>>.

for a senior

Composes nested collectors fluently, controls map type via the 3-arg form, and uses collectingAndThen to post-process groups; extracts collectors into named variables for clarity.

for a principal

Designs aggregation pipelines with parallel-safety (associativity, groupingByConcurrent) and cardinality/allocation in mind, and weighs readability versus a custom collector for very deep structures.

## Why nesting works The two-argument `groupingBy(classifier, downstream)` accepts *any* `Collector` as the downstream. A `groupingBy(...)` call *is* a collector. Therefore you can pass one `groupingBy` as the downstream of another. Each level says only: 'for the elements I receive, produce a value for my key.' Stacking these gives nested maps with no manual looping. ## A concrete two-level example Count people per city *within* each country: ```java Map<String, Map<String, Long>> byCountryThenCity = people.stream().collect( Collectors.groupingBy(Person::country, Collectors.groupingBy(Person::city, Collectors.counting()))); ``` - Outer classifier `Person::country` -> top-level keys. - For each country's elements, the **downstream** `groupingBy(Person::city, counting())` runs, producing `Map<String, Long>` (city -> count). - Result: `Map<String, Map<String, Long>>`. ## Mixing collectors at each level Any level can use a different downstream. For example, country -> set of city names: ```java Map<String, Set<String>> citiesPerCountry = people.stream().collect( Collectors.groupingBy(Person::country, Collectors.mapping(Person::city, Collectors.toSet()))); ``` Here `mapping(Person::city, toSet())` first extracts the city from each person, then collects the distinct cities into a `Set`. ## Controlling the map type: the three-argument form ```java groupingBy(classifier, Supplier<M> mapFactory, downstream) -> M ``` The `mapFactory` lets you choose the concrete map. For sorted country keys: ```java Map<String, Long> sorted = people.stream().collect( Collectors.groupingBy(Person::country, TreeMap::new, Collectors.counting())); ``` `TreeMap::new` -> keys sorted; `LinkedHashMap::new` -> insertion order; `() -> new ConcurrentHashMap<>()` is used with the concurrent variant. ## Post-processing a group with collectingAndThen `collectingAndThen(downstream, finisher)` runs a collector and then applies a final transform. A classic use is making each group immutable: ```java Map<String, List<Person>> byCity = people.stream().collect( Collectors.groupingBy(Person::city, Collectors.collectingAndThen( Collectors.toList(), Collections::unmodifiableList))); ``` Now each value is an unmodifiable list. You can also use it to unwrap an `Optional` from `maxBy`, etc. ## Reading the structure Read nested collectors **outside-in**: the outermost classifier is the top map key; each inner collector defines the value type of the level above it. The innermost collector (`counting()`, `toSet()`, `summingInt(...)`) determines the leaf value. ## Why prefer this over nested loops - **Declarative**: the shape of the result is visible in the call structure. - **Parallel-ready**: each collector defines a combiner, so the whole thing can run on a parallel stream (use `groupingByConcurrent` + an unordered stream when it pays off). - **Composable**: swapping an aggregate is a one-line change to the innermost collector. ## Caveats - Deep nesting hurts readability; extract intermediate collectors into named variables (`Collector<Person,?,Long> byCity = ...`). - Mind cardinality: a map-of-maps over high-cardinality keys allocates many inner maps. - For parallel correctness the downstream aggregation should be associative (counting, summing, toSet are fine).

  • How do you make the outer map keys sorted alphabetically?
    Use the three-argument groupingBy(classifier, TreeMap::new, downstream); the mapFactory TreeMap::new produces sorted keys.
  • How would you make each group's list unmodifiable?
    Wrap the inner toList with collectingAndThen: collectingAndThen(toList(), Collections::unmodifiableList), so the finisher seals each list after collection.
  • What property must the innermost aggregation have for a parallel stream to be safe?
    It should be associative (and the combiner correct), so partial results from different threads merge consistently — counting, summing, and toSet qualify.

saying these in an interview costs you the question

  • Resorting to manual nested loops/maps instead of nesting collectors
  • Thinking the map type cannot be changed (the 3-arg mapFactory controls it)
  • Believing collectingAndThen changes the keys rather than post-processing each value
  • Ignoring readability/cardinality costs of deeply nested map-of-maps

context