How do you count or aggregate elements per group instead of collecting lists? Explain the downstream collector.
answer
- Two-arg form: classifier + downstream
- counting() -> Map<K, Long>
- summingInt / averagingDouble for numeric aggregates
- mapping(f, toList()) transforms elements per group
- Downstream = collector applied to each group's elements
basics
~10 sUse the two-argument groupingBy(classifier, downstream). The downstream collector decides what each group becomes. For example groupingBy(Person::city, Collectors.counting()) gives Map<String, Long> of how many people are in each city.
solid answer
~40 sThe single-arg groupingBy always gives you a List per group, but usually you want a number or another shape. The two-argument form, groupingBy(classifier, downstream), applies a second collector to the elements of each group. With Collectors.counting() you get Map<K, Long> of group sizes; with summingInt/averagingDouble you get sums/averages; with mapping(f, toList()) you transform each element before collecting; with reducing or summarizingInt you get richer aggregates. Downstream collectors compose, so you can nest groupingBy inside groupingBy for multi-level grouping, e.g. by country then by city. The classifier picks the key; the downstream defines the value type. This is the stream analogue of GROUP BY with an aggregate function like COUNT(*) or SUM(x), and it avoids materializing the intermediate lists when you only need the aggregate.
code
java · 20 linesimport java.util.*;
import java.util.stream.*;
record Employee(String dept, int salary) {}
List<Employee> staff = List.of(
new Employee("ENG", 100),
new Employee("ENG", 120),
new Employee("HR", 90));
// count per dept
Map<String, Long> headcount = staff.stream()
.collect(Collectors.groupingBy(Employee::dept, Collectors.counting()));
// {ENG=2, HR=1}
// total salary per dept
Map<String, Integer> payroll = staff.stream()
.collect(Collectors.groupingBy(Employee::dept,
Collectors.summingInt(Employee::salary)));
// {ENG=220, HR=90}go deeper
Can use groupingBy(x, counting()) to count per group and read the Map<K, Long> result.
Fluently chooses among counting/summingInt/averagingDouble/mapping and knows each changes only the value type.
Explains downstream collectors as composable Collector instances, reaches for mapping/reducing/summarizing, and notes the memory/parallel benefits of aggregating without materializing lists.
Reasons about collector internals (accumulator + combiner), picks aggregation that stays associative for parallel streams, and weighs cardinality/allocation trade-offs at scale.
## The motivation `groupingBy(classifier)` always hands you `Map<K, List<T>>` — every element of each group, in a list. Often you do not want the list; you want a *summary per group*: a count, a sum, an average, a set of names, the max, etc. The **two-argument** overload exists for exactly this: ```java groupingBy(Function<T,K> classifier, Collector<T,?,D> downstream) -> Map<K, D> ``` ## What a downstream collector is A **downstream collector** is just another `Collector` applied *to the elements that fall into each group*. The outer `groupingBy` partitions elements by key; for each key it feeds that group's elements into the downstream collector and stores the downstream's result as the map value. So the value type `D` is whatever the downstream produces. ## The common downstream collectors - `Collectors.counting()` -> `Long`: number of elements in the group. - `Collectors.summingInt(f)` / `summingLong` / `summingDouble` -> the sum of `f` over the group. - `Collectors.averagingInt(f)` etc. -> a `Double` average. - `Collectors.mapping(f, downstream2)`: apply `f` to each element first, then collect with `downstream2` (e.g. `mapping(Person::name, toList())` -> the *names* per group, not whole Persons). - `Collectors.toSet()` / `toList()`: choose the container. - `Collectors.maxBy(comparator)` / `minBy` -> an `Optional<T>` per group. - `Collectors.reducing(...)`: a general fold. - `Collectors.summarizingInt(f)` -> an `IntSummaryStatistics` (count, sum, min, max, average all at once). ## Worked examples Count people per city: ```java Map<String, Long> countByCity = people.stream().collect( Collectors.groupingBy(Person::city, Collectors.counting())); ``` Total salary per department: ```java Map<String, Integer> payByDept = employees.stream().collect( Collectors.groupingBy(Employee::dept, Collectors.summingInt(Employee::salary))); ``` Names (not whole objects) per city: ```java Map<String, List<String>> namesByCity = people.stream().collect( Collectors.groupingBy(Person::city, Collectors.mapping(Person::name, Collectors.toList()))); ``` ## Why this matters beyond convenience With the downstream form, the intermediate per-group lists are **never materialized** when you only need an aggregate — `counting()` keeps just a running count. This is cheaper in memory and is friendly to parallel streams because each collector defines a combiner that merges partial results. ## The SQL mental model `groupingBy(classifier, counting())` is `SELECT key, COUNT(*) ... GROUP BY key`. Swap `counting()` for `summingInt(x)` and it becomes `SUM(x)`. The classifier is the `GROUP BY` column; the downstream is the aggregate. ## Composability Because a downstream is itself a collector, you can nest another `groupingBy` as the downstream to get multi-level maps (`Map<Country, Map<City, Long>>`). This composability is the whole design point — small collectors combine into complex aggregations without custom loops.
- You want the names per city, not the Person objects. Which downstream collector do you use?Collectors.mapping(Person::name, Collectors.toList()) — mapping transforms each element with the function before the inner collector (toList) accumulates the results.
- Why is counting() cheaper than groupingBy then calling .size() on each list?counting() keeps only a running Long per group and never builds the intermediate list, saving the memory and allocation of materializing every element.
- What type does Collectors.summarizingInt give you per group?An IntSummaryStatistics, which exposes count, sum, min, max, and average for the group in a single pass.
saying these in an interview costs you the question
- Thinking you must collect lists first and then post-process to get counts
- Using summingInt where the classifier value, not the summed field, is the integer
- Believing downstream collectors cannot be nested (they compose, enabling multi-level grouping)
- Returning Map<K, Integer> for counting() (it is Long, not Integer)