What is a downstream collector, and how do you use one to count or sum the elements within each group of a groupingBy?
answer
- groupingBy(classifier, DOWNSTREAM)
- default downstream is toList()
- counting() -> Long per group
- summingInt / averagingInt / summarizingInt
- downstream aggregates each bucket
basics
~20 sgroupingBy can take a second collector that runs on each group. So instead of getting lists, you can count items per group with Collectors.counting() or add them up with summingInt(). The result is a Map from key to the computed value.
solid answer
~40 sA downstream collector is a second collector you pass to a grouping (or partitioning) collector; it decides what to do with the elements that land in each group, instead of the default of collecting them into a List. Common ones are counting() (returns a count per group as Long), summingInt/summingLong/summingDouble (totals a numeric field), averagingInt (mean), and reducing(). For example, groupingBy(Employee::department, counting()) yields Map<String, Long> of employees per department, and groupingBy(Employee::department, summingInt(Employee::salary)) yields total salary per department. The grouping collector builds the map; the downstream collector produces each value. You can also nest groupingBy as the downstream to get a multi-level map. This avoids materializing intermediate lists and reads as a single declarative aggregation.
go deeper
Knows groupingBy takes an optional second collector and can name counting() and summingInt() to count/sum per group.
Fluently chooses among counting/summing/averaging/summarizing, knows the return types (Long, Double, IntSummaryStatistics), and can nest groupingBy as a downstream for multi-level maps.
Explains downstream collectors as a composition mechanism, reasons about avoiding intermediate lists, and picks summarizingInt when several stats are needed in one pass.
Frames downstream collectors as the general two-stage partition/aggregate pattern, discusses overflow/boxing trade-offs, and guides teams on readable aggregation idioms vs. hand-rolled loops.
## The problem The Java Streams API lets you process a sequence of elements and finish with a *terminal operation* called `collect`, which accumulates the elements into a result using a `Collector`. A very common need is to **group** elements by some key and then **summarize** each group — like a SQL `GROUP BY` with `COUNT(*)` or `SUM(col)`. ## groupingBy, briefly `Collectors.groupingBy(classifier)` takes a *classifier function* (an element → key function) and returns a collector that builds a `Map<Key, List<Element>>`: every element is placed into the list under its key. By default the values are `List`s. ## What "downstream" means A **downstream collector** is a *second* collector handed to `groupingBy`: `groupingBy(classifier, downstream)`. After the classifier sorts each element into its group, the **downstream collector decides what the group's value becomes**. The default downstream is `toList()`. By swapping it, you change `Map<Key, List<E>>` into `Map<Key, Whatever-the-downstream-produces>`. Think of it as two stages: the *upstream* `groupingBy` partitions elements into buckets by key; the *downstream* aggregates the contents of each bucket. ## The common aggregating downstreams - **`counting()`** → produces a `Long`: how many elements landed in the group. `groupingBy(Order::status, counting())` → `Map<Status, Long>`. - **`summingInt(toIntFn)` / `summingLong` / `summingDouble`** → totals a numeric property. `groupingBy(Sale::region, summingInt(Sale::amount))` → `Map<Region, Integer>`. - **`averagingInt(toIntFn)` / `averagingLong` / `averagingDouble`** → arithmetic mean, always returning `Double`. - **`summarizingInt(toIntFn)`** → an `IntSummaryStatistics` object holding count, sum, min, max, and average together. ## Why this matters Without downstream collectors you'd first build `Map<Key, List<E>>`, then loop over each list to count or sum — extra memory for the lists and extra imperative code. The downstream collector does the aggregation *as elements arrive*, so no intermediate lists are created (for counting/summing) and the whole thing stays one declarative expression. ## Term glossary - **Collector**: an object describing how to fold stream elements into a result (it bundles a supplier, accumulator, combiner, and finisher — see custom-collector material). - **Classifier**: the key-extracting function for grouping. - **Terminal operation**: the stream method that produces a result and ends the pipeline (`collect`, `count`, `reduce`…). ## Edge cases - `counting()` returns `Long`, not `int` — beware autoboxing/comparisons. - `averagingInt` returns `Double` even for an empty (it can't happen per group, since a group exists only if it has ≥1 element, but a totally empty stream yields an empty map). - Empty stream → empty map; there are no zero-count entries for keys that never appeared.
- How would you get total salary per department as a Map<String, Integer>?stream.collect(Collectors.groupingBy(Employee::getDepartment, Collectors.summingInt(Employee::getSalary))).
- What does counting() return and why?A Long. It is implemented via reducing/summing of 1L per element, so it is a 64-bit count to avoid overflow on large streams.
saying these in an interview costs you the question
- Thinking groupingBy can only produce Map<K, List<V>>
- Expecting counting() to return int instead of Long
- Building Map<K,List> first then looping to count, instead of using a downstream
- Assuming keys with zero elements appear in the map