skip to content

What is a downstream collector, and how do you use one to count or sum the elements within each group of a groupingBy?

level: juniorimportance: must knowfreq 72%

answer

  1. groupingBy(classifier, DOWNSTREAM)
  2. default downstream is toList()
  3. counting() -> Long per group
  4. summingInt / averagingInt / summarizingInt
  5. downstream aggregates each bucket

basics

~20 s

groupingBy can take a second collector that runs on each group. So instead of getting lists, you can count items per group with Collectors.counting() or add them up with summingInt(). The result is a Map from key to the computed value.

solid answer

~40 s

A downstream collector is a second collector you pass to a grouping (or partitioning) collector; it decides what to do with the elements that land in each group, instead of the default of collecting them into a List. Common ones are counting() (returns a count per group as Long), summingInt/summingLong/summingDouble (totals a numeric field), averagingInt (mean), and reducing(). For example, groupingBy(Employee::department, counting()) yields Map<String, Long> of employees per department, and groupingBy(Employee::department, summingInt(Employee::salary)) yields total salary per department. The grouping collector builds the map; the downstream collector produces each value. You can also nest groupingBy as the downstream to get a multi-level map. This avoids materializing intermediate lists and reads as a single declarative aggregation.

go deeper

for a junior

Knows groupingBy takes an optional second collector and can name counting() and summingInt() to count/sum per group.

for a middle

Fluently chooses among counting/summing/averaging/summarizing, knows the return types (Long, Double, IntSummaryStatistics), and can nest groupingBy as a downstream for multi-level maps.

for a senior

Explains downstream collectors as a composition mechanism, reasons about avoiding intermediate lists, and picks summarizingInt when several stats are needed in one pass.

for a principal

Frames downstream collectors as the general two-stage partition/aggregate pattern, discusses overflow/boxing trade-offs, and guides teams on readable aggregation idioms vs. hand-rolled loops.

## The problem The Java Streams API lets you process a sequence of elements and finish with a *terminal operation* called `collect`, which accumulates the elements into a result using a `Collector`. A very common need is to **group** elements by some key and then **summarize** each group — like a SQL `GROUP BY` with `COUNT(*)` or `SUM(col)`. ## groupingBy, briefly `Collectors.groupingBy(classifier)` takes a *classifier function* (an element → key function) and returns a collector that builds a `Map<Key, List<Element>>`: every element is placed into the list under its key. By default the values are `List`s. ## What "downstream" means A **downstream collector** is a *second* collector handed to `groupingBy`: `groupingBy(classifier, downstream)`. After the classifier sorts each element into its group, the **downstream collector decides what the group's value becomes**. The default downstream is `toList()`. By swapping it, you change `Map<Key, List<E>>` into `Map<Key, Whatever-the-downstream-produces>`. Think of it as two stages: the *upstream* `groupingBy` partitions elements into buckets by key; the *downstream* aggregates the contents of each bucket. ## The common aggregating downstreams - **`counting()`** → produces a `Long`: how many elements landed in the group. `groupingBy(Order::status, counting())` → `Map<Status, Long>`. - **`summingInt(toIntFn)` / `summingLong` / `summingDouble`** → totals a numeric property. `groupingBy(Sale::region, summingInt(Sale::amount))` → `Map<Region, Integer>`. - **`averagingInt(toIntFn)` / `averagingLong` / `averagingDouble`** → arithmetic mean, always returning `Double`. - **`summarizingInt(toIntFn)`** → an `IntSummaryStatistics` object holding count, sum, min, max, and average together. ## Why this matters Without downstream collectors you'd first build `Map<Key, List<E>>`, then loop over each list to count or sum — extra memory for the lists and extra imperative code. The downstream collector does the aggregation *as elements arrive*, so no intermediate lists are created (for counting/summing) and the whole thing stays one declarative expression. ## Term glossary - **Collector**: an object describing how to fold stream elements into a result (it bundles a supplier, accumulator, combiner, and finisher — see custom-collector material). - **Classifier**: the key-extracting function for grouping. - **Terminal operation**: the stream method that produces a result and ends the pipeline (`collect`, `count`, `reduce`…). ## Edge cases - `counting()` returns `Long`, not `int` — beware autoboxing/comparisons. - `averagingInt` returns `Double` even for an empty (it can't happen per group, since a group exists only if it has ≥1 element, but a totally empty stream yields an empty map). - Empty stream → empty map; there are no zero-count entries for keys that never appeared.

  • How would you get total salary per department as a Map<String, Integer>?
    stream.collect(Collectors.groupingBy(Employee::getDepartment, Collectors.summingInt(Employee::getSalary))).
  • What does counting() return and why?
    A Long. It is implemented via reducing/summing of 1L per element, so it is a 64-bit count to avoid overflow on large streams.

saying these in an interview costs you the question

  • Thinking groupingBy can only produce Map<K, List<V>>
  • Expecting counting() to return int instead of Long
  • Building Map<K,List> first then looping to count, instead of using a downstream
  • Assuming keys with zero elements appear in the map

context