skip to content

What is Collectors.reducing, and how does it differ from Stream.reduce and from summing/counting collectors?

level: seniorimportance: should knowfreq 40%

answer

  1. reducing = Stream.reduce as a Collector (downstream-able)
  2. 3 overloads: op -> Optional; identity,op -> T; identity,mapper,op -> U
  3. 3-arg form is the groupingBy workhorse
  4. counting/summingInt are specialized reducing — prefer them
  5. operator must be associative for parallel

basics

~20 s

Collectors.reducing folds elements into a single value using a combining function, like Stream.reduce but packaged as a collector so it can be a downstream of groupingBy. summingInt/counting are really specialized reducing collectors; reducing is the general form when no built-in fits.

solid answer

~40 s

Collectors.reducing performs a reduction (folding many elements into one) but as a Collector, so it slots into the downstream position of groupingBy/partitioningBy where Stream.reduce cannot go. It has three overloads: reducing(op) returning Optional<T> (no identity), reducing(identity, op) returning T, and reducing(identity, mapper, op) which maps each element first then reduces. The difference from Stream.reduce is purely positional/compositional — same semantics, but usable inside a collector. counting(), summingInt(), and averagingInt() are essentially convenience reducing collectors, so prefer them when they fit (clearer and avoids boxing). Reach for reducing only when you need a custom binary operation per group, e.g. the max-length string per category. Note: for grouping, the three-arg reducing(identity, mapper, op) is the one you usually want, and the operator must be associative for correctness under parallelism.

code

java · 18 lines
java
record Book(String author, String title) {}

List<Book> books = List.of(
    new Book("Ann", "Java"),
    new Book("Ann", "Effective Java"),
    new Book("Bob", "Go"));

// longest title per author via 3-arg reducing
Map<String, String> longest = books.stream()
    .collect(Collectors.groupingBy(Book::author,
        Collectors.reducing("", Book::title,
            (a, b) -> a.length() >= b.length() ? a : b)));
// {Ann="Effective Java", Bob="Go"}

// Prefer the specialized collector when it fits:
Map<String, Long> counts = books.stream()
    .collect(Collectors.groupingBy(Book::author, Collectors.counting()));
// rather than reducing(0L, b -> 1L, Long::sum)

go deeper

for a junior

Knows reducing combines elements into one value, similar to summing.

for a middle

Knows the three overloads and that reducing is Stream.reduce usable as a downstream collector.

for a senior

Picks reducing only for custom operations, uses the three-arg form in groupingBy, and recognizes counting/summing as specialized reducing collectors.

for a principal

Reasons about associativity/identity for parallel correctness, weighs reducing vs. a custom Collector, and steers teams toward the clearest built-in for each aggregation.

## What 'reducing' means **Reduction** (a.k.a. fold) combines a sequence of elements into a single result by repeatedly applying a binary operator: `((a op b) op c) op d`. Sum, product, min, max, and concatenation are all reductions. The `Stream` interface already has `reduce(...)`. `Collectors.reducing(...)` is the *same idea wrapped as a `Collector`* so it can be used wherever a collector is expected — most importantly as a **downstream** of `groupingBy`. ## The three overloads 1. **`reducing(BinaryOperator<T> op)`** → returns `Optional<T>`. No identity/start value, so an empty input has no result → `Optional.empty()`. Example: `reducing(BinaryOperator.maxBy(comparator))`. 2. **`reducing(T identity, BinaryOperator<T> op)`** → returns `T`. `identity` is the start value and the result for an empty input. Example: `reducing(0, Integer::sum)`. 3. **`reducing(U identity, Function<T,U> mapper, BinaryOperator<U> op)`** → maps each element with `mapper` first, then reduces with `op` from `identity`. This is the most useful one inside `groupingBy`, because the elements in a group are usually the full objects but you want to reduce some projected value. ## Why it exists if Stream.reduce already does this Because **position matters**. On a flat pipeline you'd just write `stream.reduce(...)`. But after `groupingBy`, the per-group elements are handed to a *downstream collector* — and `Stream.reduce` is not a collector, so it can't sit there. `Collectors.reducing` fills that slot. Example — longest title per author: ``` groupingBy(Book::author, reducing("", Book::title, (a, b) -> a.length() >= b.length() ? a : b)) // Map<Author, String> ``` ## Relationship to summing/counting `counting()`, `summingInt()`, `averagingInt()`, `minBy()`, `maxBy()` are all effectively **specialized reducing collectors** (or thin wrappers). They are clearer, communicate intent, and the numeric ones avoid repeated boxing. **Rule of thumb: use the specialized collector when one fits; use `reducing` only for a custom binary operation that has no built-in.** ## Correctness requirements - The operator should be **associative** (`(a op b) op c == a op (b op c)`) so results are deterministic, especially in **parallel** streams where the stream is split and partial results combined. - `identity` must be a true identity: `identity op x == x` for all `x` (e.g. `0` for sum, `""` for max-length string only because comparison treats it as the shortest — be careful). - The two-arg/three-arg forms with identity never produce `Optional`; the one-arg form does, to represent 'no elements'. ## Common confusion Developers sometimes use `reducing(0, e -> 1, Integer::sum)` to count — that's exactly what `counting()` does, so just use `counting()`. Likewise `reducing(0, Order::amount, Integer::sum)` is `summingInt(Order::amount)`. ## Term glossary - **BinaryOperator<T>**: a function taking two `T`s and returning a `T` (the combining op). - **Identity**: the neutral start value where `identity op x == x`. - **Associative**: grouping of operations doesn't change the result; required for correct parallel reduction. - **Downstream collector**: the collector used to aggregate each group of a grouping.

  • When should you use reducing instead of summingInt or counting?
    Only when you need a custom binary operation with no built-in equivalent; for plain sums/counts/averages the specialized collectors are clearer and avoid boxing.
  • Why must the reducing operator be associative?
    Parallel streams split the data and combine partial results in unspecified groupings; a non-associative operator would yield nondeterministic or wrong results.

saying these in an interview costs you the question

  • Using reducing where counting()/summingInt() would be clearer
  • Forgetting the no-identity overload returns Optional
  • Using a non-associative operator and getting wrong parallel results
  • Thinking reducing differs semantically from Stream.reduce (only its position differs)

context