skip to content

Compare reduce and collect: when would you use each, and why is collect preferred for building mutable containers?

level: middleimportance: must knowfreq 70%

answer

  1. reduce = immutable fold to a scalar; collect = mutable reduction into a container
  2. reduce needs an associative operator + neutral identity
  3. collect Collector = supplier + accumulator + combiner + finisher
  4. reduce(String::concat) is O(n^2); use Collectors.joining()
  5. collect is parallel-safe via per-thread containers + combiner

basics

~20 s

reduce folds elements into one immutable value (like a sum) using a combining function. collect accumulates elements into a mutable container (List, Map, String) using a Collector. Use collect for building collections because it reuses one container instead of creating new ones each step.

solid answer

~50 s

Both are terminal reduction operations. reduce performs an immutable (functional) fold: it repeatedly applies an associative BinaryOperator to produce a single value such as a sum, product, or max. Its three-arg form (identity, accumulator, combiner) supports parallel folds. collect performs a mutable reduction: it uses a Collector with a supplier (creates the result container), an accumulator (adds an element into it), and a combiner (merges two containers for parallelism), plus an optional finisher. You should use collect for building containers (toList, toSet, toMap, joining, groupingBy) because mutating one container in place is far cheaper than reduce, which would have to create a fresh immutable result on every element — O(n^2) copying for something like string concatenation. Rule of thumb: reduce for a single immutable scalar; collect (mutable reduction) for accumulating into a collection or summary.

code

java · 17 lines
java
// reduce: immutable scalar fold
int sum = Stream.of(1, 2, 3, 4)
                .reduce(0, Integer::sum);   // 10

// WRONG: quadratic, creates a new String each step
String badJoin = Stream.of("a", "b", "c")
                       .reduce("", String::concat);

// RIGHT: mutable reduction, linear
String join = Stream.of("a", "b", "c")
                    .collect(Collectors.joining(", "));  // "a, b, c"

// collect into a grouping container
Map<Integer, List<String>> byLen =
    Stream.of("ann", "bo", "cara", "de")
          .collect(Collectors.groupingBy(String::length));
// {2=[bo, de], 3=[ann], 4=[cara]}

go deeper

for a junior

Knows reduce sums/combines into one value and collect builds a List/Map; can use Collectors.toList() and reduce for a sum.

for a middle

Explains immutable vs mutable reduction, the Collector's supplier/accumulator/combiner, and why joining beats reduce(concat).

for a senior

Reasons about associativity, identity neutrality, parallel combining, and choosing the right specialized collector or primitive stream.

for a principal

Designs custom Collectors, weighs allocation/GC and parallel-merge costs, and sets team guidance on stream reduction patterns.

## The shared idea: reduction Both `reduce` and `collect` are **reduction** (or **fold**) operations: they combine a stream of many elements into a single result. They are both **terminal** operations (they consume the stream and produce a result). The difference is *how* they accumulate. ## `reduce` — immutable / functional reduction `reduce` repeatedly applies a **`BinaryOperator`** — a function taking two values and returning one — to fold elements together. Three forms: 1. `Optional<T> reduce(BinaryOperator<T> op)` — no seed; returns empty Optional for an empty stream. 2. `T reduce(T identity, BinaryOperator<T> op)` — `identity` is a seed/neutral value (0 for sum, 1 for product, "" for concat). Returned when the stream is empty. 3. `U reduce(U identity, BiFunction<U,? super T,U> accumulator, BinaryOperator<U> combiner)` — lets the result type `U` differ from the element type and supplies a **combiner** to merge partial results from parallel sub-streams. Key rule: the operator must be **associative** ( (a op b) op c == a op (b op c) ) and the identity must be neutral, otherwise sequential and parallel runs disagree. `reduce` is designed for **immutable** values: each application returns a *new* value, leaving inputs untouched. Great for `int`/`long` sums, products, min/max, boolean ANDs, etc. ## Why `reduce` is wrong for containers Suppose you wanted to build a String by concatenating with `reduce("", String::concat)`. Each step creates a **brand-new String** copying all characters accumulated so far. For n elements that is O(n^2) work and O(n) garbage strings — quadratic and wasteful. The same applies to building a List immutably (copy the whole list each element). Immutable folding of a growing container is inherently expensive. ## `collect` — mutable reduction `collect` solves this with **mutable reduction**: it accumulates into a *single mutable container* that is updated in place. It is driven by a **`Collector`**, which has up to five parts: - **supplier** — creates a new empty result container (e.g. `ArrayList::new`). - **accumulator** — folds one element into the container (e.g. `List::add`). - **combiner** — merges two containers (needed for parallel streams; e.g. `addAll`). - **finisher** — optional final transform of the container into the result type. - **characteristics** — hints (CONCURRENT, UNORDERED, IDENTITY_FINISH). The JDK ships ready-made collectors via `Collectors`: `toList()`, `toSet()`, `toUnmodifiableList()`, `toMap(...)`, `joining(", ")`, `groupingBy(...)`, `partitioningBy(...)`, `counting()`, `summingInt(...)`, `averagingDouble(...)`, etc. Because the container is mutated in place, building a List/Map/String is **linear**, not quadratic. ## Parallel safety In parallel, the stream is split, each chunk is reduced independently, and the **combiner** merges partial results. For `collect`, the combiner merges containers; the supplier guarantees each thread gets its *own* container, so there is no shared mutation. That's why `collect` is safe in parallel while ad-hoc mutation in `forEach` is not. ## Choosing between them - Single **immutable scalar** result (sum, product, min/max, boolean) → **`reduce`** (or the specialized `IntStream.sum()` etc.). - Accumulating into a **mutable container** or computing a **summary/grouping** → **`collect`**. - Concatenating strings → **`collect(Collectors.joining())`**, never `reduce(String::concat)`. ## Mental model `reduce` answers "fold these into one value with this math"; `collect` answers "pour these into this bucket and hand me the bucket."

  • Why must the reduce operator be associative?
    Parallel streams split work and combine partial results in arbitrary grouping order. Only an associative operator (and neutral identity) guarantees the same result regardless of how elements are grouped, so sequential and parallel runs agree.
  • What are the components of a Collector and what is the combiner for?
    supplier (new container), accumulator (add an element), combiner (merge two containers), optional finisher (final transform), and characteristics. The combiner merges per-thread partial containers in parallel execution.

reduce is repeatedly mixing two paint blobs into a new blob until one remains. collect is keeping a single bucket and pouring each blob into it — far less wasted paint when the result keeps growing.

saying these in an interview costs you the question

  • Using reduce("", String::concat) to build strings (quadratic) instead of Collectors.joining
  • Believing reduce can mutate a shared container safely
  • Forgetting reduce's operator must be associative for correct parallel results
  • Thinking collect always returns a List (it returns whatever the Collector builds: Map, String, summary, etc.)

context