When should you use reduce versus collect for a reduction, and why?
answer
- reduce = immutable value fold
- collect = mutable container build
- reduce-into-List = O(n^2) copy trap
- Collector = supplier + accumulator + combiner (+ finisher)
- joining/groupingBy/toList -> collect
basics
~20 sUse reduce when you combine values into a new immutable result, like summing numbers or finding a max. Use collect when you accumulate into a mutable container, like building a List, Map, or StringBuilder. reduce makes a fresh value each step; collect mutates one container, which is far more efficient for large results.
solid answer
~40 sreduce is an immutable, value-folding reduction: each step produces a brand-new result from the running result and the next element. That's ideal for small, immutable results — sums, max, boolean ANDs. But for accumulating into a container (a List, Map, StringBuilder), reduce would create a fresh copy at every step (O(n^2) work and garbage), so it's the wrong tool. collect performs a mutable reduction: it allocates one mutable result container, mutates it in place as elements flow in, and (in parallel) merges per-thread containers with a combiner — giving O(n). So: combining values immutably → reduce; building/mutating a container → collect (Collectors.toList, joining, groupingBy, etc.). Both still obey the same identity/associativity discipline; collect's Collector just packages supplier + accumulator + combiner together and is the idiomatic, performant choice for container-building.
code
java · 19 linesList<String> words = List.of("alpha", "beta", "gamma");
// reduce: immutable value fold -> total length
int totalLen = words.stream()
.reduce(0, (len, w) -> len + w.length(), Integer::sum);
// collect: mutable reduction -> build a List (O(n))
List<String> upper = words.stream()
.map(String::toUpperCase)
.collect(Collectors.toList());
// collect for strings (StringBuilder under the hood) — NOT reduce("", String::concat)
String joined = words.stream().collect(Collectors.joining(", "));
// What a Collector is, made explicit:
List<String> manual = words.stream().collect(
ArrayList::new, // supplier (fresh mutable container)
List::add, // accumulator (mutates in place)
List::addAll); // combiner (merges parallel containers)go deeper
Knows reduce is for combining numbers into a value and collect is for making a List/Map, even if not the performance reasons.
Explains that collect mutates one container while reduce creates new values, and can give Collectors.toList/joining/groupingBy as collect use cases.
Articulates the O(n^2) copy trap of reducing into a container, names the Collector's supplier/accumulator/combiner/finisher, and chooses correctly between value-fold and container-build.
Reasons about thread confinement via the per-chunk supplier, the immutability constraint that forbids in-place mutation in reduce, primitive-stream specializations to avoid boxing, and the performance/GC characteristics of each at scale.
## Two flavors of reduction Both `reduce` and `collect` are **reductions** — they collapse a stream into a result. The difference is *how the running result is updated*: - **`reduce` is an immutable / value-folding reduction.** Each application of the accumulator returns a **new** result object; the previous one is discarded. The signature `(U, T) -> U` returns a fresh `U` every step. - **`collect` is a mutable reduction.** It creates **one** mutable result container up front and **mutates it in place** as each element arrives — no new object per step. ## Why this distinction matters: the O(n^2) trap Suppose you try to build a `List` with reduce: ```java // ANTI-PATTERN — quadratic, lots of garbage List<T> list = stream.reduce( new ArrayList<>(), (acc, t) -> { var copy = new ArrayList<>(acc); copy.add(t); return copy; }, (a, b) -> { var c = new ArrayList<>(a); c.addAll(b); return c; }); ``` Because reduce's contract says the accumulator must return a value (and must not corrupt the identity, which may be reused across parallel chunks), you cannot legally mutate `acc` in place — you must copy. Copying a growing list at every one of n steps is **O(n^2)** time and allocates n throwaway lists. For a 1-million-element stream this is catastrophic. `collect` solves exactly this: it allocates a single `ArrayList` and appends in place — **O(n)**, minimal garbage: ```java List<T> list = stream.collect(Collectors.toList()); ``` ## What a Collector actually is `collect(Collector)` packages three (or four) pieces that mirror reduce's parts but for a mutable container: - **supplier** `() -> A` — creates a fresh mutable container (e.g. `ArrayList::new`). This is collect's analogue of reduce's *identity*, but it produces a *new mutable* container each time rather than reusing one immutable value. - **accumulator** `(A, T) -> void` — folds an element into the container by mutating it (e.g. `List::add`). - **combiner** `(A, A) -> A` — merges two containers for parallel execution (e.g. `addAll`). - **finisher** `(A) -> R` — optional final transform. Notice the supplier returns a new container per parallel chunk, so mutation is thread-confined: each thread mutates its own container, then combiners merge them. That's why mutation is safe here even in parallel, whereas mutating reduce's shared identity would be a data race. ## Decision rule - **Use `reduce`** when the result is an **immutable value** you fold by combining: `sum`, `product`, `max/min`, boolean `and/or`, a single concatenated number, etc. The result type is small and a new value per step is cheap. - **Use `collect`** when the result is a **container you mutate/build**: a `List`/`Set`/`Map` (`toList`, `toSet`, `groupingBy`, `toMap`), a joined `String` (`Collectors.joining` over a `StringBuilder`), partitioning, summarizing statistics, etc. ## Edge cases & nuances - For numeric sums, prefer the **primitive specializations** (`IntStream.sum()`, `mapToInt(...).sum()`) over `reduce` to avoid boxing — they're reductions under the hood but type-specialized. - String building: `reduce("", String::concat)` is the O(n^2) string-copy trap; use `collect(Collectors.joining())`, which uses a `StringBuilder`. - Both forms still require the **same laws** (neutral identity/supplier, associative/consistent accumulator and combiner) for correct parallelization; collect simply moves the identity from "a shared immutable seed" to "a freshly supplied mutable container," which is what makes in-place mutation safe. ## One-line summary reduce = combine into a new immutable value; collect = accumulate into a mutable container. Match the tool to whether your result is a value or a container, and you avoid both correctness bugs and the O(n^2) copy trap.
- Why is building a List with reduce O(n^2) but collect O(n)?reduce's accumulator must return a value and may have its identity reused across chunks, so it cannot legally mutate in place — it copies the whole list each step (n copies of growing size = O(n^2)). collect allocates one container per (parallel) chunk and mutates in place (append = amortized O(1)), giving O(n).
- How does collect stay thread-safe in parallel despite using mutation?Its supplier creates a separate mutable container per parallel chunk, so each thread mutates only its own container (thread confinement), and the combiner merges the containers afterward. There's no shared mutable state, so no data race.
reduce is like merging two balls of clay into a new ball each time — fine for a single growing ball of numbers. collect is like writing onto one shared notebook: you keep adding lines to the same notebook instead of recopying the whole thing every line. For a thousand lines, recopying (reduce) is absurdly slow; appending (collect) is the obvious choice.
saying these in an interview costs you the question
- Building a List/Map/String with reduce by copying the accumulator each step.
- Claiming reduce can mutate its identity/accumulator in place (it must not — the identity may be reused).
- Using reduce("", String::concat) for joining instead of Collectors.joining.
- Thinking collect is just reduce with extra syntax — missing that it's a mutable reduction with a per-chunk supplier.