Why must lambdas in a parallel stream be pure, stateless, and non-interfering, and what goes wrong if they aren't?
answer
- Parallel = Spliterator split + ForkJoinPool + merge
- Lambdas must be stateless, non-interfering, side-effect-free (pure)
- Shared mutable side-effect => races / lost updates / CME / AIOOBE
- Accumulate via reduce/collect, not by mutating shared state
- reduce needs identity + ASSOCIATIVE accumulator + combiner
basics
~20 sA parallel stream splits the work across threads. If your lambda changes shared variables or the source collection, multiple threads collide, giving wrong or random results and crashes. So the lambdas must only depend on their input and not change shared state.
solid answer
~50 sA parallel stream decomposes the source into chunks, runs the pipeline on multiple ForkJoinPool threads, and merges partial results. For that to be correct, the per-element lambdas must be: stateless (output depends only on the element, no accumulating shared variable), non-interfering (they don't modify the stream's source during execution), and effectively side-effect-free / pure. If a lambda writes to a shared mutable variable or external collection (e.g. a non-thread-safe list or a counter), concurrent threads race: you get lost updates, ConcurrentModificationException, ArrayIndexOutOfBounds inside ArrayList, or simply nondeterministic answers. The correct way to accumulate is a reduction (reduce/collect) using an identity, an associative combiner, and a thread-safe Collector — the framework handles safe splitting and merging. Ordering, the cost of splitting, boxing, and whether the operation is associative all affect whether parallel is even worth it; purity is the non-negotiable precondition before any of that.
code
java · 24 linesimport java.util.*;
import java.util.stream.*;
class ParallelPurity {
// WRONG: shared mutable side-effect across threads -> races, lost updates, exceptions
List<Integer> broken() {
List<Integer> out = new ArrayList<>();
IntStream.range(0, 1_000).parallel()
.forEach(out::add); // data race on a non-thread-safe list
return out; // size often != 1000, may even throw
}
// RIGHT: stateless mapping + framework-managed reduction
List<Integer> correct() {
return IntStream.range(0, 1_000).parallel()
.boxed()
.collect(Collectors.toList()); // safe split + merge
}
// RIGHT: associative reduction with an identity
long sum() {
return IntStream.range(0, 1_000).parallel().sum(); // associative, no shared state
}
}go deeper
Knows that parallel streams use multiple threads and that changing shared variables inside the lambda can cause wrong results.
Can name stateless/non-interfering/side-effect-free as the rules and recognizes that forEach-into-a-shared-list is unsafe, preferring collect.
Explains the Spliterator/ForkJoinPool/merge model, demonstrates the ArrayList race and its symptoms (lost updates, CME, AIOOBE), and accumulates via reduce/collect with an associative accumulator.
Reasons about associativity/identity laws, split quality per source type, common-pool starvation and isolation, and makes a measured call on whether parallelization pays off versus a sequential pipeline.
## Setup: what parallel streams actually do Calling `.parallel()` (or `collection.parallelStream()`) asks the Streams framework to run your pipeline concurrently. Mechanically it: 1. **Splits** the source into pieces using a **Spliterator** (a splittable iterator that can hand off half its elements). 2. Runs the intermediate operations on each piece **on different threads** of the common **ForkJoinPool** (a shared thread pool sized to the CPU count by default). 3. **Combines** the partial results back into one (for reductions/collectors). This is the fork/join 'divide and conquer' model. It only produces correct results if the per-element work obeys three rules. ## The three required properties - **Stateless**: an operation is stateless if its result for an element depends *only on that element*, not on any state that changes during the run (a counter, a previously-seen value, a shared accumulator). `map(x -> x*2)` is stateless. A lambda that increments a shared `int count` is **stateful** and unsafe. - **Non-interfering**: the lambda must not modify the *source* of the stream while the pipeline runs. Adding to or removing from the very list you're streaming over is interference and typically throws `ConcurrentModificationException` (even sequentially). - **Pure / side-effect-free**: ideally the lambda produces no observable side-effect at all — no writing to external collections, fields, files, or counters. (Side-effects aren't *forbidden* outright for sequential streams, but they're discouraged, and in parallel they're a correctness hazard unless the target is properly thread-safe.) '**Pure**' is the umbrella term: a pure function returns the same output for the same input and changes nothing outside itself. Pure lambdas automatically satisfy stateless + non-interfering + side-effect-free. ## What goes wrong when you violate this Consider the wrong way to collect results in parallel: ```java List<Integer> out = new ArrayList<>(); // NOT thread-safe IntStream.range(0, 1_000).parallel() .forEach(i -> out.add(i)); // shared-mutable side-effect == BUG ``` Multiple threads call `out.add` simultaneously. `ArrayList` is not synchronized, so you can get: - **Lost updates** — `out` ends up with fewer than 1000 elements because two threads wrote the same slot. - **Corruption exceptions** — `ArrayIndexOutOfBoundsException` or `NullPointerException` from inside ArrayList's resize, because its internal size/array got into an inconsistent state. - **Nondeterminism** — different counts on different runs; passes in tests, fails in prod. A shared `int count` incremented in a lambda has the classic **read-modify-write race**: `count++` is not atomic, so increments are lost. Modifying the source mid-stream yields `ConcurrentModificationException`. ## The correct way: reductions and collectors Instead of mutating shared state, you **reduce**: express accumulation as a combinable operation the framework can split and merge safely. - `reduce(identity, accumulator, combiner)` needs an **identity** (a neutral starting value, e.g. 0 for sum), an **associative** accumulator (grouping doesn't change the result: (a+b)+c == a+(b+c)), and a **combiner** to merge partial results. Associativity is what makes parallel merging valid. - `collect(Collectors.toList())` / `groupingBy` use a `Collector` that the framework runs per-partition into separate containers and then merges, so no shared mutable container is touched by multiple threads. ```java List<Integer> out = IntStream.range(0, 1_000).parallel() .boxed() .collect(Collectors.toList()); // safe, framework-managed long sum = IntStream.range(0, 1_000).parallel().sum(); // safe associative reduction ``` Here each thread accumulates into its own container/partial and the framework merges — no races. ## When is parallel even worth it (the senior+ judgment) Purity is the *correctness* precondition; performance is a separate question. Parallel streams help only when: the data set is large, the per-element work is non-trivial, the source splits cheaply (arrays/ArrayList split well; LinkedList and IO-bound sources split poorly), the operation is associative, and you're not paying heavy boxing. They also share the **common ForkJoinPool** with the rest of the JVM, so a blocking or long task in a parallel stream can starve unrelated parallel work. Often a plain sequential stream or loop is faster and always simpler. Levels: senior = states the three rules, shows the ArrayList race, and uses reduce/collect instead; principal = reasons about associativity/identity, Spliterator split quality, common-pool contention, and whether parallelization actually pays off versus its hazards.
- What's the correct way to build a list of results from a parallel stream?Use collect(Collectors.toList()) (or toList()). The Collector accumulates each partition into its own container and the framework merges them, so no shared mutable list is touched by multiple threads. Never forEach into a shared non-thread-safe collection.
- Why must the accumulator in a parallel reduce be associative?Because the framework splits the data, reduces each chunk independently, and merges the partials in an unspecified grouping. Associativity ((a∘b)∘c == a∘(b∘c)) guarantees the result is independent of how the elements were grouped, so parallel and sequential give the same answer.
A parallel stream is a team splitting a deck of cards to count them. If each person keeps their own tally and you add the tallies at the end (reduction), it's fast and correct. If everyone instead scribbles onto one shared tally sheet at once (shared mutable state), they overwrite each other and the total is wrong — that's the ArrayList-in-forEach bug.
saying these in an interview costs you the question
- Using forEach to add to a shared ArrayList in a parallel stream
- Incrementing a shared counter in a lambda and expecting a correct total
- Modifying the stream's source collection mid-pipeline
- Assuming .parallel() is always faster (ignoring split cost, boxing, common-pool contention)
- Forgetting that the accumulator must be associative for parallel reduction to be correct