What correctness requirements must your stream operations meet to be safe in parallel, and what are the classic hazards?
answer
- lambdas stateless + thread-safe (depend only on inputs)
- non-interference: don't mutate the source mid-stream
- reductions must be associative with a true identity
- shared counter/ArrayList from forEach = data race
- fix = reduce / collect, let the framework merge
basics
~20 sYour lambdas must be stateless and thread-safe — they can't read or write shared mutable variables. Reductions must be associative so the order of combining partial results doesn't change the answer. Never mutate a shared collection or counter from a parallel stream; use a proper reduce/collect instead.
solid answer
~50 sParallel correctness comes down to a few rules. First, the functions you pass — to map, filter, reduce — must be stateless: their result must depend only on their input, never on a mutable field or shared variable, because they run concurrently on many threads. Second, they must be non-interfering: they must not modify the stream's source while it is being processed. Third, reduction operations must be associative and have a proper identity, because the framework groups and combines partial results in an unspecified order; a non-associative combiner gives different answers in parallel vs. sequential. The classic hazards are shared mutable state — for example incrementing a shared counter or adding to a plain ArrayList from forEach — which causes data races, lost updates, or corrupted collections (a non-thread-safe list can throw or silently drop entries). The fix is to express the computation as a reduce or collect with the right identity/accumulator/combiner (or a concurrent collector), so the framework manages the merging safely instead of you sharing state.
code
java · 11 lines// BROKEN: shared mutable ArrayList from a parallel forEach — data race
List<String> out = new ArrayList<>();
items.parallelStream().forEach(i -> out.add(i.toUpperCase())); // may throw or lose data
// CORRECT: let the framework merge per-chunk results
List<String> safe = items.parallelStream()
.map(String::toUpperCase)
.collect(Collectors.toList());
// CORRECT reduction: associative combiner + identity
int total = nums.parallelStream().reduce(0, Integer::sum); // 0 is identity, + is associativego deeper
Knows you shouldn't change a shared variable from inside a parallel stream and should use collect instead.
Can explain stateless lambdas and that an ArrayList shared across a parallel forEach is unsafe.
Articulates statelessness, non-interference, and associativity-with-identity, and refactors shared-state bugs into reduce/collect.
Reviews APIs and team conventions for parallel-safety, reasons about determinism/ordering guarantees, and chooses concurrent collectors vs. atomic accumulators deliberately.
## Why parallel adds new rules In a *sequential* stream everything runs on one thread in a single order, so even sloppy code (mutating an outside variable from a lambda) usually "works." In a *parallel* stream the same lambda runs **simultaneously on multiple threads** over different chunks, and partial results are combined in an **unspecified order**. Code that quietly assumed single-threaded, in-order execution now breaks. Three properties keep you safe. ## 1. Stateless behavioural parameters A *behavioural parameter* is the lambda/function you hand to `map`, `filter`, `reduce`, etc. It must be **stateless**: its result depends only on its input arguments, not on any mutable state that could change during execution. Reading or writing a shared field from inside the lambda is a *data race* — two threads touching the same memory without synchronization — which yields undefined, non-deterministic results. ## 2. Non-interference The lambdas must **not modify the stream's source** while the pipeline runs. Adding to or removing from the backing collection mid-stream can corrupt the traversal (often surfacing as `ConcurrentModificationException`, or worse, silent corruption under parallelism). ## 3. Associative reductions (with identity) A *reduction* collapses many elements into one result (sum, max, concatenation) via a combine function. Because the framework splits the data, computes partial results per chunk, and merges them in an **arbitrary grouping/order**, the combine function must be **associative**: `(a∘b)∘c` must equal `a∘(b∘c)`. Addition and `max` are associative; subtraction and *average-of-pairs* are not, so a non-associative reducer can produce a *different answer in parallel than sequentially*. `reduce` also needs a true **identity** value `i` such that `combine(i, x) == x`; a wrong identity gets folded in once per chunk and corrupts the total. ## The classic hazard: shared mutable state The single most common parallel-stream bug is mutating shared state from a lambda: ```java // BROKEN under parallel: data race + non-thread-safe list List<Result> results = new ArrayList<>(); items.parallelStream().forEach(i -> results.add(transform(i))); // DON'T int[] counter = {0}; items.parallelStream().forEach(i -> counter[0]++); // DON'T (lost updates) ``` `ArrayList` is not thread-safe; concurrent `add` can throw `ArrayIndexOutOfBoundsException`, drop elements, or corrupt internal state. The shared counter loses updates because `++` is not atomic. ### The fix: let the framework merge Express the computation as a **reduction/collection** so the framework owns the merging: ```java List<Result> results = items.parallelStream() .map(MyApp::transform) .collect(Collectors.toList()); // safe, merges per-chunk long count = items.parallelStream().filter(MyApp::matches).count(); ``` `collect` uses a *supplier / accumulator / combiner* (or a thread-safe concurrent collector such as `groupingByConcurrent`) so each thread accumulates into its *own* container and the framework combines them — no shared mutation. If you genuinely must share a counter, use an atomic/`LongAdder`, but a proper reduction is almost always cleaner. ## Summary checklist - Lambdas: **stateless**, depend only on inputs. - Do **not** mutate the source mid-stream (non-interference). - Reductions: **associative** combiner + correct **identity**. - Never write shared mutable variables/collections from a parallel lambda — use `reduce`/`collect`.
- Why must a reduce combiner be associative, and give a function that is NOT?Because the framework computes partial results per chunk and merges them in an unspecified grouping/order; only an associative operation gives the same answer regardless of grouping. Subtraction is not associative — (a−b)−c ≠ a−(b−c) — so it yields different results in parallel vs. sequential.
- What is the recommended fix for code that adds results into a shared ArrayList inside a parallel forEach?Replace it with a map().collect(Collectors.toList()) (or another collector). collect gives each thread its own container and lets the framework combine them safely, eliminating the shared-mutation data race.
saying these in an interview costs you the question
- Mutating a shared ArrayList or counter from a parallel forEach instead of using collect/reduce.
- Using a non-associative combiner (e.g. subtraction, average) and expecting consistent results in parallel.
- Passing a wrong identity to reduce, which gets folded in once per chunk.
- Thinking 'it worked sequentially' proves parallel correctness — single-threaded order hides the races.