skip to content

Parallel Streams

parallelStream splits the source through Spliterator.trySplit and runs on the shared common ForkJoinPool, so splittability, statelessness and associativity all matter. Interviewers ask when parallelism actually pays off, and the honest answer involves both the NQ model and the shared pool.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a parallel stream in Java, and how do you create one from a collection?

level: juniorimportance: must knowfreq 70%

answer

  1. parallelStream() on a Collection, or .parallel() on a Stream
  2. .sequential() switches back; last call before terminal wins
  3. mode applies to the whole pipeline, not one stage
  4. runs on the common ForkJoinPool by default
  5. only worth it for big data + heavy per-element work

basics

~20 s

A parallel stream splits the data into chunks and processes them on multiple threads at once, instead of one element at a time. You create one with collection.parallelStream() or by calling .parallel() on an existing stream.

solid answer

~40 s

A sequential stream processes elements one at a time on the calling thread; a parallel stream splits the source into pieces and processes those pieces concurrently across several threads, then combines the partial results. You get one in two ways: call parallelStream() on a Collection, or call parallel() on any existing stream pipeline (and sequential() switches back). The whole pipeline runs in the mode set when the terminal operation begins, so a single parallel()/sequential() call decides the mode for the entire pipeline, not just the stage it sits on. By default the work runs on the JVM-wide common ForkJoinPool. Parallelism only helps when there is enough data and enough per-element work to outweigh the splitting, thread-coordination, and merging overhead; for small or cheap workloads a sequential stream is usually faster.

go deeper

for a junior

Can name both ways to create a parallel stream (parallelStream(), parallel()) and explain it splits work across threads.

for a middle

Knows the mode applies to the whole pipeline and that the last parallel()/sequential() call wins; knows it is not automatically faster.

for a senior

Explains the default common ForkJoinPool, sizing (cores − 1), and frames the decision as a cost-benefit (overhead vs. saved work).

for a principal

Reasons about when parallel streams fit the system as a whole — shared-pool contention with other JVM tasks, and when an explicit executor or different concurrency model is the better tool.

## What a stream is A **stream** in Java is a pipeline that carries elements from a *source* (a collection, an array, a generator) through zero or more *intermediate operations* (like `map`, `filter`) to a single *terminal operation* (like `collect`, `count`, `forEach`) that produces a result or side effect. Nothing actually runs until the terminal operation is reached — intermediate operations are *lazy*. ## Sequential vs parallel A **sequential** stream processes elements one after another on the single thread that called the terminal operation. A **parallel** stream instead splits the source into several chunks, hands each chunk to a different worker thread so they run *at the same time* (on multiprocessor hardware), and then merges the partial results back into one answer. This is an example of *data parallelism*: the same operations applied to different slices of the data simultaneously. ## How to create one Two entry points: 1. `collection.parallelStream()` — a `Collection` method that gives you a stream already in parallel mode. 2. `.parallel()` — a method on any existing `Stream` that flips it to parallel mode. Its opposite is `.sequential()`. A stream is always in exactly **one** mode. The mode is whatever it was last set to **when the terminal operation starts**; calling `parallel()` or `sequential()` anywhere in the pipeline sets the mode for the *whole* pipeline, not just downstream stages. So `list.stream().parallel().filter(...).sequential().count()` runs sequentially — the last call wins. ## What runs the work By default parallel streams execute on the **common ForkJoinPool**, a single thread pool shared by the entire JVM. Its size defaults to (number of available processors − 1) worker threads, plus the submitting thread also helps, so total parallelism roughly equals the core count. ## When it pays off Splitting, scheduling threads, and merging results all cost time. Parallelism wins only when the data set is large and/or the per-element work is expensive enough that doing the work concurrently saves more time than the overhead costs. For small collections or trivial operations, a plain sequential stream is faster and simpler. Always measure rather than assume. ## Minimal example ```java List<Integer> nums = List.of(1, 2, 3, 4, 5); long evens = nums.parallelStream() .filter(n -> n % 2 == 0) .count(); ``` Here the work is far too small to benefit from parallelism — it is shown only to illustrate the API. Real candidates have thousands+ of elements and meaningful per-element cost.

  • If you write list.stream().parallel().map(f).sequential().forEach(g), does it run in parallel?
    No. The mode is whatever it was last set to when the terminal operation begins; sequential() is the last mode-setting call, so the entire pipeline runs sequentially.
  • Does parallelStream() guarantee a speedup?
    No. It only helps when there is enough data and enough per-element work to outweigh the splitting, thread-coordination, and merging overhead; otherwise it is often slower than a sequential stream.

saying these in an interview costs you the question

  • Thinking parallel is always faster — overhead often makes it slower for small/cheap work.
  • Believing parallel()/sequential() only affects the operation it is chained after (it sets the whole pipeline's mode).
  • Assuming each element gets its own thread — the source is split into a few chunks, not one-thread-per-element.

context

open as a page

Where do parallel streams run their work by default, and why does the shared common ForkJoinPool matter?

level: middleimportance: must knowfreq 68%

basics

~20 s

By default parallel streams run on a single thread pool shared by the whole JVM, called the common ForkJoinPool. Its size is about the number of CPU cores minus one. Because it is shared, a slow or blocking task can starve everything else using it.

open as a page

What correctness requirements must your stream operations meet to be safe in parallel, and what are the classic hazards?

level: seniorimportance: must knowfreq 64%

basics

~20 s

Your lambdas must be stateless and thread-safe — they can't read or write shared mutable variables. Reductions must be associative so the order of combining partial results doesn't change the answer. Never mutate a shared collection or counter from a parallel stream; use a proper reduce/collect instead.

open as a page

How does a parallel stream divide its source for parallel processing, and why does the source's data structure affect performance?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A parallel stream uses a Spliterator, whose trySplit() repeatedly cuts the source into halves so different threads can process them. Sources that split cheaply and evenly — arrays and ArrayList — parallelize well; LinkedList and Stream.iterate split poorly, so parallelism barely helps.

open as a page

How do you decide whether parallelism will actually pay off, and how do ordering-sensitive operations like findAny and forEach behave in parallel?

level: principalimportance: should knowfreq 50%

basics

~20 s

Parallelism pays off when there are many elements (N) and each costs real work (Q) — a big N×Q. For tiny data or cheap operations the overhead loses. In parallel, findAny may return any matching element (not the first), and forEach runs in no guaranteed order; use findFirst or forEachOrdered if order matters.

open as a page