skip to content

What is a parallel stream in Java, and how do you create one from a collection?

level: juniorimportance: must knowfreq 70%

answer

  1. parallelStream() on a Collection, or .parallel() on a Stream
  2. .sequential() switches back; last call before terminal wins
  3. mode applies to the whole pipeline, not one stage
  4. runs on the common ForkJoinPool by default
  5. only worth it for big data + heavy per-element work

basics

~20 s

A parallel stream splits the data into chunks and processes them on multiple threads at once, instead of one element at a time. You create one with collection.parallelStream() or by calling .parallel() on an existing stream.

solid answer

~40 s

A sequential stream processes elements one at a time on the calling thread; a parallel stream splits the source into pieces and processes those pieces concurrently across several threads, then combines the partial results. You get one in two ways: call parallelStream() on a Collection, or call parallel() on any existing stream pipeline (and sequential() switches back). The whole pipeline runs in the mode set when the terminal operation begins, so a single parallel()/sequential() call decides the mode for the entire pipeline, not just the stage it sits on. By default the work runs on the JVM-wide common ForkJoinPool. Parallelism only helps when there is enough data and enough per-element work to outweigh the splitting, thread-coordination, and merging overhead; for small or cheap workloads a sequential stream is usually faster.

go deeper

for a junior

Can name both ways to create a parallel stream (parallelStream(), parallel()) and explain it splits work across threads.

for a middle

Knows the mode applies to the whole pipeline and that the last parallel()/sequential() call wins; knows it is not automatically faster.

for a senior

Explains the default common ForkJoinPool, sizing (cores − 1), and frames the decision as a cost-benefit (overhead vs. saved work).

for a principal

Reasons about when parallel streams fit the system as a whole — shared-pool contention with other JVM tasks, and when an explicit executor or different concurrency model is the better tool.

## What a stream is A **stream** in Java is a pipeline that carries elements from a *source* (a collection, an array, a generator) through zero or more *intermediate operations* (like `map`, `filter`) to a single *terminal operation* (like `collect`, `count`, `forEach`) that produces a result or side effect. Nothing actually runs until the terminal operation is reached — intermediate operations are *lazy*. ## Sequential vs parallel A **sequential** stream processes elements one after another on the single thread that called the terminal operation. A **parallel** stream instead splits the source into several chunks, hands each chunk to a different worker thread so they run *at the same time* (on multiprocessor hardware), and then merges the partial results back into one answer. This is an example of *data parallelism*: the same operations applied to different slices of the data simultaneously. ## How to create one Two entry points: 1. `collection.parallelStream()` — a `Collection` method that gives you a stream already in parallel mode. 2. `.parallel()` — a method on any existing `Stream` that flips it to parallel mode. Its opposite is `.sequential()`. A stream is always in exactly **one** mode. The mode is whatever it was last set to **when the terminal operation starts**; calling `parallel()` or `sequential()` anywhere in the pipeline sets the mode for the *whole* pipeline, not just downstream stages. So `list.stream().parallel().filter(...).sequential().count()` runs sequentially — the last call wins. ## What runs the work By default parallel streams execute on the **common ForkJoinPool**, a single thread pool shared by the entire JVM. Its size defaults to (number of available processors − 1) worker threads, plus the submitting thread also helps, so total parallelism roughly equals the core count. ## When it pays off Splitting, scheduling threads, and merging results all cost time. Parallelism wins only when the data set is large and/or the per-element work is expensive enough that doing the work concurrently saves more time than the overhead costs. For small collections or trivial operations, a plain sequential stream is faster and simpler. Always measure rather than assume. ## Minimal example ```java List<Integer> nums = List.of(1, 2, 3, 4, 5); long evens = nums.parallelStream() .filter(n -> n % 2 == 0) .count(); ``` Here the work is far too small to benefit from parallelism — it is shown only to illustrate the API. Real candidates have thousands+ of elements and meaningful per-element cost.

  • If you write list.stream().parallel().map(f).sequential().forEach(g), does it run in parallel?
    No. The mode is whatever it was last set to when the terminal operation begins; sequential() is the last mode-setting call, so the entire pipeline runs sequentially.
  • Does parallelStream() guarantee a speedup?
    No. It only helps when there is enough data and enough per-element work to outweigh the splitting, thread-coordination, and merging overhead; otherwise it is often slower than a sequential stream.

saying these in an interview costs you the question

  • Thinking parallel is always faster — overhead often makes it slower for small/cheap work.
  • Believing parallel()/sequential() only affects the operation it is chained after (it sets the whole pipeline's mode).
  • Assuming each element gets its own thread — the source is split into a few chunks, not one-thread-per-element.

context