skip to content

What is the relationship between Spliterator, Collection.spliterator(), Stream, and StreamSupport?

level: middleimportance: should knowfreq 40%

answer

  1. Source → Spliterator → Stream → terminal op
  2. Collection.spliterator() is a Java 8 default factory
  3. stream()/parallelStream() are defined on top of spliterator()
  4. StreamSupport.stream(spliterator, parallel) bridges custom sources
  5. boolean = sequential vs parallel switch

basics

~20 s

A Spliterator is the low-level thing that walks (and can split) the elements of a source. Every Collection can make one via spliterator(), and Stream/parallelStream are built on top of it. StreamSupport.stream(...) is the helper that turns any Spliterator into a Stream.

solid answer

~40 s

Spliterator is the traversal/decomposition engine; Stream is the high-level functional pipeline you actually use. The bridge is: Collection.spliterator() (a Java 8 default method) produces a Spliterator for the collection, and Collection.stream()/parallelStream() are defined in terms of it. For your own data sources, StreamSupport.stream(Spliterator, boolean) wraps a Spliterator into a Stream, the boolean choosing sequential or parallel execution. So the layering is: source → Spliterator (how to traverse and split) → Stream (map/filter/collect) → terminal operation that drives tryAdvance/forEachRemaining (sequentially) or recursive trySplit into the common ForkJoinPool (in parallel). You normally never touch Spliterator directly; you call stream(). You reach for Spliterator/StreamSupport only when streaming a non-Collection source, like an Iterable or a custom generator.

go deeper

for a junior

Can say collections give you a Spliterator via spliterator() and that stream() uses it; StreamSupport makes a Stream from a Spliterator.

for a middle

Explains the source→Spliterator→Stream layering, that stream()/parallelStream() are defined on spliterator(), and uses StreamSupport.stream for custom sources.

for a senior

Describes sequential (tryAdvance) vs parallel (trySplit into ForkJoinPool) execution driven by the terminal op, and how characteristics tune it; knows the Iterable vs Collection spliterator() override.

for a principal

Reasons about where to inject a custom Spliterator in the layering, the parallel cost model, and designs library APIs that expose streams over non-Collection sources cleanly.

## The layered picture Think of it as four layers, from concrete data up to the API you write: ``` Data source (ArrayList, array, file, custom...) │ produces ▼ Spliterator ← LOW-LEVEL: how to traverse (tryAdvance) and split (trySplit) │ wrapped by ▼ Stream / IntStream ← HIGH-LEVEL: map/filter/reduce/collect pipeline │ driven by ▼ Terminal operation ← pulls elements: sequentially, or splits into ForkJoinPool ``` ## The pieces and how they connect **`Spliterator`** — the **low-level** mechanism. It knows *how to walk* a source one element at a time (`tryAdvance`/`forEachRemaining`) and *how to divide it* for parallel work (`trySplit`), plus metadata (`estimateSize`, `characteristics`). It is not something you usually call directly. **`Collection.spliterator()`** — a **default method** added to `Collection` in Java 8 (it actually lives on `Iterable`/`Collection`). Every collection can hand you a Spliterator describing itself, with the right characteristics (an `ArrayList` reports `ORDERED|SIZED|SUBSIZED`, a `HashSet` reports `DISTINCT|SIZED`, etc.). This is the standard factory. **`Collection.stream()` and `parallelStream()`** — the methods you actually call. They are **defined in terms of `spliterator()`**: internally they do roughly `StreamSupport.stream(spliterator(), false)` (or `true` for parallel). So a Stream *is* a high-level pipeline sitting on top of a Spliterator. **`StreamSupport`** — the bridge for **non-Collection** or **custom** sources. Its key method: ```java Stream<T> StreamSupport.stream(Spliterator<T> s, boolean parallel) ``` given any Spliterator, it produces a Stream; the boolean selects sequential vs parallel execution. This is how you stream things that aren't Collections — for example an `Iterable`: ```java public static <T> Stream<T> streamOf(Iterable<T> it) { return StreamSupport.stream(it.spliterator(), false); } ``` or a custom Spliterator over a file, a paged API, or a generator. ## What happens when the terminal operation runs - **Sequential** (`stream()`): the terminal op repeatedly calls `tryAdvance` (or `forEachRemaining`) on the single Spliterator, on the calling thread. - **Parallel** (`parallelStream()`): the framework recursively calls `trySplit` to build a tree of chunks, submits them to the **common ForkJoinPool**, runs `forEachRemaining` on each leaf in a worker thread, and combines results. The Spliterator's `characteristics` (SIZED, ORDERED, DISTINCT, …) tune this. ## Practical takeaways - For a normal collection, just call `.stream()` / `.parallelStream()` — never touch the Spliterator. - To stream an `Iterable`, an array section, or a custom source, get/build a Spliterator and pass it to `StreamSupport.stream(...)`. - `Arrays.stream(...)` and `IntStream.range(...)` are similar convenience wrappers that build Spliterators for you under the hood. - The boolean in `StreamSupport.stream` is your sequential/parallel switch — but parallel only pays off when the Spliterator splits well (see the trySplit question).

  • How do you turn an Iterable (not a Collection) into a Stream?
    Iterable has a default spliterator() method, so call StreamSupport.stream(iterable.spliterator(), false). Use true for a parallel stream, though Iterable's default Spliterator usually has unknown size and weak splitting, so parallel rarely helps.
  • Is spliterator() on Collection or Iterable?
    The default spliterator() is declared on Iterable (returning an unknown-size, ORDERED-by-default Spliterator), and Collection overrides it to provide a better one with accurate characteristics like SIZED.

saying these in an interview costs you the question

  • Thinking Stream and Spliterator are unrelated — Stream is built on Spliterator
  • Saying you must implement a Spliterator to use streams — collections give you one for free
  • Forgetting StreamSupport when streaming an Iterable/custom source
  • Believing the StreamSupport boolean controls ordering — it controls sequential vs parallel execution

context