skip to content

Spliterator

Spliterator is the traversal abstraction underneath streams: tryAdvance for one element, trySplit for parallel decomposition, plus size estimates and characteristic flags. It is why some sources parallelize well and others do not.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the relationship between Spliterator, Collection.spliterator(), Stream, and StreamSupport?

level: middleimportance: should knowfreq 40%

answer

  1. Source → Spliterator → Stream → terminal op
  2. Collection.spliterator() is a Java 8 default factory
  3. stream()/parallelStream() are defined on top of spliterator()
  4. StreamSupport.stream(spliterator, parallel) bridges custom sources
  5. boolean = sequential vs parallel switch

basics

~20 s

A Spliterator is the low-level thing that walks (and can split) the elements of a source. Every Collection can make one via spliterator(), and Stream/parallelStream are built on top of it. StreamSupport.stream(...) is the helper that turns any Spliterator into a Stream.

solid answer

~40 s

Spliterator is the traversal/decomposition engine; Stream is the high-level functional pipeline you actually use. The bridge is: Collection.spliterator() (a Java 8 default method) produces a Spliterator for the collection, and Collection.stream()/parallelStream() are defined in terms of it. For your own data sources, StreamSupport.stream(Spliterator, boolean) wraps a Spliterator into a Stream, the boolean choosing sequential or parallel execution. So the layering is: source → Spliterator (how to traverse and split) → Stream (map/filter/collect) → terminal operation that drives tryAdvance/forEachRemaining (sequentially) or recursive trySplit into the common ForkJoinPool (in parallel). You normally never touch Spliterator directly; you call stream(). You reach for Spliterator/StreamSupport only when streaming a non-Collection source, like an Iterable or a custom generator.

go deeper

for a junior

Can say collections give you a Spliterator via spliterator() and that stream() uses it; StreamSupport makes a Stream from a Spliterator.

for a middle

Explains the source→Spliterator→Stream layering, that stream()/parallelStream() are defined on spliterator(), and uses StreamSupport.stream for custom sources.

for a senior

Describes sequential (tryAdvance) vs parallel (trySplit into ForkJoinPool) execution driven by the terminal op, and how characteristics tune it; knows the Iterable vs Collection spliterator() override.

for a principal

Reasons about where to inject a custom Spliterator in the layering, the parallel cost model, and designs library APIs that expose streams over non-Collection sources cleanly.

## The layered picture Think of it as four layers, from concrete data up to the API you write: ``` Data source (ArrayList, array, file, custom...) │ produces ▼ Spliterator ← LOW-LEVEL: how to traverse (tryAdvance) and split (trySplit) │ wrapped by ▼ Stream / IntStream ← HIGH-LEVEL: map/filter/reduce/collect pipeline │ driven by ▼ Terminal operation ← pulls elements: sequentially, or splits into ForkJoinPool ``` ## The pieces and how they connect **`Spliterator`** — the **low-level** mechanism. It knows *how to walk* a source one element at a time (`tryAdvance`/`forEachRemaining`) and *how to divide it* for parallel work (`trySplit`), plus metadata (`estimateSize`, `characteristics`). It is not something you usually call directly. **`Collection.spliterator()`** — a **default method** added to `Collection` in Java 8 (it actually lives on `Iterable`/`Collection`). Every collection can hand you a Spliterator describing itself, with the right characteristics (an `ArrayList` reports `ORDERED|SIZED|SUBSIZED`, a `HashSet` reports `DISTINCT|SIZED`, etc.). This is the standard factory. **`Collection.stream()` and `parallelStream()`** — the methods you actually call. They are **defined in terms of `spliterator()`**: internally they do roughly `StreamSupport.stream(spliterator(), false)` (or `true` for parallel). So a Stream *is* a high-level pipeline sitting on top of a Spliterator. **`StreamSupport`** — the bridge for **non-Collection** or **custom** sources. Its key method: ```java Stream<T> StreamSupport.stream(Spliterator<T> s, boolean parallel) ``` given any Spliterator, it produces a Stream; the boolean selects sequential vs parallel execution. This is how you stream things that aren't Collections — for example an `Iterable`: ```java public static <T> Stream<T> streamOf(Iterable<T> it) { return StreamSupport.stream(it.spliterator(), false); } ``` or a custom Spliterator over a file, a paged API, or a generator. ## What happens when the terminal operation runs - **Sequential** (`stream()`): the terminal op repeatedly calls `tryAdvance` (or `forEachRemaining`) on the single Spliterator, on the calling thread. - **Parallel** (`parallelStream()`): the framework recursively calls `trySplit` to build a tree of chunks, submits them to the **common ForkJoinPool**, runs `forEachRemaining` on each leaf in a worker thread, and combines results. The Spliterator's `characteristics` (SIZED, ORDERED, DISTINCT, …) tune this. ## Practical takeaways - For a normal collection, just call `.stream()` / `.parallelStream()` — never touch the Spliterator. - To stream an `Iterable`, an array section, or a custom source, get/build a Spliterator and pass it to `StreamSupport.stream(...)`. - `Arrays.stream(...)` and `IntStream.range(...)` are similar convenience wrappers that build Spliterators for you under the hood. - The boolean in `StreamSupport.stream` is your sequential/parallel switch — but parallel only pays off when the Spliterator splits well (see the trySplit question).

  • How do you turn an Iterable (not a Collection) into a Stream?
    Iterable has a default spliterator() method, so call StreamSupport.stream(iterable.spliterator(), false). Use true for a parallel stream, though Iterable's default Spliterator usually has unknown size and weak splitting, so parallel rarely helps.
  • Is spliterator() on Collection or Iterable?
    The default spliterator() is declared on Iterable (returning an unknown-size, ORDERED-by-default Spliterator), and Collection overrides it to provide a better one with accurate characteristics like SIZED.

saying these in an interview costs you the question

  • Thinking Stream and Spliterator are unrelated — Stream is built on Spliterator
  • Saying you must implement a Spliterator to use streams — collections give you one for free
  • Forgetting StreamSupport when streaming an Iterable/custom source
  • Believing the StreamSupport boolean controls ordering — it controls sequential vs parallel execution

context

open as a page

What is a Spliterator in Java, and how does it differ from a classic Iterator?

level: middleimportance: should knowfreq 55%

basics

~20 s

A Spliterator is an object that walks over the elements of a source (like a list) one at a time, and can also split itself in two so the halves can be processed in parallel. An Iterator can only walk forward, never split.

open as a page

What are Spliterator characteristics, and how does the Streams API use flags like ORDERED, SIZED, DISTINCT, SORTED, and IMMUTABLE?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Characteristics are flags a Spliterator reports about its data — for example that the elements are in order, that the exact count is known, that they are all unique, or that the source can't change. The stream pipeline reads these flags to skip unnecessary work.

open as a page

How does trySplit() work, and what makes a good split for parallel streams?

level: seniorimportance: should knowfreq 48%

basics

~20 s

trySplit() tries to hand off about half of the remaining elements to a brand-new Spliterator and keeps the rest for itself, so two threads can work on the two halves at once. If it can't usefully split, it returns null.

open as a page

How would you write a custom Spliterator, and when is doing so actually worth it?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

You implement tryAdvance (process one element), trySplit (hand off about half for parallelism), estimateSize (how many are left), and characteristics (the flags). You only bother when you have a custom data source that the built-in Collection/Stream tools can't traverse or parallelize well.

open as a page