What is a Spliterator, and what role does it play as the underlying source of every Java stream?
answer
- Spliterator = splittable iterator under every stream
- tryAdvance (traverse) + trySplit (partition)
- characteristics: SIZED/ORDERED/SORTED/DISTINCT/NONNULL
- Split quality → parallel performance (ArrayList vs LinkedList/iterate)
- StreamSupport.stream(spliterator, parallel) to build one
basics
~20 sA Spliterator is the object that actually feeds a stream: it knows how to traverse elements one by one (tryAdvance) and how to split itself into chunks (trySplit) so a parallel stream can process pieces on different threads.
solid answer
~50 sEvery stream is backed by a Spliterator ("splittable iterator") — the abstraction that describes how the source is traversed and partitioned. Its two core methods are tryAdvance(consumer), which processes the next element if any (like Iterator.hasNext+next fused), and trySplit(), which tries to hand off a portion of the remaining elements as a second Spliterator, enabling parallel decomposition. It also reports characteristics — flags like SIZED, ORDERED, SORTED, DISTINCT, NONNULL, IMMUTABLE — that let the framework optimize: a SIZED source can presize collectors and split evenly; ORDERED constrains result order; SORTED can skip a sort. estimateSize() helps balance splits. Most users never write one — Collection, arrays, and ranges supply good Spliterators — but you can wrap any Iterator via Spliterators.spliterator(...) and build a stream with StreamSupport.stream(spliterator, parallel). Understanding it explains why some streams parallelize well (sized, splittable, like ArrayList/range) and others poorly (iterate, LinkedList, I/O).
code
java · 15 linesimport java.util.*;
import java.util.stream.*;
// Adapt a legacy Iterator into a (sequential) stream via its Spliterator:
Iterator<String> legacy = List.of("a", "b", "c").iterator();
Spliterator<String> sp =
Spliterators.spliteratorUnknownSize(legacy, Spliterator.ORDERED);
Stream<String> stream = StreamSupport.stream(sp, /* parallel = */ false);
System.out.println(stream.map(String::toUpperCase).toList()); // [A, B, C]
// Inspect characteristics of a SIZED source:
Spliterator<Integer> arrSp = List.of(1, 2, 3, 4).spliterator();
System.out.println(arrSp.hasCharacteristics(Spliterator.SIZED)); // true
System.out.println(arrSp.estimateSize()); // 4
Spliterator<Integer> half = arrSp.trySplit(); // splits off a prefix for parallel workgo deeper
Aware that a stream has an underlying source object and doesn't need to implement it.
Can name Spliterator's tryAdvance and trySplit and that it enables parallel splitting, unlike a plain Iterator.
Explains characteristics (SIZED/ORDERED/SORTED/DISTINCT) and how they plus split cost determine parallel performance across ArrayList/LinkedList/iterate/I-O sources, and wraps Iterators via StreamSupport.
Designs custom Spliterators for novel sources with correct characteristics and balanced trySplit, and reasons about when parallelism pays given source decomposition, ordering constraints, and contention.
## The thing under every stream When you write `list.stream()` or `IntStream.range(0, 100)`, something has to actually *produce* the elements and, for parallel streams, *divide the work*. That something is a **`Spliterator`** (`java.util.Spliterator`) — the word is a portmanteau of **splittable + iterator**. Every stream is created from a Spliterator; the stream API is essentially a fluent layer over it. ## Why not just an Iterator? A classic `Iterator` only does sequential traversal (`hasNext()` / `next()`). That's fine for a `for-each` loop, but it **cannot be parallelized** — there's no way to say "you take the first half, I'll take the second." Streams needed an abstraction that supports *both* sequential traversal *and* partitioning. Hence Spliterator. ## The two core methods 1. **`boolean tryAdvance(Consumer<T> action)`** — if an element remains, perform `action` on it and return `true`; otherwise return `false`. It fuses `hasNext` and `next` into one call (and supports `null` elements cleanly). There's also `forEachRemaining` for bulk sequential traversal. 2. **`Spliterator<T> trySplit()`** — attempt to split off a *prefix* of the remaining elements into a **new** Spliterator and return it (the original keeps the rest); return `null` if it can't or won't split. This is the heart of parallelism: the framework recursively calls `trySplit` to break the source into chunks that worker threads process independently. A good split is roughly *even*; a `null` from `trySplit` means "can't divide further — process sequentially." ## Characteristics: hints that drive optimization A Spliterator reports a bitmask of **characteristics** via `characteristics()`: - **`SIZED`** — the exact element count is known (`estimateSize()` is exact). Lets collectors presize arrays and lets the splitter divide evenly. Ranges and `ArrayList` are SIZED; `Stream.iterate`, a `HashSet`-filtered stream, or a file stream are not. - **`SUBSIZED`** — every child from `trySplit` is also SIZED. - **`ORDERED`** — elements have a meaningful encounter order (lists, arrays); affects whether operations must preserve order. - **`SORTED`** — elements come pre-sorted by a known order; a `sorted()` step can be skipped. - **`DISTINCT`** — no duplicates (e.g. from a `Set`); `distinct()` becomes a no-op. - **`NONNULL`** — no null elements. - **`IMMUTABLE` / `CONCURRENT`** — whether the source can change during traversal; governs fail-fast vs concurrent behavior. `estimateSize()` returns the (possibly estimated) remaining count, used to decide whether further splitting is worthwhile. ## Why this explains parallel performance The practical payoff: **how well a stream parallelizes is governed by its Spliterator.** - `ArrayList` / arrays / `IntStream.range`: SIZED + cheap, even `trySplit` → **excellent** parallel scaling. - `LinkedList`: must walk nodes to split → **poor** splitting. - `Stream.iterate`: each element depends on the previous, so it can't hand off an independent chunk → essentially **unsplittable**. - `Files.lines` / `BufferedReader.lines`: I/O-bound, not SIZED, splits awkwardly → **limited** parallel benefit. So when someone asks "why is my `list.parallelStream()` slow," the answer often lies in the underlying Spliterator's characteristics and split cost. ## Using and building one You rarely write a Spliterator, because `Collection.spliterator()`, `Arrays.spliterator()`, and the range factories supply solid ones. But you can: - Adapt an `Iterator`: `Spliterators.spliterator(iterator, size, characteristics)` or `spliteratorUnknownSize(iterator, characteristics)`. - Turn a Spliterator into a stream: `StreamSupport.stream(spliterator, parallel)`. ```java Iterator<String> it = legacyApi(); Spliterator<String> sp = Spliterators.spliteratorUnknownSize(it, Spliterator.ORDERED); Stream<String> stream = StreamSupport.stream(sp, false); ``` Writing a custom Spliterator is the right move when you have a source the JDK doesn't model and you want it to parallelize well — you implement `tryAdvance`, `trySplit`, `estimateSize`, and accurate `characteristics`. ## Mental model An `Iterator` is a one-lane road (sequential only). A `Spliterator` is a road that can **fork**: it still drives forward one element at a time, but at any point it can split into two roads for two drivers. The honesty of its characteristics (am I SIZED? ORDERED? DISTINCT?) is what lets the stream engine plan the fastest route.
- Why does a parallel stream over an ArrayList usually outperform one over a LinkedList?ArrayList's Spliterator is SIZED and array-backed, so trySplit cheaply divides the index range evenly; LinkedList must traverse nodes to split, producing uneven, costly splits with poor parallel scaling.
- How do you build a stream from a legacy Iterator?Wrap it with Spliterators.spliteratorUnknownSize(iterator, characteristics) (or spliterator(it, size, characteristics) if the count is known), then pass it to StreamSupport.stream(spliterator, false/true).
- What does the SORTED characteristic let the pipeline do?If the source reports SORTED with a matching comparator, a subsequent sorted() step can be elided as a no-op, saving the sort.
saying these in an interview costs you the question
- Calling a Spliterator just an Iterator — it adds splitting (trySplit) for parallelism.
- Assuming every stream parallelizes equally well regardless of source.
- Thinking you must write a Spliterator to use streams (the JDK supplies them).
- Reporting wrong characteristics (e.g. SIZED when you can't size) — it corrupts optimization and results.
- Believing trySplit must always split — returning null (sequential) is valid.