skip to content

What is a Spliterator in Java, and how does it differ from a classic Iterator?

level: middleimportance: should knowfreq 55%

answer

  1. Split + iterator = parallel-friendly traversal
  2. tryAdvance fuses hasNext+next into one call
  3. trySplit hands off ~half the elements
  4. estimateSize + characteristics = metadata for the pipeline
  5. engine under Streams; Collection.spliterator()

basics

~20 s

A Spliterator is an object that walks over the elements of a source (like a list) one at a time, and can also split itself in two so the halves can be processed in parallel. An Iterator can only walk forward, never split.

solid answer

~40 s

Spliterator ("splittable iterator"), added in Java 8, is the traversal engine behind the Streams API. Like an Iterator it visits elements of a source, but it adds two things. First, splitting: trySplit() hands off roughly half the remaining elements to a new Spliterator, enabling parallel decomposition for fork/join. Second, it exposes metadata the framework optimizes on: estimateSize() (how many elements remain) and characteristics flags (ORDERED, SIZED, DISTINCT, SORTED, IMMUTABLE, etc.). Its traversal methods are tryAdvance (one element, returns false when exhausted) and forEachRemaining (bulk). Compared to Iterator's hasNext()/next() two-call model, tryAdvance fuses the check-and-fetch into one call, which is cheaper and avoids the inconsistent-state window. Every Collection can produce one via spliterator(); stream() is built on top of it.

go deeper

for a junior

Can say a Spliterator walks elements like an Iterator but can also split itself for parallel work, and that streams use it.

for a middle

Names tryAdvance/trySplit/estimateSize/characteristics, explains the one-call traversal vs hasNext/next, and that Collection.spliterator() backs stream().

for a senior

Articulates why splitting plus metadata enables the fork/join parallel pipeline, when trySplit returns null, and the primitive OfInt/OfLong/OfDouble specializations.

for a principal

Reasons about balanced-split quality, late-binding vs early-binding semantics, fail-fast behavior, and the cost trade-offs that make custom Spliterators worth writing for a particular source.

## The problem Spliterator solves Before Java 8, the standard way to traverse a collection was the `Iterator`: ```java Iterator<T> it = list.iterator(); while (it.hasNext()) { // 1) check T x = it.next(); // 2) fetch } ``` This works for sequential, single-threaded loops, but it has two limitations that matter for the Streams API and parallelism: 1. **It cannot be divided.** An `Iterator` only moves forward, one element at a time, from front to back. There is no way to say "you take the first half, I'll take the second half" — which is exactly what you need to spread work across multiple CPU cores. 2. **It carries no metadata.** The framework cannot ask "how many elements are left?" or "are these already sorted/distinct?" so it cannot make smart optimization decisions. A **`Spliterator`** (the name is a portmanteau of *split* + *iterator*) was introduced in Java 8 (`java.util.Spliterator`) to fix both. It is the low-level traversal mechanism that the **Streams API** is built on top of. ## The two ways it traverses - **`boolean tryAdvance(Consumer<? super T> action)`** — process exactly **one** remaining element by passing it to `action`, then return `true`. If no elements remain, do nothing and return `false`. This is the single-element step. Notice it combines the *check* ("is there an element?") and the *fetch* ("give it to me") into **one** method call, unlike `Iterator`'s separate `hasNext()`/`next()`. - **`void forEachRemaining(Consumer<? super T> action)`** — process **all** remaining elements in bulk. The default implementation just loops on `tryAdvance`, but a concrete Spliterator can override it with a tighter loop (e.g. a plain array index walk), which is why bulk traversal is often faster. ## The splitting ability — `trySplit` `Spliterator<T> trySplit()` is the defining feature. It attempts to **partition** the elements: it removes roughly the first half of the remaining elements from *this* Spliterator and returns a **new** Spliterator covering them; `this` keeps the rest. If splitting is not worthwhile or not possible (too few elements, or an inherently sequential source), it returns `null`. This lets the parallel-stream machinery recursively split a source into chunks, hand each chunk to a worker thread in the common ForkJoinPool, and combine the results — *without the source having to know anything about threads*. A well-behaved `trySplit` produces balanced halves so the work spreads evenly. ## Metadata: `estimateSize` and characteristics - **`long estimateSize()`** — an estimate of how many elements remain (or `Long.MAX_VALUE` if unknown). The framework uses it to decide whether splitting further is worth the overhead and to size intermediate buffers. - **`int characteristics()`** — a bitmask of flags describing the source (covered in its own question): `ORDERED`, `SIZED`, `SUBSIZED`, `DISTINCT`, `SORTED`, `NONNULL`, `IMMUTABLE`, `CONCURRENT`. They let the stream pipeline skip work — e.g. `distinct()` on an already-`DISTINCT` source is a no-op. ## How it relates to the rest of the platform Every `Collection` gained a `default Spliterator<E> spliterator()` method in Java 8, and `Collection.stream()` / `parallelStream()` are defined in terms of it. You rarely write or call a Spliterator by hand; you use streams, and the Spliterator is the engine underneath. `StreamSupport.stream(spliterator, parallel)` is the bridge for turning a custom Spliterator into a Stream. There are also primitive specializations — `Spliterator.OfInt`, `OfLong`, `OfDouble` — that avoid boxing for `IntStream` and friends. ## Summary of the contrast | | Iterator | Spliterator | |---|---|---| | Traverse one | `hasNext()`+`next()` (two calls) | `tryAdvance()` (one call) | | Traverse all | manual loop | `forEachRemaining()` | | Split for parallelism | impossible | `trySplit()` | | Size hint | none | `estimateSize()` | | Metadata | none | `characteristics()` | | Introduced | Java 1.2 | Java 8 |

  • Why does tryAdvance return a boolean instead of having a separate hasNext?
    Fusing the check and the fetch into one call is cheaper (one virtual call, not two) and removes the window where the iterator is in a half-advanced state, which simplifies correct implementation and improves performance for the stream pipeline.
  • How do you turn a custom Spliterator into a Stream?
    Call StreamSupport.stream(spliterator, parallel), where the boolean chooses sequential vs parallel execution.

saying these in an interview costs you the question

  • Saying Spliterator replaces Iterator everywhere — Iterator is still fine for plain sequential loops
  • Claiming trySplit always succeeds — it returns null when splitting is not worthwhile
  • Thinking you normally call Spliterator methods by hand — you use streams, it is the engine underneath
  • Confusing it with parallelStream itself — Spliterator is the decomposition mechanism, not the executor

context