skip to content

What are the three Collector characteristics (CONCURRENT, UNORDERED, IDENTITY_FINISH), and how does each affect parallel collection?

level: seniorimportance: should knowfreq 45%

answer

  1. IDENTITY_FINISH = skip finisher, cast container
  2. UNORDERED = order doesn't matter, enables concurrent path
  3. CONCURRENT = shared single container across threads
  4. concurrent path needs CONCURRENT + parallel + unordered
  5. groupingByConcurrent / toConcurrentMap are CONCURRENT+UNORDERED

basics

~20 s

They are hints a Collector gives the stream. CONCURRENT means all threads can safely share one container. UNORDERED means element order doesn't matter so the stream can skip preserving it. IDENTITY_FINISH means the accumulation container is already the final result, so the finisher can be skipped.

solid answer

~50 s

A Collector advertises a set of characteristics that let the stream optimize. CONCURRENT means the accumulator can be called concurrently on a single shared result container from multiple threads — the stream can use a concurrent reduction instead of split-and-merge, but only if the stream is parallel and (for non-UNORDERED collectors) the stream is also unordered. UNORDERED means the collector doesn't depend on encounter order, so the stream may relax ordering for speed and is free to use the concurrent path. IDENTITY_FINISH means the finisher is the identity function, so the framework can skip calling it and cast the accumulation container directly to the result type. The big interaction: a CONCURRENT collector only actually performs a concurrent reduction (shared container, no merging) when the stream is parallel AND unordered — otherwise it falls back to the ordinary split/accumulate/combine strategy. Collectors.toConcurrentMap and groupingByConcurrent are CONCURRENT+UNORDERED; toList is IDENTITY_FINISH; toSet is UNORDERED+IDENTITY_FINISH.

go deeper

for a junior

May not know characteristics exist; can still use toList/toSet without them.

for a middle

Can name the three characteristics and roughly what each means; knows toConcurrentMap exists for parallel work.

for a senior

Explains each characteristic precisely and the CONCURRENT+parallel+unordered triple gate for the concurrent path; knows which standard collectors carry which flags.

for a principal

Reasons about when concurrent collection actually pays off (contention vs merge cost), audits custom collectors' characteristics for correctness, and knows the failure modes of mis-declaring CONCURRENT.

## What characteristics are A **Collector** describes a mutable reduction. Besides its supplier/accumulator/combiner/finisher functions, it returns a `Set<Collector.Characteristics>` from `characteristics()`. These are **optimization hints** the stream framework reads to decide *how* to run the collection. There are exactly three flags. ### IDENTITY_FINISH The **finisher** is the function that turns the intermediate accumulation type `A` into the final result type `R`. When `A` and `R` are the same type and no transform is needed, the finisher is the identity function. Declaring **IDENTITY_FINISH** tells the framework: *don't bother calling the finisher; just cast the accumulated container to `R`.* This saves a function call and is purely an optimization. `Collectors.toList()` is IDENTITY_FINISH because the `ArrayList` it accumulates into *is* the returned `List`. `Collectors.joining()` is **not** IDENTITY_FINISH — it accumulates into a `StringBuilder` (type `A`) but finishes with `.toString()` to produce a `String` (type `R`). ### UNORDERED Many collections don't care about **encounter order** — the order in which elements appear in the source. A `Set` doesn't; a sum doesn't; a `Map` (by key) doesn't. Declaring **UNORDERED** tells the framework the result is the same regardless of element order, so on a parallel stream it may skip the work of preserving order, which is faster. `Collectors.toSet()` is UNORDERED; `Collectors.toList()` is **not** (a List preserves order). ### CONCURRENT Normally, parallel collection uses **split-and-merge**: the stream splits into chunks, each thread builds its *own* container (one supplier call per chunk), and the combiner merges them pairwise. **CONCURRENT** declares that the accumulator is safe to call from **multiple threads on the same single shared container** — so the framework can do a **concurrent reduction**: call the supplier *once*, share that container, and have all threads accumulate into it directly, with **no merging step**. This requires a thread-safe container (e.g. `ConcurrentHashMap`). `Collectors.toConcurrentMap()` and `Collectors.groupingByConcurrent()` are CONCURRENT. ## The crucial interaction CONCURRENT alone is **not** enough to trigger the concurrent path. The framework uses concurrent reduction only when **all** of these hold: 1. the collector is **CONCURRENT**, **and** 2. the stream is **parallel**, **and** 3. the stream is **unordered** — either the collector is also **UNORDERED** or the stream pipeline itself is unordered (e.g. `.unordered()` was called, or the source has no defined encounter order). Why the unordered requirement? Sharing one container across threads loses any guarantee about the order in which elements are inserted. If encounter order mattered, a concurrent reduction could produce a differently-ordered result, so the framework refuses it and falls back to ordered split-and-merge. That is why the standard concurrent collectors are declared **CONCURRENT + UNORDERED** together. If any condition fails (sequential stream, or ordered, or non-concurrent collector), the framework silently uses the ordinary supplier-per-chunk + combiner strategy. A CONCURRENT collector still works correctly there — it just doesn't get the shared-container speedup. ## Worked example of consequences ```java // CONCURRENT + UNORDERED: shares one ConcurrentHashMap across threads when parallel Map<Dept, List<Emp>> m = emps.parallelStream() .collect(Collectors.groupingByConcurrent(Emp::dept)); // not CONCURRENT: each thread builds a HashMap, then they are merged Map<Dept, List<Emp>> m2 = emps.parallelStream() .collect(Collectors.groupingBy(Emp::dept)); ``` With `groupingByConcurrent`, all threads insert into one `ConcurrentHashMap`. With plain `groupingBy`, each thread builds its own `HashMap` and the combiner merges them, which can dominate cost for large maps. ## Summary table - IDENTITY_FINISH → skip the finisher (micro-optimization). - UNORDERED → may relax order; required for the concurrent path. - CONCURRENT → may share one container across threads, but only when parallel + unordered. Knowing these explains *why* `toConcurrentMap` can beat `toMap` under parallelism and why a hand-written three-arg `collect` (empty characteristics) never gets the concurrent path.

  • Why must a stream be unordered for a CONCURRENT collector to use its concurrent reduction?
    Sharing one container across threads gives no control over insertion order, so encounter order can't be guaranteed. The framework only takes the concurrent path when order is irrelevant — i.e. the collector or stream is unordered — otherwise it falls back to ordered split-and-merge.
  • Is Collectors.joining() IDENTITY_FINISH? Why or why not?
    No. It accumulates into a StringBuilder (type A) and its finisher calls toString() to produce a String (type R), so A != R and a real finisher runs.

saying these in an interview costs you the question

  • Thinking CONCURRENT alone makes the collection run concurrently — it also needs a parallel and unordered stream
  • Believing UNORDERED reorders or shuffles results (it only permits the framework to not preserve order)
  • Assuming IDENTITY_FINISH changes the result — it only skips a redundant finisher call
  • Sharing a non-thread-safe container under a CONCURRENT collector

context