skip to content

When you build collections via the Streams API (collect, toList, groupingBy), can you still benefit from pre-sizing, and what are the limits of doing so?

level: seniorimportance: nice to knowfreq 22%

answer

  1. Result count unknown after filter/flatMap → no pre-size
  2. toList()/toMap collectors grow from default capacity
  3. Supplier overloads pick type, not size (zero-arg)
  4. Known size → manual sized loop or sized-supplier collector
  5. Stream.toList() is unmodifiable, can't size; don't fight it

basics

~20 s

Streams usually can't pre-size their result because the element count isn't known until the stream finishes — collectors append into a default-capacity collection that grows as usual. If you know the size and it's a hot path, build the collection manually with a pre-sized constructor, or use a collector supplier that creates a pre-sized container.

solid answer

~50 s

Most stream pipelines don't know how many elements will survive filtering, so terminal operations like collect(toList()) or Collectors.toMap append into a default-capacity collection that resizes as it fills — you get the same grow-and-copy/rehash overhead as a manual loop, with no place to pass a capacity. Collectors.toMap and groupingBy don't expose a sizing hint; even the supplier-based overloads only let you choose the map type, not its initial capacity (the lambda runs with no size argument). So for known-large results on hot paths, the practical options are: drop to a manual loop or forEach into a pre-sized collection; or use a collector whose supplier news up a pre-sized container when you can bound the size. Streams over sized sources (collections, arrays) do know the source size and can split well for parallelism, but the surviving result size after filter/flatMap is still unknown. The honest summary: stream collectors trade a bit of sizing control for expressiveness; pre-size manually only where profiling shows the resize churn matters.

code

java · 16 lines
java
// Result size unknown (filter) → collector grows from default capacity
List<String> filtered = items.stream()
    .filter(Item::active)
    .map(Item::name)
    .collect(Collectors.toList());   // resizes as it fills; no sizing hook

// Size KNOWN (pure map over a sized source) → pre-size manually
int n = items.size();
List<String> names = new ArrayList<>(n);          // one allocation
items.stream().map(Item::name).forEach(names::add); // no resize (sequential)

// Or a sized supplier when you can bound the count (JDK 19+)
Map<Long, Item> byId = items.stream().collect(Collectors.toMap(
    Item::id, i -> i,
    (a, b) -> a,
    () -> HashMap.newHashMap(n)));

go deeper

for a junior

Knows streams build a collection at the end and that you don't normally pass it a size.

for a middle

Explains that filter/flatMap make the result count unknown, so collectors grow from default capacity like an un-pre-sized loop.

for a senior

Knows the supplier overloads pick type not size, can pre-size via a manual loop or sized-supplier when the count is independently known, and weighs clarity vs the micro-gain.

for a principal

Reasons about parallel-collector merge allocation profiles, sets guidance on when to leave streams expressive vs drop to sized manual builds, and ties any such change to measured allocation hotspots.

## The Streams API in one paragraph The **Streams API** lets you express data processing as a pipeline: a **source** (collection, array, generator), zero or more **intermediate** operations (`filter`, `map`, `flatMap`…), and a **terminal** operation (`collect`, `forEach`, `reduce`…). A common terminal is `collect(...)`, which accumulates elements into a result container using a **Collector** (e.g. `Collectors.toList()`, `toMap`, `groupingBy`). ## Why collectors usually can't pre-size Pre-sizing requires knowing the **final element count** up front. In a stream, that count is generally **unknown until the pipeline finishes**, because intermediate operations change it: - `filter` removes an unpredictable fraction, - `flatMap` can multiply elements, - `distinct` deduplicates. So a collector like `toList()` starts with a **default-capacity** container and **grows it** as elements arrive — the exact grow-and-copy (list) or double-and-rehash (map) cycle you'd get from a manual loop without pre-sizing. There's simply no count to hand it. ### Even the 'supplier' overloads don't fix it `Collectors.toMap(keyFn, valFn, mergeFn, mapSupplier)` and `toCollection(supplier)` let you choose the *container type*, but the supplier is a **zero-argument** `Supplier` — it's invoked with no size information: ```java // You pick the type, but you cannot pass a size into the supplier from the stream. .collect(Collectors.toCollection(ArrayList::new)); ``` You *can* close over a known size yourself if you have one independently of the stream (see below), but the collector framework won't compute or pass it for you. `groupingBy` similarly has no per-group sizing hint. ## Where a size *is* known Sometimes you legitimately know (or bound) the result size: - A `map`-only pipeline (no `filter`/`flatMap`/`distinct`) over a sized source yields **exactly** the source size. - A business rule bounds the output (e.g. 'at most one row per user'). In those cases you can pre-size — but you usually have to step outside the standard collectors: ```java // Known size → pre-size manually, then fill from the stream int n = source.size(); List<R> out = new ArrayList<>(n); source.stream().map(this::transform).forEach(out::add); // no resize // Or a supplier that closes over the known size: Map<K,V> m = source.stream().collect(Collectors.toMap( Item::key, Item::value, (a, b) -> a, () -> HashMap.newHashMap(n))); // pre-sized backing map (JDK 19+) ``` Note `forEach(out::add)` after a `map` is fine, but using `forEach` to mutate shared state in a **parallel** stream is unsafe — use a manual sized loop or a proper concurrent/merging collector there. ## `toList()` and other quirks - `Stream.toList()` (JDK 16+) returns an **unmodifiable** list built once the count is known internally; you can't pass it a capacity, and it's already reasonably efficient — don't try to micro-optimize it. - `Collectors.toList()` returns a mutable `ArrayList` that grows from default capacity. - Parallel collectors merge per-thread partial containers, which has its own (often larger) allocation profile; pre-sizing intuition from the sequential case doesn't transfer cleanly. ## The limits and the judgment call Pre-sizing inside streams is **constrained by design**: the API favors composability over manual capacity control, and the surviving count is frequently genuinely unknown. So: - For **unknown-size** results, accept the default growth — it's already amortized-efficient — and don't contort the pipeline. - For **known-large** results on a **hot path** where profiling shows resize/rehash churn, prefer a manual pre-sized build (loop or sized supplier) over a plain collector. - Don't add sizing machinery speculatively; it complicates otherwise-clean stream code for usually-negligible gains. Measure first (JFR/async-profiler), then optimize the specific hotspot. ## Bottom line Standard stream collectors append into default-capacity containers and can't be told the result size, because that size usually isn't known until the stream ends. You benefit from pre-sizing in streams only when you independently know the count — and then typically by building the collection yourself or supplying a pre-sized container — and only where it's measurably worth the loss of stream clarity.

  • Why can't Collectors.toMap pre-size its result map for you?
    Because the number of surviving entries isn't known until the stream is fully consumed (filters/flatMaps change it), and the collector framework's supplier is a zero-argument factory with no size input. The collector accumulates into whatever the supplier produces, growing it as needed. You can only pre-size by supplying a container you've already sized using a count you know independently of the stream.
  • If a pipeline is just map over a sized source (no filter), is the result size known?
    Yes — a pure map preserves the element count, so the result equals the source size. There you could legitimately pre-size: build an ArrayList with new ArrayList<>(source.size()) and fill it, or use a sized supplier in a collector. The catch is that any filter/flatMap/distinct breaks this guarantee.

saying these in an interview costs you the question

  • Assuming Collectors.toList() pre-sizes from the source size automatically
  • Thinking toCollection/toMap's supplier receives the element count
  • Using forEach into shared state in a parallel stream to 'pre-size'
  • Contorting clean stream code to pre-size when the size is unknown or the path is cold

context