skip to content

When is pre-sizing collections a premature or counterproductive optimization, and how would you decide where it's actually worth applying?

level: seniorimportance: should knowfreq 35%

answer

  1. Win is constant-factor + GC, not big-O
  2. Worth it: known size + large + hot path
  3. Over-estimate wastes heap + cache locality
  4. Capacity literal = magic number to maintain
  5. Measure (JFR/async-profiler/JMH) before applying

basics

~20 s

Pre-sizing only helps when the final size is known and the code is hot. If sizes are small, the path is cold, or the size is unknown, the constructor argument adds clutter without measurable benefit — and a bad over-estimate wastes memory. Measure first; apply it on proven hot, known-size build paths.

solid answer

~50 s

Pre-sizing is a constant-factor, low-risk optimization, but it's not free of judgment. It pays off when three things hold: the final element count is known or tightly estimable, the collection is large, and the build path runs often or is latency-sensitive. Outside that, it's noise: a 5-element list resizing once costs nothing, and a wild over-estimate (sizing for millions when you have dozens) wastes heap, hurts cache locality, and can pressure the GC more than the resizes it avoided. The danger is treating it as a blanket rule and sprinkling capacity arguments everywhere, which adds clutter and a hidden 'magic number' to maintain. The disciplined approach is to default to readable code, profile to find real allocation/GC hotspots (JFR, async-profiler, allocation flame graphs), and pre-size the specific hot, known-size builders that show up. On JDK 19+, prefer HashMap.newHashMap(n) so the intent is explicit and the math can't be wrong.

go deeper

for a junior

Understands pre-sizing helps for big known-size collections and that tiny lists don't benefit.

for a middle

Can list the conditions (known size, large, hot) and notes that over-estimating wastes memory.

for a senior

Frames it as a constant-factor/GC optimization, warns against blanket application and magic-number capacities, and ties the decision to profiling evidence.

for a principal

Sets team norms (prefer factory helpers, justify with data), reasons about allocation-pressure budgets and GC-pause SLAs, and balances micro-optimization effort against readability and maintenance across the codebase.

## The optimization and its real cost **Pre-sizing** means passing an expected capacity to a collection constructor so its backing storage is allocated once instead of growing repeatedly (see the companion questions for the mechanics). It removes the **constant-factor** overhead of grow-and-copy (ArrayList ~1.5x) and double-and-rehash (HashMap), plus the garbage those produce. Crucially it does **not** change asymptotic complexity — building a list is O(n) either way — so the win is in constant factors and **allocation/GC pressure**, not big-O. ## Why 'always pre-size' is wrong Treating pre-sizing as a universal rule misapplies effort and can backfire: 1. **Small collections:** A list that holds 3 or 30 elements resizes a handful of times at most, copying a tiny array. The cost is in the nanoseconds and utterly dominated by everything else. Adding `new ArrayList<>(30)` here buys nothing measurable. 2. **Cold paths:** Code that runs once at startup, or rarely, has no hot-loop cost to save. Optimizing it trades readability for zero throughput. 3. **Unknown size:** If you can't estimate the count, you either guess (risking a bad estimate) or you can't pre-size meaningfully. Default growth is already efficient — don't force a number. 4. **Over-estimation harms:** Sizing for the worst case when the common case is tiny allocates a large, mostly-empty array on **every** call. That wastes heap, worsens **cache locality** (a sparse big array touches more memory pages), and can create *more* GC pressure than the resizes you avoided. An over-sized `HashMap` also spreads entries across more buckets than needed. 5. **Maintenance cost:** A literal capacity is a **magic number**. If the data's typical size shifts later, the hint silently becomes wrong (too small → resizes return; too big → waste). It's one more thing to keep honest. 6. **The Knuth caveat:** 'Premature optimization is the root of all evil.' Micro-optimizing before measuring spends your attention on code that isn't the bottleneck and obscures intent. ## A decision procedure Apply pre-sizing when **all three** are true, ideally confirmed by data: - **Known/estimable size:** you have the count (e.g. `list.size()` of a source, a query row count, a batch size) or a tight upper bound. - **Large enough to matter:** thousands-plus elements, where the resize/rehash work and garbage are non-trivial. - **Hot or latency-sensitive:** the builder runs in a tight loop, per-request, or on a path with a latency SLA where GC pauses hurt. ### How to find those spots Don't guess — **measure**: - **Allocation profiling:** Java Flight Recorder (JFR), async-profiler in allocation mode, or VisualVM to see which call sites allocate the most (resizes show up as collection-internal array allocations). - **GC logs / pause analysis:** if collection churn is driving young-gen GC, the throwaway arrays from resizing are a visible contributor. - **Microbenchmarks (JMH):** to validate a specific hot builder actually improves, including warm-up so the JIT has compiled the path. ## Prefer expressive forms When you do pre-size, make intent obvious and the math safe: - `new ArrayList<>(source)` / `new ArrayList<>(expectedSize)`. - `HashMap.newHashMap(n)` / `HashSet.newHashSet(n)` (JDK 19+) instead of hand-rolling `n/0.75+1`. - A named constant or a comment explaining where the size comes from, so the 'magic number' is justified. ## Bottom line Pre-sizing is a sharp, cheap tool for the **specific** case of building a large collection of known size on a hot path. As a blanket habit it adds clutter, risks wasteful over-allocation, and distracts from real bottlenecks. Default to clear code, profile, and pre-size the few places that earn it.

  • How would you justify a specific pre-sizing change in a code review?
    Point to evidence: this builder is on a per-request hot path, the size is known (e.g. equals the source row count), it's large (thousands of entries), and a profile/benchmark shows the resize allocations were a measurable fraction of churn. Then show the change uses an expressive form (HashMap.newHashMap(n)) with the size source documented, so it's not a stray magic number.
  • Could pre-sizing ever increase GC pressure?
    Yes, if you over-estimate badly. Allocating a large backing array on a frequently-called path puts a big object into the heap each time; if most are barely used, you've created more allocation volume than the small resizes you avoided. The fix is to size to a realistic typical/upper bound, not the absolute worst case.

saying these in an interview costs you the question

  • Claiming pre-sizing should be applied to every collection unconditionally
  • Believing it improves algorithmic complexity rather than constant factors
  • Ignoring that a large over-estimate can hurt memory/GC more than it helps
  • Optimizing without profiling — guessing at the bottleneck

context