skip to content

How do you efficiently build a large String inside a loop, and what role does pre-sizing the StringBuilder's capacity play?

level: middleimportance: must knowfreq 60%

answer

  1. create → append in loop → toString once
  2. default capacity 16; grow ≈ doubles + copies
  3. amortized O(1) append → O(n) loop
  4. new StringBuilder(expectedLen) avoids grows/garbage
  5. String.join / Collectors.joining for simple joins
  6. StringBuffer = synchronized; default to StringBuilder

basics

~20 s

Use a StringBuilder: create one, call append in the loop, then toString at the end. If you know roughly how big the result will be, give the StringBuilder that size up front so it doesn't have to keep growing and copying its internal buffer.

solid answer

~50 s

Replace the += accumulation with a single StringBuilder: build it once before the loop, append each piece inside, and call toString() after. That turns O(n^2) into O(n) because append writes into a growable char[] in amortized O(1). Pre-sizing matters because the buffer has a default capacity (16); when it fills, StringBuilder allocates a larger array (roughly double) and copies everything over. Those resizes are amortized O(1) overall, but if you can estimate the final length, passing it to new StringBuilder(capacity) avoids the intermediate grow-and-copy cycles and the garbage they create, shaving real time on hot paths. Alternatives for simple cases: String.join, a Stream with Collectors.joining, or StringBuilder reuse. Note StringBuffer is the synchronized (thread-safe) sibling — avoid it unless the builder is genuinely shared across threads, since its locking costs you for nothing in the common single-threaded case.

go deeper

for a junior

Can convert a += loop to create/append/toString with StringBuilder and knows toString goes after the loop.

for a middle

Explains the growable buffer, default capacity 16, amortized O(1) append, and when pre-sizing helps; knows StringBuilder vs StringBuffer.

for a senior

Reasons about geometric growth → amortized O(1), estimates capacity from the data, and picks String.join/Collectors.joining where they read better.

for a principal

Treats it as a constant-factor/allocation optimization to justify with profiling, weighs readability vs micro-optimization, and sets team conventions for builders vs joiners.

## The problem being solved Concatenating with `+=` in a loop is **O(n^2)** because String is immutable, so each step re-copies the entire accumulated text. The standard remedy is `StringBuilder`, a *mutable* string accumulator. ## StringBuilder mechanics Internally a `StringBuilder` holds: - a backing `char[]` buffer (its **capacity** = how many chars it can hold before it must grow), and - a `count`/length = how many chars are actually used. `append(x)` copies `x` into the unused tail of the buffer and bumps the length. No re-copy of existing characters — that's why one append is **O(length of x)**, and appending single pieces across the loop is **amortized O(1)** each, giving **O(n)** for the whole loop. ```java StringBuilder sb = new StringBuilder(); for (String part : parts) { sb.append(part); } String result = sb.toString(); // one final String built here ``` `toString()` produces the final immutable String once, at the end. ## Why and when the buffer 'grows' The default capacity of `new StringBuilder()` is **16 characters**. When you append past the current capacity, StringBuilder must **grow**: it allocates a new, larger array (the growth policy is roughly *double the old capacity, plus a bit*), copies all existing characters into it, and discards the old array. Each individual grow is an O(current-length) copy, but because the capacity roughly *doubles* each time, the grows happen geometrically less often, and the **total** copying across all grows is still O(n) — this is the meaning of **amortized O(1)** per append. (Amortized = averaged over the whole sequence of operations; a few operations are expensive but they're rare enough that the average stays constant.) ## Pre-sizing If you can estimate the final length, pass it to the constructor: ```java StringBuilder sb = new StringBuilder(expectedLength); ``` This allocates a buffer big enough up front, so the loop performs **zero** grow-and-copy cycles and creates **zero** intermediate throwaway arrays. The asymptotic complexity is O(n) either way, but pre-sizing removes a constant-factor overhead and reduces garbage — meaningful on hot paths or very large outputs. Over-estimating wastes a little memory; under-estimating just means a few grows. A reasonable estimate (e.g. `count * averagePartLength`) is usually enough. ## Alternatives for simpler cases - `String.join(", ", list)` — when you're joining a known collection with a separator. - `list.stream().collect(Collectors.joining(", "))` — same idea in a stream pipeline (it uses a StringBuilder under the hood, often a `StringJoiner`). - These are clearer than a manual loop when you just need delimiter-separated values. ## StringBuilder vs StringBuffer `StringBuffer` is the older, **synchronized** (thread-safe) version: every method is locked. `StringBuilder` is the unsynchronized version. In the overwhelmingly common single-threaded build, the locking in `StringBuffer` buys you nothing and costs you time, so the modern default is **StringBuilder**; reach for `StringBuffer` only if the same builder is genuinely shared and mutated across multiple threads (rare — usually you'd redesign instead).

  • What is StringBuilder's default capacity, and what happens when you exceed it?
    16 characters. When you append past capacity it allocates a larger array (roughly double), copies the existing characters over, and discards the old one — a grow-and-copy step.
  • If both pre-sized and default StringBuilder are O(n), why bother pre-sizing?
    To eliminate the constant-factor cost of repeated grows and the intermediate garbage arrays they create — it can noticeably speed up hot paths and reduce GC pressure for large outputs.

saying these in an interview costs you the question

  • Calling toString() inside the loop (defeats the purpose by materializing the String repeatedly).
  • Believing pre-sizing changes the Big-O — it only removes constant-factor grow/copy overhead and garbage.
  • Using StringBuffer 'to be safe' in single-threaded code — the synchronization is pure overhead there.
  • Reusing/sharing a StringBuilder across threads without synchronization (it is not thread-safe).

context