What are the memory implications of String's internal representation, including object overhead and the savings from compact strings?
answer
- Two objects: String wrapper + backing array
- Each has a header (~12-16B) + 8-byte alignment padding
- Compact strings ~halve ASCII payload (2B/char -> 1B)
- Short strings dominated by fixed overhead
- Levers: interning, -XX:+UseStringDeduplication (G1), measure with JOL
basics
~20 sA String costs more than just its text: the String object header, fields, and the separate backing array each have memory overhead. Compact strings cut the text bytes roughly in half for ASCII, but very short strings are dominated by fixed per-object overhead.
solid answer
~50 sA String's memory is two objects: the String wrapper (object header, the array reference, the coder and hash ints) and the backing array (its own header plus the payload bytes). On a 64-bit JVM with compressed oops, each object header is about 12-16 bytes and allocations align to 8 bytes. Before Java 9 the payload was 2 bytes per char; since Java 9 it is 1 byte per char for Latin-1 text and 2 for UTF-16, so ASCII strings roughly halved their payload, which materially reduced heap and GC pressure in typical apps. The catch is fixed overhead: for short strings the two headers plus padding dominate, so the relative saving shrinks. This is why interning, deduplication (-XX:+UseStringDeduplication with G1), and avoiding gratuitous small strings matter. For exact sizing you measure with a tool like JOL rather than guess.
code
java · 13 lines// Estimating real size with JOL (Java Object Layout)
// dependency: org.openjdk.jol:jol-core
import org.openjdk.jol.info.GraphLayout;
public class StringSize {
public static void main(String[] args) {
String s = "hello"; // 5 ASCII chars
// totalSize includes the String wrapper AND its backing byte[]
System.out.println(GraphLayout.parseInstance(s).totalSize());
// Typically ~48 bytes on 64-bit + compressed oops:
// ~24 (wrapper) + ~24 (byte[] header + 5 padded payload)
}
}go deeper
Understands that a String uses more memory than just its letters and that ASCII strings got smaller in modern Java.
Can explain the two-object layout (wrapper + array), per-char payload before/after Java 9, and that compact strings roughly halve ASCII payload.
Quantifies fixed overhead (headers + alignment + compressed oops), explains why short strings are overhead-bound, and knows interning, deduplication, and measuring with JOL.
Reasons about heap/GC at fleet scale: when deduplication vs interning pays off, the memory/throughput trade-offs, compressed-oops thresholds, and how to drive decisions from heap-dump evidence rather than rules of thumb.
## Why a String costs more than its letters A `String` is **two heap objects**, not one: 1. the `String` instance itself (a wrapper), and 2. the backing array it points to (`char[]` before Java 9, `byte[]` since). Each heap object carries an **object header** the JVM uses for bookkeeping (identity hash, lock state, garbage-collection marks, and for arrays a length field), and every allocation is **padded** up to an 8-byte boundary (object alignment). So even an empty String is not free. ## Terms, defined - **Object header**: per-object metadata. On HotSpot 64-bit with **compressed ordinary object pointers (compressed oops)**, a normal object header is ~12 bytes (mark word + compressed class pointer); an array adds a 4-byte length, so ~16 bytes. Without compressed oops the references and headers are larger. - **Compressed oops**: a JVM trick (default below ~32 GB heaps) that stores object references as 32-bit values instead of 64-bit, shrinking footprint. - **Alignment / padding**: object sizes are rounded up to a multiple of 8 bytes, so a few bytes are often wasted at the end. - **Payload**: the actual text bytes in the backing array. - **Retained size**: total memory kept alive by an object, including objects only it references (here: the wrapper + its array). ## A worked size estimate (64-bit, compressed oops, post-Java 9) String wrapper fields: a reference to `value` (4 bytes with compressed oops), `int hash` (4), `byte coder` (1), plus `~hashIsZero` flag. With a ~12-byte header and 8-byte alignment, the wrapper is about **24 bytes**. The backing `byte[]` has a ~16-byte header plus payload, rounded to 8. So for the 5-char ASCII string "hello": wrapper 24 + array (16 header + 5 payload -> 24 padded) = roughly **48 bytes** for 5 letters. The fixed overhead (~40 bytes) dwarfs the 5 bytes of text. ## The compact-strings saving, quantified For a long ASCII string the payload dominates and dropping from 2 bytes/char to 1 byte/char roughly **halves the array payload**. For a 100-char ASCII string: pre-Java 9 ~200 payload bytes; post-Java 9 ~100. That is the heap-wide win JEP 254 delivered. But for short strings the two object headers plus padding are fixed, so the proportional benefit shrinks toward zero - a 1-char string saves only 1 byte of payload out of ~40+ total. ## Consequences and mitigations - **Avoid swarms of tiny strings**; their fixed overhead is the killer, not their text. - **Interning** (`String.intern()` or literal pooling) lets identical strings share one object, but the pool itself uses memory and is GC-rooted - use deliberately. - **String deduplication**: `-XX:+UseStringDeduplication` (with G1) makes the GC detect equal backing arrays across distinct String objects and point them at one shared array, saving payload without code changes. It deduplicates the array, not the wrapper. - **Measure, do not guess**: use JOL (Java Object Layout) or a heap profiler for exact per-object sizes, because numbers depend on JVM, compressed-oops state, and alignment. ## How to derive an answer at any level From "a String is wrapper + array, each with a header and padding," reason that (a) short strings are overhead-bound, (b) long ASCII strings are payload-bound and compact strings halve that payload, and (c) the practical levers are fewer small strings, interning, and -XX:+UseStringDeduplication, validated with JOL.
- Why does compact strings help much less for a Map with thousands of 2-3 character keys?Because such strings are overhead-bound: the two object headers plus 8-byte padding (~40 bytes) dominate, and the payload (2-3 bytes) is tiny. Halving 3 payload bytes to ~2 is negligible against the fixed cost, so the percentage saving is small.
- How does -XX:+UseStringDeduplication differ from interning?Interning shares the whole String object for equal content and is application-driven (intern()/literals). Deduplication is a GC feature (G1) that, after objects survive a while, makes distinct String objects share one identical backing array - it dedups the array payload, not the wrapper, and requires no code changes.
Each String is a parcel inside a parcel: the inner box (array) holds the letters, the outer box (wrapper) holds a label pointing to it. Two boxes means two sets of packaging tape and padding - cheap for a big shipment of text, expensive when you mail one letter at a time.
saying these in an interview costs you the question
- Estimating String size as just 2*length (or length) bytes, ignoring two object headers and alignment padding.
- Claiming compact strings help short strings as much as long ones; short strings are overhead-bound.
- Saying interning always saves memory; the pool itself costs memory and can pin objects.
- Confusing String deduplication (shares the array) with interning (shares the whole object).
- Quoting exact byte counts as universal; they depend on compressed oops, JVM, and alignment - measure with JOL.