skip to content

What are compact strings introduced in Java 9, and how do they change String's internal layout?

level: middleimportance: must knowfreq 55%

answer

  1. JEP 254, Java 9
  2. char[] -> byte[] + coder flag
  3. coder: LATIN1 (0) or UTF16 (1)
  4. ASCII string ~half the memory
  5. on by default; -XX:-CompactStrings disables

basics

~20 s

Since Java 9, String stores text in a byte array instead of a char array, with a small flag saying whether the bytes are Latin-1 (one byte per char) or UTF-16 (two bytes). Pure-ASCII text now uses half the memory.

solid answer

~50 s

Compact strings (JEP 254, Java 9) changed String's backing from char[] to byte[] plus a one-byte 'coder' flag. The coder is LATIN1 (0) or UTF16 (1). If every character of the String fits in Latin-1 (code points 0-255, which covers ASCII and Western European text), the JVM stores one byte per character; otherwise it falls back to two bytes per character, exactly like the old UTF-16 char[]. Since most real-world strings are ASCII, this roughly halves their footprint and cut overall heap usage and GC pressure significantly in typical apps. The change is internal: the public String API is unchanged, charAt and length still behave as if UTF-16, and the coder flag is consulted on each access. It is on by default; you can disable it with -XX:-CompactStrings, which forces all strings back to UTF-16.

code

java · 16 lines
java
// Conceptual shape of String since Java 9
public final class String {
    static final byte LATIN1 = 0;
    static final byte UTF16  = 1;

    private final byte[] value;  // 1 or 2 bytes per char
    private final byte coder;    // which encoding value uses

    public char charAt(int i) {
        if (coder == LATIN1) {
            return (char) (value[i] & 0xFF);          // widen 1 byte
        }
        int j = i << 1;                                // 2 bytes per char
        return (char) ((value[j] & 0xFF) | (value[j + 1] << 8));
    }
}

go deeper

for a junior

Knows that since Java 9 String uses a byte array and ASCII strings take less memory than before.

for a middle

Names JEP 254, the byte[]+coder layout, LATIN1 vs UTF16, and that a single non-Latin-1 char forces the whole string to two bytes per char.

for a senior

Explains why a fixed-width coder flag (not UTF-8) preserves O(1) charAt, the heap/GC motivation, and the -XX:-CompactStrings toggle and its uses.

for a principal

Reasons about platform-wide impact: heap and GC reduction at scale, interaction with intrinsics/vectorized String methods, and when to measure or disable for atypical UTF-16-heavy workloads.

## The problem compact strings solved Before Java 9, every `String` stored its text as a `char[]`, and every `char` is **two bytes** (a UTF-16 code unit). But the overwhelming majority of strings in real programs are plain ASCII or Western European text whose characters all have small code points. For such text the high byte of every `char` was always zero - so roughly **half of all String storage was wasted zero bytes**. Heap-dump analyses repeatedly showed Strings (and their char arrays) as the single largest consumer of live heap. That is the motivation for **JEP 254: Compact Strings**, shipped in Java 9. ## Terms, defined - **Latin-1 (ISO-8859-1)**: a single-byte encoding covering code points 0 through 255. This includes all ASCII (0-127) plus common accented Western European letters (128-255). One character = one byte. - **UTF-16**: a two-byte-per-code-unit encoding (the old behavior); needed for anything outside Latin-1, e.g. Cyrillic, CJK, emoji. - **coder**: a one-byte flag stored in each String saying which of the two encodings the byte[] uses. - **footprint / heap**: how much memory the object occupies; smaller footprint means less garbage-collection work. ## The new layout String's storage became: ```java private final byte[] value; // raw bytes, not chars private final byte coder; // 0 = LATIN1, 1 = UTF16 ``` plus the usual cached `hash`. When a String is constructed, the JVM scans the input: if **every** character fits in Latin-1, it packs one byte per character and sets `coder = LATIN1`. If even one character needs more, it stores two bytes per character (UTF-16 layout) and sets `coder = UTF16`. There is no per-string mixing: a single non-Latin-1 character forces the whole String into UTF16 mode. ## Why a flag instead of always-byte UTF-8 (a variable-length encoding) was rejected because random access - `charAt(i)` - must stay O(1). With a fixed width (1 byte in Latin-1 mode, 2 in UTF16 mode) plus the coder flag, the JVM can index directly: the position is `i` or `2*i`. So the design keeps constant-time indexing while saving memory on the common case. ## Behavior is unchanged The public API behaves exactly as before. `length()` still returns the number of UTF-16 code units; `charAt(i)` still returns a `char`. In Latin-1 mode the byte is widened back to a char on read; in UTF16 mode two bytes are combined. All the work happens behind the field accesses, gated on `coder`. ## Controlling it Compact strings are **on by default**. The flag `-XX:-CompactStrings` disables the optimization and forces every String to UTF16 (the old char[]-equivalent layout), which is occasionally useful to measure the feature's impact or to work around pathological all-UTF16 workloads where the branch on `coder` adds overhead without saving memory. ## How to derive an answer at any level Start from "old String wasted a zero byte per ASCII char," then "Java 9 stores byte[] + a coder flag (LATIN1 or UTF16)," then "one byte per char when it fits, else two," then "halves ASCII memory, API unchanged, -XX:-CompactStrings disables it." That chain rebuilds the full answer.

  • If a 1000-character string contains 999 ASCII chars and one emoji, how is it stored?
    Entirely in UTF16 mode - two bytes per character for all 1000 chars (the emoji actually needs a surrogate pair, so it is two chars). A single non-Latin-1 character forces the whole string to UTF16; there is no per-character mixing.
  • Does compact strings change the result of length() or charAt()?
    No. The public API is identical and still speaks in UTF-16 code units. Only the internal storage and an extra branch on the coder flag changed; results are the same as before Java 9.

Like shipping a box: if everything is small (Latin-1), you use small slots and fit twice as much; if even one oversized item shows up, you switch the whole box to large slots (UTF-16). A label on the box (the coder) tells the reader which slot size to expect.

saying these in an interview costs you the question

  • Saying strings are stored as UTF-8 internally - they are Latin-1 or UTF-16, never UTF-8.
  • Claiming a string can mix Latin-1 and UTF-16 bytes; the coder is per-string, all-or-nothing.
  • Thinking the optimization changed the public API or charAt semantics.
  • Believing it is off by default; it is on, and -XX:-CompactStrings turns it off.
  • Assuming any non-ASCII text defeats it - Latin-1 covers code points up to 255, not just ASCII.

context