skip to content

String Internal Representation

Since Java 9, String stores a byte array with a coder flag, using Latin-1 where possible instead of always UTF-16. Compact strings cut the memory footprint of typical applications noticeably, which is why the change is worth knowing.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

How does a Java String store its characters internally, and what was the original backing data structure?

level: juniorimportance: must knowfreq 60%

answer

  1. private final char[] value (Java 8 and earlier)
  2. char = 2 bytes = one UTF-16 code unit
  3. final class + final array = immutable
  4. immutable -> safe to share, cache, use as key
  5. substring/replace return new String, never mutate

basics

~20 s

A String keeps its characters in a private array inside the String object. Originally that was a char array with two bytes per character. The String is immutable, so the array never changes after the String is created.

solid answer

~40 s

A String wraps a private final array holding its characters plus a cached hash code. From Java 1.0 through Java 8 the backing was a final char[], where each char is a 16-bit UTF-16 code unit, so every character cost two bytes regardless of its value. Because String is immutable, this array is set once in the constructor and never mutated; operations like substring or replace return brand-new String objects. The class is final so no subclass can break that guarantee, which is why Strings are safe to share, cache, and use as map keys, and why interning and the string pool work. In Java 9 the backing changed to a byte[] plus a coder flag (compact strings), but the immutable, encapsulated-array design stayed identical.

code

java · 12 lines
java
// Conceptual shape of String pre-Java 9
public final class String {        // final: no subclass
    private final char[] value;    // 2 bytes per char, set once
    private int hash;              // cached hashCode (0 until computed)

    String concat(String other) {  // does NOT mutate; returns new String
        char[] buf = new char[value.length + other.value.length];
        System.arraycopy(value, 0, buf, 0, value.length);
        System.arraycopy(other.value, 0, buf, value.length, other.value.length);
        return new String(buf);    // brand-new object
    }
}

go deeper

for a junior

Knows a String holds characters in an internal array and that Strings are immutable, so 'changing' one makes a new object.

for a middle

Can name the exact field (final char[] value pre-Java 9), explain that each char is two bytes / UTF-16, and tie immutability to thread-safety and map-key safety.

for a senior

Articulates why final class + private final field guarantee immutability, the role of the cached hash, and the memory cost of 2-byte chars that motivated later optimization.

for a principal

Frames immutability as an API/security/concurrency contract (defensive copying at the boundary, safe publication, interning trade-offs) and can reason about the consequences for the platform, not just one object.

## What a String actually is A `String` in Java is an ordinary **object**, not a primitive. Internally it is a thin wrapper around an array that holds the text, plus a couple of bookkeeping fields. Two declarations are central: - the class is `final` (no subclass is allowed), and - its storage field is `private final` (cannot be reassigned, cannot be read from outside). Together these create **immutability**: once a `String` exists, the characters it represents can never change. ## Terms, defined - **char**: a Java primitive that is a 16-bit (2-byte) unsigned value. It holds one **UTF-16 code unit**. - **UTF-16**: a way of encoding Unicode text where most characters fit in one 16-bit unit; rare characters (emoji, some scripts) use two units called a *surrogate pair*. - **byte**: an 8-bit (1-byte) value. - **immutable**: the object's observable state never changes after construction. - **code unit** vs **code point**: a code unit is one storage slot (16 bits in UTF-16); a code point is one logical Unicode character, which may need one or two code units. ## The original backing: char[] From Java 1.0 through Java 8 the field was: ```java private final char[] value; ``` plus an `int hash` field caching the hash code. Because each `char` is two bytes, a String of N characters used at least 2*N bytes for text, regardless of whether the text was plain ASCII ("hello") or contained non-Latin characters. ASCII text therefore wasted half its storage: "h" only needs the value 104, which fits in one byte, but was stored in two. ## Why immutable The constructor copies the incoming characters into `value` and that array is never handed out or modified afterward. Any method that seems to "change" a String actually returns a new String. This makes Strings: - **thread-safe** with no locking (no one can mutate shared text), - safe as **HashMap keys** (the hash code can be cached and never goes stale), - **poolable / internable** (identical literals can share one object). ## What changed later (preview) Java 9 replaced `char[]` with `byte[]` plus a `byte coder` flag — the "compact strings" feature — to stop wasting space on Latin-1 text. The wrapper-around-an-immutable-array model did not change; only the element width and an extra flag did. That follow-on optimization is the subject of the other questions in this leaf. ## How to derive an answer at any level Start from "String is an object wrapping a private final char array (originally)," add "each char is 2 bytes / UTF-16," then "final class + final field = immutable," then list the consequences (sharing, keys, pooling). That chain reconstructs every level of answer.

  • If String is immutable, why does it have a non-final 'hash' field?
    hash is a lazily computed cache of hashCode(). It starts at 0 and is filled on first call. Writing it does not change the observable text, so it is a benign internal cache and does not violate immutability (it is even safe under data races because all threads compute the same value).
  • Does immutability mean a String can never be garbage collected?
    No. Immutability is about not changing contents, not about lifetime. An unreferenced String is collected normally. (Interned strings live as long as the pool references them, which is a separate concern.)

Think of a String as a sealed envelope: you fill it once (the array), seal it (final), and from then on you can only read it or make a fresh copy with edits, never reopen and rewrite the original.

saying these in an interview costs you the question

  • Saying String stores chars as one byte each before Java 9 (it was two bytes per char, UTF-16).
  • Claiming you can modify a String in place via reflection 'safely' - it breaks the JVM's immutability assumptions (pool, cached hash) and is undefined behavior.
  • Confusing String with StringBuilder/StringBuffer, which are mutable by design.
  • Saying a char always equals one user-visible character (surrogate pairs need two chars).

context

open as a page

What are compact strings introduced in Java 9, and how do they change String's internal layout?

level: middleimportance: must knowfreq 55%

basics

~20 s

Since Java 9, String stores text in a byte array instead of a char array, with a small flag saying whether the bytes are Latin-1 (one byte per char) or UTF-16 (two bytes). Pure-ASCII text now uses half the memory.

open as a page

Given the UTF-16-based String API, what subtle bugs arise from char-based operations on text with supplementary characters, and how do you handle them correctly?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A Java char is one UTF-16 unit, but some characters (like many emoji) need two chars (a surrogate pair). So length() and charAt() count units, not real characters, and slicing or counting by char can split a character in half. Use code-point methods or codePoints() for correctness.

open as a page

What are the memory implications of String's internal representation, including object overhead and the savings from compact strings?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A String costs more than just its text: the String object header, fields, and the separate backing array each have memory overhead. Compact strings cut the text bytes roughly in half for ASCII, but very short strings are dominated by fixed per-object overhead.

open as a page

Why was String's internal representation able to change from char[] to byte[] without breaking applications, and what does that teach about API/implementation boundaries at platform scale?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

The internal array was private, and the public methods (length, charAt, etc.) kept their exact behavior. So the JVM could swap char[] for byte[] without changing anything visible to programs - well-behaved code never depended on the hidden field.

open as a page