Why was String's internal representation able to change from char[] to byte[] without breaking applications, and what does that teach about API/implementation boundaries at platform scale?
answer
- Storage was private; contract is UTF-16 behavior, not array type
- char[] -> byte[]+coder, identical length/charAt/hashCode/equals
- Only boundary-violators broke: reflection on value, JNI, heap-dump parsers
- Validated by huge test corpus; gated by -XX:-CompactStrings
- Lesson: narrow stable contract enables fleet-scale internal change
basics
~20 sThe internal array was private, and the public methods (length, charAt, etc.) kept their exact behavior. So the JVM could swap char[] for byte[] without changing anything visible to programs - well-behaved code never depended on the hidden field.
solid answer
~50 sString's character storage was always a private field, and the public API was specified in UTF-16 code units, independent of how those units were stored. Compact strings (Java 9) replaced char[] with byte[]+coder while preserving every observable behavior: length(), charAt(), hashCode(), and equals() returned identical results, so programs against the contract were unaffected. The few things that broke were code reaching past the boundary - reflection on the value field, JNI assuming a char[], or serialization formats that read internals - which were never supported. The lesson is that disciplined encapsulation plus a stable, implementation-independent contract is what lets a platform make a fleet-wide performance change (halving String heap) transparently. It also shows the value of intrinsics and benchmarking: the JVM could optimize the hottest type in the language only because nothing public had ever promised the storage shape, and behavior parity was validated by an enormous test corpus before flipping the default.
go deeper
Understands that programs kept working because they only used String's public methods, not its hidden array.
Explains that the array was private and the public methods behaved identically, so well-behaved code was unaffected.
Identifies the contract-vs-implementation boundary, names which boundary-violating code broke (reflection/JNI), and why those were never supported.
Generalizes to platform strategy: narrow stable contracts enable fleet-scale internal evolution; can discuss validation corpus, intrinsics, kill-switch flags, and the engineering economics of changing the hottest type safely.
## The question behind the question String is arguably the most-used reference type in Java. Changing its internals is terrifying precisely because so much code touches it. The fact that Java 9 could swap the backing array from `char[]` to `byte[]+coder` and have virtually every program keep working is a case study in **API vs implementation boundaries**. ## Terms, defined - **Encapsulation**: hiding internal state (here, a `private` field) so callers cannot depend on it. - **Contract / specification**: the documented, observable behavior of the public methods - what they return, not how. String's contract is defined in **UTF-16 code units**, never in terms of a storage array. - **Observable behavior**: anything a conforming program can witness through the public API (return values, exceptions, ordering, hash codes). - **Intrinsic**: a method the JIT compiler recognizes and replaces with hand-tuned machine code (many String methods are intrinsified). - **Safe publication / immutability**: because Strings are immutable and their fields final, the storage choice has no concurrency-visible effect. ## Why the change was transparent Three properties made it safe: 1. **The storage was private.** No public method ever exposed the `char[]`. Callers could only observe via `length()`, `charAt()`, `getChars()`, `toCharArray()`, etc. 2. **The contract was storage-independent.** Those methods promise UTF-16 *code units*, not a particular array type. `byte[]`+coder can produce the exact same code units, so every return value matched. 3. **Behavioral parity was preserved and tested.** `equals`, `hashCode` (the documented s[0]*31^(n-1)+... formula), comparison, and iteration all returned identical results. The JVM team validated this against a vast test corpus before making compact strings the default. Thus any program written **to the contract** could not tell the difference. ## What did break - and why it was always unsupported - **Reflection** reading `String.value` as a `char[]` (e.g. some serialization or 'fast' hacks) broke, because the field is now `byte[]`. Reaching into private fields was never part of the contract. - **JNI / native code** assuming a `char[]` layout could break; the supported JNI String functions (`GetStringChars`, etc.) kept working because they too are defined on the abstract UTF-16 view. - **Code parsing heap dumps** with hard-coded String shapes needed updating. The pattern: only code that **violated the boundary** was affected. That is the strongest possible argument for encapsulation. ## The platform-scale lesson This is why a runtime can deliver a fleet-wide win (cutting one of the largest heap consumers roughly in half) without a migration: the boundary was respected for two decades, so the implementation had freedom to evolve. It also illustrates the JVM's broader strategy - keep public contracts narrow and stable, optimize aggressively underneath (compact strings, string intrinsics, indify string concatenation in Java 9's `invokedynamic`-based concat), and gate changes behind a flag (`-XX:-CompactStrings`) plus massive compatibility testing. The takeaway for any large system: invest in a minimal, behavior-defined public surface, and you buy yourself the right to change everything behind it later. ## How to derive an answer at any level From "the array was private and the API promised UTF-16 behavior, not a storage type," conclude "identical observable results -> conforming code unaffected; only boundary-violators (reflection/JNI/heap-dump parsers) broke," then generalize to "narrow stable contracts enable fleet-scale internal evolution."
- What concrete code patterns broke when char[] became byte[], and were they ever supported?Reflection reading the private 'value' field as char[] (used by some serialization or performance hacks), native/JNI code assuming a char[] memory layout, and heap-dump tools with hard-coded String shapes. None were part of the supported contract - they reached past the encapsulation boundary, which is exactly why they were fragile.
- How did the JVM team de-risk flipping the default to compact strings?By preserving exact observable behavior (same length/charAt/hashCode/equals results), validating against a very large compatibility test corpus, intrinsifying hot paths to avoid regressions, and shipping a kill switch (-XX:-CompactStrings) so anyone hitting a problem could revert without recompiling.
Like renovating a building's plumbing without changing the taps: residents (programs) only ever touched the taps (public API), so swapping copper for PEX behind the walls was invisible - except to anyone who had illegally drilled into the wall to tap a pipe directly (reflection/JNI).
saying these in an interview costs you the question
- Claiming the change required application code changes for conforming programs.
- Saying nothing broke at all - reflection/JNI/heap-dump-parsing code that breached the boundary did.
- Attributing the safety to luck rather than to a private field plus a storage-independent contract.
- Confusing the storage change with an API change to length()/charAt() semantics (those were preserved).
- Believing immutability is what allowed the swap - immutability helps concurrency, but encapsulation + a stable contract is the real enabler.