When is calling intern() a good idea at scale, and what are the risks and alternatives?
answer
- Win: high duplication + low cardinality
- Lose: high/unbounded cardinality, per-call lookup cost
- StringTable is a fixed native hash table (-XX:StringTableSize)
- Alternatives: app-level Interner, G1 UseStringDeduplication, enums
- Never intern untrusted/unbounded input
basics
~20 sintern() helps when you have many duplicate runtime strings drawn from a small set, because it collapses them to one shared object and saves memory. It hurts when values are mostly unique, because the pool grows, lookups cost time, and GC suffers. Often a plain HashMap cache or JVM string deduplication is safer.
solid answer
~40 sintern() pays off for low-cardinality, high-duplication runtime strings: parsing large datasets where the same few tokens (status codes, column names, enum-like values) repeat millions of times, so collapsing to a single pooled instance saves heap and enables fast == comparison. It backfires for high-cardinality data: each unique value adds a pool entry, every call costs a hash lookup, and a large native-managed pool can pressure GC and CPU. Alternatives are usually preferable: a bounded ConcurrentHashMap-based canonicalizing cache gives you control over eviction and capacity; the JVM's G1 String Deduplication (-XX:+UseStringDeduplication) deduplicates char arrays in the background without code changes; and enums or interned constants for truly fixed vocabularies. As an architect, I would benchmark before adopting intern(), avoid interning untrusted/unbounded input, and document any code that depends on string identity since it is fragile.
go deeper
Understands intern() can save memory for repeated strings but should not be applied blindly.
Identifies the duplication-vs-cardinality trade-off and knows equals() is the safe default.
Compares intern() with an application-level canonicalizing cache and explains StringTable/GC costs.
Owns a measured decision framework, picks the least-global tool, forbids interning untrusted input, considers G1 string dedup, and treats identity dependence as a documented invariant.
## The trade-off in one sentence `intern()` trades a **per-call lookup cost** and **pool growth** for **memory savings** (shared instances) and **fast `==` comparison** — so it wins only when duplication is high and cardinality is low. ## Background recap The **string pool** holds one canonical `String` per distinct value; **literals** are auto-interned. `s.intern()` finds the pooled value-equal string (or adds `s`) and returns the canonical reference. Since Java 7 the pool lives in the heap, but it is backed by a fixed-size **native hash table** (`StringTable`) whose default bucket count can be tuned with `-XX:StringTableSize`. ## When intern() is a good idea 1. **High duplication, low cardinality.** Parsing a 10-million-row CSV where a `status` column has 5 distinct values: without interning you may hold millions of separate `String` objects with identical content; interning collapses them to 5 shared instances, saving large amounts of heap. 2. **Fast equality in hot paths.** Once a fixed vocabulary is interned, `==` can replace `.equals()` for cheaper comparison (e.g. in a parser dispatch). 3. **Stable, bounded vocabularies** known at design time. ## The risks 1. **Lookup cost.** Every `intern()` hashes the string and probes the `StringTable`. In a tight loop over mostly-unique strings, this overhead can exceed any benefit. 2. **Pool/StringTable pressure.** A high default-size native table degrades to long bucket chains as you add many entries, slowing every intern and even literal resolution across the JVM (a global side effect). 3. **GC and footprint.** Although interned strings can be collected when unreferenced, an actively-referenced large pool is long-lived data that adds to old-gen pressure. 4. **Untrusted input = unbounded growth.** Interning user-supplied or unbounded-cardinality strings can be a memory-exhaustion vector. 5. **Fragile `==` semantics.** Code relying on identity after interning breaks the moment one path forgets to intern; this is a maintenance hazard. ## Alternatives (usually better) - **Application-level canonicalizing cache:** a `ConcurrentHashMap<String,String>` (or Guava `Interner`) you control — bounded size, eviction policy, isolated from the global pool. `map.computeIfAbsent(s, k -> k)` gives intern-like dedup without touching the JVM `StringTable`. - **G1 String Deduplication:** `-XX:+UseStringDeduplication` (with G1) deduplicates the backing `char[]`/`byte[]` of equal strings in the background — saves memory with **no code changes** and no `==` guarantee dependency. Good default for memory-heavy services. - **Enums / constants:** for truly fixed vocabularies, an `enum` or a set of `static final` constants is clearer and safer than interning runtime strings. - **Just use `.equals()`:** if comparison speed is not actually a bottleneck, skip interning entirely. ## A principal-level decision framework 1. **Measure first** — profile heap and confirm duplicate strings dominate. 2. **Estimate cardinality** — low and bounded favors interning/canonicalizing cache; high or unbounded rules it out. 3. **Prefer the least-global tool** — app-level interner or `-XX:+UseStringDeduplication` over `String.intern()`. 4. **Never intern untrusted/unbounded input.** 5. **If using intern(), tune `-XX:StringTableSize`** and document the identity dependency. ## Summary `intern()` is a sharp, niche tool. Reach for it only with measured, low-cardinality duplication; otherwise prefer a bounded application cache or JVM-level string deduplication, and treat any reliance on string identity as a documented, tested invariant.
- What does -XX:+UseStringDeduplication do and how does it differ from intern()?With G1 GC it scans surviving strings and shares their backing char/byte arrays when values match, saving memory automatically with no code changes. Unlike intern(), it does not make distinct String objects ==; it only deduplicates the underlying array.
- Why can interning high-cardinality strings hurt JVM-wide performance?The native StringTable has a fixed bucket count; adding many entries lengthens bucket chains, slowing every intern and even literal resolution globally, and the long-lived entries add old-gen GC pressure.
saying these in an interview costs you the question
- Treating intern() as a free, always-beneficial optimization
- Interning unbounded user input
- Forgetting the per-call lookup cost and global StringTable impact
- Depending on == identity without documenting/testing it
- Not considering -XX:+UseStringDeduplication as a zero-code alternative