skip to content

What does sys.intern() do to a Python string, and when is calling it worth it?

level: middleimportance: should knowfreq 40%

answer

  1. One copy instead of many
  2. Runtime-built strings are never shared
  3. A table lookup returns the canonical object
  4. Repetition pays, uniqueness does not
  5. Exact str only, no un-intern

basics

~20 s

sys.intern() stores one canonical copy of a string in the interpreter's intern table and returns it, so equal strings built at runtime collapse to a single object. It pays off when a few distinct values repeat many times.

solid answer

~40 s

CPython shares some strings for free: equal literals in one compilation unit, and identifier-like literals, which are interned at compile time. Strings created at *runtime* - from `str.split()`, from decoding bytes, from a parser - are always fresh objects, so a million rows can hold a million equal copies of the same key. `sys.intern(s)` looks the value up in the interpreter's intern table and returns the object already there, letting your duplicate be freed; the result is one object per distinct value instead of one per occurrence. Equality also gets a pointer fast path, though hashes are cached on every `str` anyway. Intern low-cardinality repeated values only - on near-unique data you pay a hash and a probe per call and save nothing. It accepts exact `str` only; `bytes` and `str` subclasses raise `TypeError`.

code

python · 12 lines
python
import sys

x = "".join(["customer", "_id"])
y = "".join(["customer", "_id"])
print(x == y, x is y)

print(sys.intern(x) is sys.intern(y))

try:
    sys.intern(b"customer_id")
except TypeError as exc:
    print(exc)

go deeper

for a junior

Be ready to say what the call does in one sentence: it returns one shared copy of a string so equal strings stop being separate objects. Knowing that strings made at runtime are not shared automatically is enough at this level.

for a middle

An interviewer expects the mechanics: compile-time sharing versus runtime allocation, what the intern table lookup returns, why the duplicate is then freed, and the fact that only exact str is accepted. Be able to name a case where interning is the wrong call.

for a senior

Show that you decide on evidence: profile first, intern per column or per field on a measured repeat rate, and place the call where the string is created so peak memory drops rather than only steady state. Say out loud what it does not fix.

for a principal

Own the framing that interning is a cheap local optimisation, not an architecture. If a job only fits in memory because of it, the real decision is streaming, aggregating during the read, or a compact record layout - and you should be able to say when the one-line fix buys enough runway to defer that.

### Two different mechanisms wear the same word "Interning" in CPython covers two things that are easy to conflate. The first is **automatic sharing**, which you never ask for: the compiler stores each distinct constant of a compilation unit once, string literals that look like identifiers (ASCII letters, digits and underscores) are placed in the interpreter's intern table at compile time, and the integers from -5 through 256 are preallocated once at start-up. The second is **explicit interning** with `sys.intern()`, which is the only tool you have for strings that did not exist when the code was compiled. That second category is where the memory goes. A string produced at runtime is always a brand new object: the result of `str.split()`, of decoding bytes, of concatenating two variables, of a parser handing you a field name. CPython does not check whether an equal string already exists — that check would cost a hash and a lookup on every string creation, which would be a bad trade. So a loop that reads a million records and builds a dict per record allocates a million separate copies of every key, each with its own object header, its own cached-hash slot and its own character buffer, all holding identical bytes. ### What `sys.intern()` actually does `sys.intern(s)` hashes the value, looks it up in the interpreter-wide intern table, and returns the object already stored there if there is one; otherwise it stores `s` and returns `s`. Rebinding your name to the returned object means the duplicate you just built has no references left and is freed immediately by reference counting. The net effect on a repetitive data set is one object per *distinct* value rather than one per *occurrence*. There is a second, smaller win. String equality in CPython starts with a pointer check, so comparing two interned strings that are equal answers on the first instruction instead of walking the characters. Dictionary lookups benefit on collisions and on the final key comparison. Do not oversell it: every `str` caches its hash after the first use whether it is interned or not, so interning does not make hashing cheaper, and the comparison win only shows up in hot loops that compare or look up the same long strings many times. ### The rules and the limits `sys.intern()` accepts an exact `str` only. Passing `bytes` raises `TypeError: intern() argument must be str, not bytes`, and a `str` subclass raises `TypeError: can't intern` — worth knowing when a parsing layer hands you subclassed strings. There is no public way to un-intern a value, and there is no "intern everything" switch. The decision rule is cardinality, not size. Interning pays when a small number of distinct values appears many times: column names, category codes, status flags, symbol names, repeated dict keys. It actively costs you when values are near-unique — request ids, hashes, free text — because you pay a hash and a table probe per call and gain nothing, and the table now carries an entry per value. Measure the repeat rate before you reach for it. On lifetime, CPython 3.14 is friendlier than the folklore. Small integers and a set of statically allocated short strings became *immortal* in 3.12 under PEP 683 — their reference counts never change. Strings you intern yourself at runtime are not in that group: on 3.14 the intern table does not add a counted reference, so such a string is still freed once your last reference goes away. The old claim that "anything you intern lives until the interpreter exits" should not be repeated as a blanket fact on modern CPython. ### The correctness line Interning is a memory optimisation and nothing more. Which values happen to be shared is a CPython implementation detail that varies by version, by build and by other implementations, so no program should behave differently because two equal values turned out to be one object. If you need a guaranteed single instance — a sentinel, a flyweight, a cache — build it yourself with a module-level object or your own dict, where the sharing is a property of your code rather than of the runtime. If profiling says the repeated *keys* rather than the repeated *values* dominate, interning is the cheap 20% fix and a change of record layout is the structural one; that layout question belongs to object overhead, not here. Start with interning because it is a one-line change at the point where the strings are created, measure, and only then restructure.

  • Does interning speed up dictionary lookups, or only cut memory?
    Mostly memory. Every `str` caches its hash after first use whether interned or not, so hashing is unchanged. What interning adds is a pointer-equality fast path in string comparison, which shortens the final key comparison and collision handling. That is measurable in a hot loop comparing long strings and negligible elsewhere - claim the memory win, not a general speedup.
  • Can you un-intern a string, and does the intern table grow forever?
    There is no public un-intern call. On CPython 3.14 the intern table does not hold a counted reference to a runtime-interned string, so the string can still be freed when your last reference goes away - it is not a leak. Interning near-unique values is still wasted work: a hash, a probe and a table entry per value, for no sharing.
  • Why not just dedupe with your own dict instead of sys.intern()?
    That is a legitimate option and sometimes better, because you control the lifetime and can drop the whole map when a batch finishes. `sys.intern()` is cheaper to write, is interpreter-wide so unrelated parts of the program share the same objects, and is implemented in C. Use your own map when you want the deduplication scoped to one job, and intern when the values are process-wide.

It is a coat check for strings: hand in your copy, get back the one ticket everyone else with the same coat already holds, and the duplicate is thrown away.

saying these in an interview costs you the question

  • Thinks all equal strings in CPython are already one object
  • Interns every string, including unique ids
  • Claims sys.intern() accepts bytes
  • Says interning makes string hashing cheaper
  • Relies on interning for identity checks in program logic
  • Interns after building the list, expecting a lower peak

context