How can sys.intern cut memory in a nightly report generator holding millions of parsed rows, and where does it stop helping?
answer
- Profile before optimising
- Parsed fields are always new objects
- Keys repeat once per row per column
- Intern at creation, not afterwards
- Decide per column on repeat rate
basics
~20 sParsed fields are freshly allocated, so repeated keys and category values exist once per row. Interning them as the parser produces them leaves one object per distinct value. It does nothing for unique values, numbers or the row containers.
solid answer
~40 sConfirm first with a `tracemalloc` snapshot at peak that string construction, not the row containers, dominates. Every string a parser returns is a new object, so ten columns across two million rows can hold twenty million copies of ten distinct key strings. Calling `sys.intern()` on keys and low-cardinality values **inside the parse loop** returns the canonical object and lets each duplicate be freed immediately, cutting live string objects from occurrences to distinct values; doing it after the list is built still frees them but you have already paid the peak. Decide per column on a measured repeat rate - a column hitting an existing value 83% of the time pays, a column of unique ids does not. It cannot help with numbers, timestamps, one large unique string, or the per-row containers themselves.
code
python · 15 linesimport sys, tracemalloc
def build(n, shared):
rows = []
for i in range(n):
key = "".join(["customer", "_id"])
rows.append({sys.intern(key) if shared else key: i})
return rows
for shared in (False, True):
tracemalloc.start()
data = build(200_000, shared)
print("interned" if shared else "plain ", tracemalloc.get_traced_memory()[0] // 1024, "KiB")
tracemalloc.stop()
del datago deeper
Know the shape of the problem: strings that a parser builds are separate objects even when equal, so repeated keys cost memory once per row. Being able to name sys.intern() as the tool that collapses them is enough here.
Explain the mechanics you would apply: intern in the parse loop rather than afterwards, only for columns whose values repeat, and know that it accepts exact str only. Be able to sketch a before and after measurement.
Demonstrate the discipline: profile at peak first, decide per column on a measured repeat rate, place the call where the string is created, and report memory and wall clock from the same input. Say plainly what interning cannot fix so nobody expects it to save a job that needs streaming.
Own the decision about when a memory optimisation is the wrong answer. If the nightly job only fits because of interning, the sustainable options are streaming or aggregating during the read, and your call is whether the one-line change buys enough runway to schedule the bigger one deliberately.
### First, prove strings are the problem A nightly report generator that parses a large extract and holds the rows in memory usually grows in one of three places: the per-row container, the field values, and the field *keys*. Before changing anything, take a `tracemalloc` snapshot at peak and look at which allocation sites dominate, and sample a few rows with `sys.getsizeof()` on the keys and values. If the top line is the parser's string construction, interning is on the table. If it is the row containers themselves, interning will not move the number and the fix is a different record layout — a question of object overhead rather than of interning. ### Why parsed data duplicates so aggressively Every string a parser hands back is freshly allocated. Split a line into fields and you get new objects for the field text; build a dict per row and the key strings are new objects too, once per row per column. Ten columns across two million rows is twenty million string objects holding ten distinct values. Equality is fine, memory is not: each carries an object header, a length, a cached hash slot and its own character buffer. `sys.intern()` collapses that. Interning a value returns the canonical object already in the interpreter's intern table, so the duplicate you just allocated loses its last reference and is freed at once. Applied to keys and low-cardinality values as they leave the parser, the count of live string objects drops from *occurrences* to *distinct values*. ### Measure the repeat rate before you commit Interning is only a win where values repeat. Run the loader over a sample, count occurrences per column with a `collections.Counter`, and compute what share of lookups would have hit an existing entry. A column showing an **83% cache-hit rate** — five sixths of its values already seen — is exactly the profile that pays. A column of transaction ids with a near-zero hit rate is the opposite: you would pay a hash and a table probe per row, add a table entry per value, and save nothing. Intern per column, on evidence, not across the whole record. Placement matters as much as the decision. Call `sys.intern()` at the point the string is created, inside the parse loop, not on a finished list afterwards. Interning late still frees the duplicates, but you have already paid the peak, and peak RSS is what makes the nightly job get killed. ### The traps The first is hand-rolled interning. A helper that dedupes with its own dict keyed by, say, the first 32 characters of each value looks like a cheap version of the same idea and is a correctness bug: two distinct long values sharing a prefix map to one entry, and the report silently comes out with the wrong figures — a **silent truncation** with no exception anywhere. If you keep your own dedupe map, key it on the full value. `sys.intern()` never has this problem. The second is type. `sys.intern()` accepts an exact `str`; `bytes` and `str` subclasses raise `TypeError`. A loader that reads raw bytes and defers decoding, or a parsing layer that returns a subclassed string, needs a decode or a `str()` conversion first. The third is scope. Interning does nothing for the numbers, the timestamps or the row containers, and it does not shrink one large unique string. If the profile shows those dominating, the honest answers are streaming instead of materialising the whole extract, aggregating during the read, or holding bulk numeric columns in a compact array rather than in per-row objects. The fourth is lifetime discipline. There is no un-intern call, and interning genuinely unique values grows the interpreter's table for the life of the process. On CPython 3.14 a runtime-interned string is still freed once your last reference drops, so this is a wasted-work problem rather than a leak — but it is wasted work in the hottest loop of the job. ### Reporting the result Land the change with a before/after number from the same input: peak traced memory or peak RSS, plus the wall clock, because a per-field intern call is not free. A 15-25% cut on a string-heavy extract is a typical outcome, it costs one line at the parse site, and it is reversible. If that is not enough to fit the budget, the next step is structural — stop storing a key per row at all — and that is a bigger change with a bigger payoff.
- How do you decide which columns to intern before shipping the change?Run the loader over a sample and count occurrences per column with a `collections.Counter`. Compare distinct values against total rows: a column where 83% of reads would have hit a value already seen is worth interning, and one approaching a distinct value per row is not. Do it per column rather than blanket-interning every field, and keep the measurement in a script so it can be rerun when the extract changes shape.
- A teammate replaced the interning with a dict keyed by each value's first 32 characters. What is wrong with that?It is not deduplication, it is a collision. Two distinct values sharing a 32-character prefix map to the same entry, so one silently replaces the other and the report comes out wrong with no exception raised anywhere. If you keep your own dedupe map, key it on the full value; `sys.intern()` compares full values and cannot do this.
- The parse layer hands back subclassed strings. What happens when you intern them?`sys.intern()` accepts an exact `str` only, so a `str` subclass raises `TypeError: can't intern`, and raw `bytes` raises a `TypeError` naming the type. Convert at the boundary - decode the bytes, or build a plain `str` from the subclass - and intern the plain value. That conversion allocates, which is fine because the duplicate is freed as soon as interning returns the canonical object.
It is warehouse consolidation: twenty million identical labels printed on twenty million separate cards, replaced by one card everybody points at - which does nothing for the boxes or for the labels that are all different.
saying these in an interview costs you the question
- Interns every field without profiling first
- Interns after materialising the whole list and expects a lower peak
- Expects interning to shrink numbers or row containers
- Hand-rolls dedupe on a truncated key prefix
- Reports the change with no before and after measurement
- Assumes interning leaks because the table never releases anything