When does swapping per-record lists for tuples actually cut memory in a search-index rebuilder holding millions of records?
answer
- Same five fields, two containers
- One allocation versus two
- Spare capacity you never asked for
- Count the bytes, then multiply by records
- The referenced objects usually dominate
basics
~20 sIt pays when records are numerous, small and fixed-shape. On CPython 3.14 a five-item list costs 104 bytes against 88 for the tuple, 120 if grown by appends — real at ten million records, noise at ten thousand.
solid answer
~50 sMeasure before switching. On a 64-bit CPython 3.14 build `sys.getsizeof` reports 88 bytes for a five-item tuple, 104 for the same list written as a literal, and 120 when that list was grown by appends and kept its spare capacity. So the saving is roughly 16–32 bytes per record: about 300 MB across ten million records, and nothing at all across ten thousand. The catch is that `sys.getsizeof` is shallow — it counts the container, not the strings it points at — so if each record holds five distinct strings the containers may be a minority of the real footprint, and the switch disappoints. Confirm the shape of the heap with `tracemalloc` snapshots before and after rather than reasoning from the per-object numbers. The second, often larger win is that tuple records become hashable, so deduplication with a set becomes possible at all.
code
python · 11 linesimport sys
rec_t = ("index-07", 4, True, 1024, "ok")
rec_l = ["index-07", 4, True, 1024, "ok"]
grown = []
for field in rec_t:
grown.append(field)
print(sys.getsizeof(rec_t), sys.getsizeof(rec_l), sys.getsizeof(grown))
# 88 104 120 on a 64-bit CPython 3.14 build
print(sys.getsizeof(rec_t) + sum(sys.getsizeof(f) for f in rec_t)) # shallow vs totalgo deeper
Know the direction of the difference: for the same items a tuple is a smaller object than a list, because a list keeps spare room for growth. You are not expected to quote byte counts or to plan a migration.
Explain where the difference comes from — one exactly-sized allocation versus a header plus an over-allocated pointer array — and use sys.getsizeof to measure it. Be clear that it reports the container only, never the objects it references.
Demonstrate the diagnosis, not the folklore: rule out simpler causes such as an unclosed resource, take before-and-after tracemalloc snapshots, work out what fraction of the heap the containers even represent, and say plainly when the saving is not worth the change.
Own the tradeoff for the team: whether a memory win of this size justifies less readable positional records across a service several people maintain, when to reach for a named-field type instead, and whether the real answer is to stop holding every record in memory at all.
This is a measurement question wearing a language question's clothes, and the honest answer starts by refusing to answer it from theory. ### The setting A search-index rebuilder streams documents out of shard files and holds one small fixed-shape record per document — id, revision, shard name, byte length, status — while it builds postings. Resident memory climbs through the run. A four-person team owns the service, and their first theory is a resource left unclosed: shard file handles accumulating across batches. That is worth ruling out in minutes (count open descriptors, check that every reader is opened under a context manager), but a handle leak produces a small, linear drip and eventually a descriptor-exhaustion error, not gigabytes of resident memory. When the profile says the memory is in ordinary Python objects and the object count tracks the document count, the per-record container becomes a fair suspect. ### The numbers, on a 64-bit CPython 3.14 build ```python import sys rec_t = ("index-07", 4, True, 1024, "ok") rec_l = ["index-07", 4, True, 1024, "ok"] grown = [] for field in rec_t: grown.append(field) print(sys.getsizeof(rec_t), sys.getsizeof(rec_l), sys.getsizeof(grown)) # 88 104 120 ``` Three numbers, three reasons. The tuple is a single object holding its five item pointers inline at exactly the right size. The list is a header object plus a separately allocated array of pointers — two allocations, more header. And a list assembled by appending carries spare capacity it asked for on your behalf, so it is larger again than the literal. Multiply honestly. Sixteen bytes saved per record is 160 MB at ten million records, thirty-two is over 300 MB; at ten thousand records it is a third of a megabyte and not worth a code review. There is also a smaller construction win: a literal tuple of constants is stored in the code object's constants and merely loaded, while a literal list must be rebuilt on every call because a mutable object cannot be shared between calls. In a per-document hot loop that is measurable, though it is rarely the reason anyone makes the change. ### Why the switch often disappoints `sys.getsizeof` is shallow. It reports the size of the container object alone, never the objects it references. If each record holds five distinct strings averaging forty bytes, the referenced objects are several hundred bytes and the container is the minority of the total — shaving 16 bytes off a 300-byte record is a five percent win, not the one the ticket promised. The interventions that actually move that needle are different in kind: sharing repeated string values instead of holding a distinct object per record, storing ids as integers rather than text, or not holding all records at once and streaming through them instead. So the sequence is: take a `tracemalloc` snapshot at a steady point, take another after the change, and diff them. Per-object arithmetic tells you the ceiling of a change; only a snapshot tells you what fraction of the heap you were even aiming at. ### The benefit that is not about bytes Often the stronger reason to make records tuples is that a tuple of hashable items is hashable. That makes `seen = set()` deduplication of records possible at all, allows a record to key a dict directly, and lets the rebuilder detect re-emitted documents without inventing a separate key. If the rebuilder currently builds a string key by joining fields, the tuple both removes that work and removes the ambiguity of a chosen separator. And there is a correctness benefit: a record that cannot be appended to cannot accidentally grow a sixth field halfway through a pipeline, which is exactly the sort of bug that survives review in a positional data structure. ### The costs to weigh Positional records get less readable as they travel. `record[3]` at the far end of a pipeline is a review hazard, and on a four-person team where not everyone touches the rebuilder weekly, that is a real maintenance cost against a memory win you should already have measured. If the records are long-lived and widely handled, a named-field record type keeps the fixed shape and the readability, at a slightly larger footprint than a bare tuple. Also: if any stage genuinely mutates a record in place, converting to tuples means rebuilding a whole record for every edit, and that trade can cost more than it saves. ### The answer to give Worth it when records are numerous, small, fixed-shape and never mutated after construction — and when a snapshot shows containers, not their contents, are a meaningful slice of the heap. Not worth it as a blanket style rule, and never worth it as a substitute for measuring.
- The team measures a 5% drop after converting records to tuples. Where did the rest of the memory go?Into the objects the records point at. `sys.getsizeof` is shallow, so per-container arithmetic only ever bounded the container slice of the heap. If each record holds several distinct strings, those dominate. The next moves are about the contents, not the container: share repeated string values rather than holding one object per record, store ids as integers, or stop holding every record at once and stream instead.
- Beyond memory, what does making the record a tuple buy the rebuilder?Hashability. A tuple of hashable fields can go straight into a set for deduplication or key a dict, so re-emitted documents can be detected without building a joined string key and choosing a separator that no field may contain. It also freezes the record's shape: no stage downstream can append a sixth field mid-pipeline, which is a bug class that survives review in positional data.
- When would you refuse to make this change?When the records are few — the saving is a rounding error under a hundred thousand records. When a pipeline stage genuinely edits a record in place, since every edit then rebuilds the whole record. And when the record has enough fields that positional access hurts readability for the people maintaining it; there, keep the fixed shape but give the fields names rather than trading review safety for bytes.
saying these in an interview costs you the question
- Switches to tuples without measuring the heap first
- Treats sys.getsizeof as reporting the full recursive size
- Claims tuples save memory proportional to the data they reference
- Expects a large win at a few thousand records
- Ignores readability cost of long positional records
- Assumes an unclosed file handle explains gigabytes of RSS