How does CPython's key-sharing dict make instance attributes cheaper than a dict?
answer
- The names repeat; the values do not
- One table, many instances
- Named in a PEP about dictionaries
- Since 3.11 it starts inline in the object
- Why attributes belong in __init__
basics
~10 sAlmost every instance of a class carries the same attribute names, so CPython stores those names once in a keys table owned by the class and gives each instance only its own values array.
solid answer
~40 sA dict's memory is dominated by its key table: the sparse index array plus the key pointers and hashes. Instances of one class nearly always share the same attribute names, so PEP 412 (Python 3.3) let many instance dicts point at a single shared key table while each holds only its own values array. Since 3.11 CPython goes further: the values live inline in the object itself and no dict object exists until something actually touches `__dict__`. On 3.14, an object with three attributes costs about 136 bytes against roughly 224 for an equivalent three-key dict, and touching `__dict__` adds about 64 more per instance to materialize the dict header. Sharing depends on instances agreeing on their attribute set, which is the real reason for the old advice to assign every attribute in `__init__`.
code
python · 18 linesimport tracemalloc
class Point:
def __init__(self, x, y, z):
self.x = x
self.y = y
self.z = z
N = 100_000
tracemalloc.start()
objs = [Point(i, i, i) for i in range(N)]
per_obj = tracemalloc.get_traced_memory()[0]
dicts = [{"x": i, "y": i, "z": i} for i in range(N)]
per_dict = tracemalloc.get_traced_memory()[0] - per_obj
tracemalloc.stop()
print(round(per_obj / N), round(per_dict / N))go deeper
Know only that attributes on an object are stored more cheaply than the same data in a dictionary, and that setting all of an object's attributes in __init__ is the conventional style. The mechanism is not expected of you.
Explain that instances of one class share a single table of attribute names while each holds just its own values, and that since 3.11 those values start out inline in the object with __dict__ created only on demand.
Show that you would measure rather than assert. Compare per-instance cost against the dict equivalent, explain what makes an instance lose the shared layout, and place __slots__ correctly as the next step once counts justify it.
Own the representation decision at scale: when per-object attributes stop being the right container at all and a columnar or record-oriented layout wins, and how to keep such a choice measurable rather than folklore in the codebase.
### The observation behind the optimization Every instance of a class tends to carry the same attribute names. A million points all have `x`, `y` and `z`. If each instance owned a full dict, each would separately store an index array, three key pointers and three cached hashes — a million identical copies of the same three strings' bookkeeping. PEP 412, *Key-Sharing Dictionary*, landed in Python 3.3 to remove that duplication. ### Split tables A dict's storage divides cleanly into two parts: the **keys** part (the sparse index array plus the key pointers and hashes) and the **values**. In a split table the keys part is owned by the class and shared by every instance, while each instance holds only a compact array of value pointers, positionally matched to the shared keys. Adding `self.x` to a new instance does not insert a key anywhere — it writes into slot 0 of that instance's values array. CPython 3.11 pushed this further. The values array is stored *inline in the object* — no separate dict object is allocated at all until something asks for `__dict__`, at which point a real dict is materialized lazily and still points at the shared keys. ### What it measures ```python import tracemalloc class Point: def __init__(self, x, y, z): self.x = x self.y = y self.z = z N = 100_000 tracemalloc.start() objs = [Point(i, i, i) for i in range(N)] per_obj = tracemalloc.get_traced_memory()[0] dicts = [{"x": i, "y": i, "z": i} for i in range(N)] per_dict = tracemalloc.get_traced_memory()[0] - per_obj tracemalloc.stop() print(round(per_obj / N), round(per_dict / N)) # 136 224 ``` About 136 bytes per instance — object header, type pointer, refcount and three inline value slots — against about 224 for a plain three-key dict holding the same data, which pays for its own key table. At a million objects that gap is roughly 90 MB. Materializing `__dict__` is not free either. Touching it on every instance adds about 64 bytes each for the dict headers, and the resulting dict reports only 96 bytes rather than the 184 a standalone three-key dict reports — because the shared keys table is not counted against it. That difference *is* the sharing, visible from Python. ### How sharing is lost Sharing holds while instances agree on their attribute set. When one instance's layout diverges from the shared table — attributes added conditionally long after construction, an attribute deleted from one object, `__dict__` replaced wholesale — CPython can fall back to giving that instance a private, combined table, and its cost jumps to that of a normal dict. This is the mechanical reason behind advice that usually gets stated as style: * **Assign every attribute in `__init__`**, including the ones that are `None` for now. A consistent set keeps the shared layout. * **Do not attach attributes conditionally** to objects you create in the millions. `if flag: self.extra = ...` is a memory decision at that scale. * **Avoid touching `instance.__dict__`** on hot paths over many objects; reading it forces a dict object into existence that would otherwise never be allocated. ### Where it sits among the alternatives Key sharing is automatic, invisible and requires nothing of you. When it is not enough — genuinely millions of small, fixed-shape objects — `__slots__` removes the per-instance mapping entirely and stores attributes in fixed slots on the object, which measures smaller again, at the cost of dynamic attributes and some inheritance flexibility. The sequence to reason about is: shared-key instance attributes by default, slots when the object count justifies the rigidity, and a different representation entirely (arrays, columns, records) when even that is too heavy. ### Why this is a differentiator, not a gate Nobody is turned down for not knowing PEP 412 by name. What the question really probes is whether a candidate has ever measured object memory rather than guessed, and whether they can explain *why* the familiar advice about setting attributes in `__init__` exists. An answer that connects the style rule to a concrete interpreter mechanism, and that reaches for a measurement rather than a rule of thumb, is doing exactly what senior work on memory looks like.
- Why is `sys.getsizeof` on an instance's `__dict__` smaller than on a plain dict with the same three keys?Because a split dict does not own its keys table — the class does — and the size report only counts the table a dict owns. So the instance dict reports the header plus its values array, around 96 bytes on 3.14, while a standalone three-key dict reports about 184. The difference is exactly the shared key table you are not paying for per instance.
- What does assigning every attribute in `__init__` have to do with memory?It keeps instances agreeing on one attribute layout, which is what lets them share a single keys table. Attributes attached later on only some objects make a layout diverge, and CPython can then give that instance a private combined table costing as much as an ordinary dict. At a handful of objects this is irrelevant; at millions it is the difference between tens of megabytes.
- When would you move from key-sharing to `__slots__`?When the object count is large enough that the remaining per-instance mapping matters and the shape is genuinely fixed. `__slots__` stores attributes in fixed positions on the object with no instance mapping at all, so it measures smaller again — but it costs dynamic attribute assignment and some multiple-inheritance freedom. Reach for it after measuring, not by default.
It is a printed form: the field labels are typed once on the master template and every copy carries only the handwritten answers, so a thousand submissions do not each reprint the labels.
saying these in an interview costs you the question
- Thinks every instance owns a private copy of its attribute names
- Says an instance's attributes always cost the same as a dict
- Believes `__dict__` always exists as a real dict object
- Claims attaching attributes later is free at any scale
- Confuses key sharing with string interning of the names
- Guesses object memory instead of measuring it