skip to content

Why does a Python instance with three attributes cost far more memory than three raw values?

level: juniorimportance: should knowfreq 45%

answer

  1. Nothing in Python is a raw value
  2. Two machine words before your data
  3. Refcount plus type pointer, per object
  4. Attributes are pointers to boxed objects
  5. Since 3.11 values sit inline

basics

~20 s

Every Python object starts with a header - a reference count and a pointer to its type - and a normal instance carries attribute storage on top of that. The values are separate objects too, each with its own header.

solid answer

~40 s

On a 64-bit CPython build the object header alone is 16 bytes: eight for the reference count, eight for the type pointer. A class that does not declare `__slots__` gives every instance its own attribute storage plus a pointer per attribute, and the values are not stored inline as raw numbers - each `int` or `float` is a separate heap object of roughly 24-28 bytes. Since 3.11 CPython keeps a plain instance's attribute values in a compact array laid out with the object and only materializes a real dictionary when something asks for `obj.__dict__`, so the per-instance figure is smaller than older articles claim. Measured on 3.14, a two-attribute class still costs about 96 bytes per instance. The cost is charged per object, which is why it only bites at millions of them.

code

python · 11 lines
python
import sys

class Reading:
    def __init__(self, ts, value):
        self.ts = ts
        self.value = value

r = Reading(1.0, 2.0)
print(sys.getsizeof(object()))   # 16: refcount + type pointer
print(sys.getsizeof(r))          # shallow: values not included
print(sys.getsizeof(r.value))    # every float is its own object

go deeper

for a junior

Be ready to say that every Python object carries a header - a reference count and a type pointer - and that numbers are objects too, so an attribute is a pointer to another object rather than a value stored in place.

for a middle

Explain where a plain instance keeps its attributes, that the values array is laid out with the object since 3.11 and only becomes a real dictionary on demand, and why a list of a million floats costs about 32 bytes per element.

for a senior

Show you can turn this into a decision: know how to measure per-instance cost with tracemalloc instead of quoting blog figures, and know which shapes in a running service actually multiply the header a million times.

for a principal

Own the framing that per-object overhead is a data-modelling choice, not a tuning knob: whether the system stores rows as objects or columns as buffers decides the memory ceiling long before any micro-optimisation does.

### Nothing in Python is a raw value At the C level every Python object begins with a fixed header. On a 64-bit build that header is two machine words, 16 bytes: a reference count and a pointer to the object's type. `sys.getsizeof(object())` returns exactly that 16. Everything else an object needs — its attribute storage, its variable-length payload, the bookkeeping the cycle collector wants — is laid out after those two words. That header is charged **per object**, and Python creates far more objects than a systems-language programmer expects. A `float` attribute is not eight bytes of IEEE-754 sitting in the instance; it is a separate heap object of 24 bytes (header plus the double) that the instance merely points at. A small `int` is 28 bytes. So "three numbers" is really one instance plus up to three other objects plus three pointers. ### Where a plain instance keeps its attributes For a class that does not declare `__slots__`, each instance gets its own attribute storage. Historically that meant a full private dictionary per instance, which is why the folklore figures for instance overhead are so large. Two changes shrank it: * **3.3 (PEP 412) key-sharing dictionaries** — instances of the same class share one copy of the key table, so each instance stores only the values. * **3.11 inline values** — CPython stores those values in a compact array laid out with the object itself and only materializes a real dictionary object when something actually asks for `obj.__dict__` (or `vars(obj)`, or sets an attribute in a way that forces it). The consequence for interviews: a plain instance on 3.14 is *much* cheaper than a 2015 blog post claims, but it is still nowhere near "the size of its data". Measured on 3.14, a two-attribute class costs about 96 bytes per instance including the list slot that holds it; the same class with `__slots__` costs about 56. ### The container adds its own tax Objects almost never live alone. A `list` holding a million floats stores a million *pointers* — eight bytes each, in one contiguous array — and each pointed-at float is its own 24-byte object. That is roughly 32 bytes of memory for eight bytes of information, a 4x tax before you count the list's own header and its over-allocation slack. A `dict` is worse per entry, because it stores a hash, a key pointer and a value pointer. ### Why `sys.getsizeof` will not show you this `sys.getsizeof` is shallow. It asks the object's own `__sizeof__` and adds the garbage-collector bookkeeping, and it stops there: it never follows a pointer. `sys.getsizeof` on the instance above reports 48 while the instance genuinely costs about twice that once you count the values array and the objects it points at. To see the real figure you either walk the object graph yourself (deduping by `id`, since two attributes can point at one object) or you allocate a large number of instances under `tracemalloc` and divide. ### What follows from all this Three practical consequences, and interviewers are usually after at least one of them. 1. **Small objects in bulk are the expensive shape.** A million rows modelled as a million tiny instances costs the headers a million times. The same data in a handful of typed buffers, or in columns, costs the header a handful of times. 2. **Per-object cost is what `__slots__` attacks.** Declaring the attribute names up front lets the class store them at fixed offsets in the instance and drops the per-instance dictionary machinery entirely. 3. **Measure on the build you deploy.** The free-threaded build, officially supported from 3.14 (PEP 779), carries extra per-object fields so that reference counting can be thread-safe; object sizes there are not the ones you measured on the default build. ### The answer that sounds senior "Every object pays a two-word header, and Python boxes numbers, so an instance is a header plus attribute storage plus a pointer per attribute plus a separate boxed object per value. On 3.14 the attribute values live inline in the object until something asks for `__dict__`, so the per-instance figure is smaller than it used to be — but the multiplier is still per-object, which is why the cost only ever bites at scale."

  • How much of that per-instance cost does `sys.getsizeof` on the instance report?
    Very little of it. `sys.getsizeof` is shallow: it reports the object's own storage plus garbage-collector bookkeeping and never follows a pointer, so it excludes the objects the attributes point at, and on 3.11+ it does not reflect the inline values array either. For a real figure, allocate many instances under `tracemalloc` and divide, or walk the graph yourself.
  • Why does an instance of an empty class still cost more than the 16-byte header?
    Instances of a normal class are tracked by the cycle collector, which attaches bookkeeping ahead of the object, and the class reserves attribute storage laid out with the instance so that setting an attribute does not require a fresh allocation. Only a class that declares `__slots__` and defines no attributes gets close to the bare header.

Shipping one screw per parcel: the screw weighs nothing, the box and the label are the shipment. Python posts every value in its own box.

saying these in an interview costs you the question

  • Claims Python stores integers and floats as raw machine words
  • Thinks an object costs only the sum of its attribute values
  • Believes sys.getsizeof returns an instance's total memory cost
  • Says a list of floats stores the doubles contiguously
  • Assumes the header exists only on container objects

context