skip to content

What does pickle actually store for a class instance, and what can it not store?

level: middleimportance: should knowfreq 50%

answer

  1. Two halves: which class, what state
  2. The class travels as a name
  3. No code is ever serialized
  4. __init__ does not run on load
  5. Live OS handles have nothing to store

basics

~20 s

For an ordinary instance pickle stores a reference to its class by module and qualified name, plus the instance state, normally its dict. No code is stored, so live OS resources and unnamed callables cannot be pickled.

solid answer

~50 s

Pickling an instance writes two things: a *reference* to its class — the module name and the qualified name, resolved by import on the way back in — and the instance's state, which by default is its `__dict__` (plus a second dict for `__slots__` values). No source code is stored, and `__init__` is not called on reload; the state is installed directly, or handed to `__setstate__` if the class defines one. That is why two families of objects fail. Anything holding live C-level or OS state — `threading.Lock`, an open file, a socket, a generator — raises `TypeError`. Anything not findable by qualified name — a lambda, a function or class defined inside another function, a dynamically built class — raises `PicklingError`. The fix for the first family is `__getstate__` to drop the attribute and `__setstate__` to rebuild it.

code

python · 20 lines
python
import pickle, threading


class Index:
    def __init__(self, path):
        self.path = path
        self.lock = threading.Lock()          # not picklable

    def __getstate__(self):
        state = self.__dict__.copy()
        del state["lock"]
        return state

    def __setstate__(self, state):
        self.__dict__.update(state)           # __init__ is never called
        self.lock = threading.Lock()


restored = pickle.loads(pickle.dumps(Index("/data/genomes")))
print(restored.path, type(restored.lock).__name__)

go deeper

for a junior

Know that pickle saves an object's data and a pointer to its class, not the class itself, and that lambdas, open files and locks simply cannot be pickled.

for a middle

Explain the two halves — class by qualified name, state usually dict — and demonstrate getstate and setstate dropping and rebuilding an attribute, noting that init is skipped on load.

for a senior

Show the operational consequence: class names and module paths become part of the stored format, so refactoring breaks old artefacts, and long-lived data deserves an explicit schema instead of pickled objects.

for a principal

Own where object-graph serialization is allowed at all — internal caches and process pools yes, durable archives and cross-service payloads no — and the migration plan for artefacts already written against class names.

### The reduce protocol, applied to an ordinary object `pickle` has built-in encodings for a fixed set of types: `None`, bools, ints, floats, `str`, `bytes`, tuples, lists, dicts, sets, and a few more. For everything else it asks the object how to rebuild itself, by calling `__reduce_ex__(protocol)`. Instances of ordinary classes do not define that method, so they inherit the default, which produces roughly: *call this reconstructor with the class object, then install this state*. That has two halves worth separating, because most pickle bugs live in one or the other. **The class is stored by reference.** What goes into the stream is the module name and the qualified name — the module where the class was defined, plus `Record` — and on load the unpickler imports that module and looks the attribute up. No class body, no methods, no source is serialized. The practical consequences are immediate: the reading process must be able to import the same name; renaming or moving a class silently invalidates every pickle you have already written; and a class defined inside a function has a qualified name containing `<locals>`, which cannot be looked up, so it cannot be pickled at all. **The state is a dict.** By default it is `object.__getstate__()`, which returns the instance `__dict__` (and, for a class using `__slots__`, a two-tuple of the dict and a slots-values dict). Since Python 3.11 `object.__getstate__` exists as a real default implementation you can call and extend; before that the default lived only inside the pickle machinery. On load, the instance is created **without calling `__init__`** and the state is applied — merged into `__dict__`, or passed to `__setstate__` when the class defines one. Forgetting that `__init__` is skipped is the classic source of "my object came back with an attribute missing that the constructor always sets": the attribute was absent from the state, and nothing re-ran to create it. ### What cannot be pickled, and why the error differs Two distinct failures, with two distinct exception types: ```python import pickle, threading pickle.dumps(threading.Lock()) # TypeError: cannot pickle '_thread.lock' object pickle.dumps(lambda x: x) # PicklingError: ... not found as __main__.<lambda> ``` `TypeError` means the object holds state that has no meaning in another process or another run: a lock owned by this interpreter, an open file with an OS file descriptor and a position, a socket, a running generator or a frame. There is nothing to serialize; a copy would be a lie. `PicklingError` means the object *is* fine conceptually but is not reachable by a qualified name. Functions and classes are always pickled by reference, so a lambda, a closure, a class produced by a factory, or anything defined inside a function has no name for the stream to carry. ### Controlling it: `__getstate__`, `__setstate__`, `__reduce__` The idiomatic fix for the first family is to define `__getstate__` to return a state dict with the unpicklable attributes removed, and `__setstate__` to rebuild them: ```python def __getstate__(self): state = self.__dict__.copy() del state["lock"] return state def __setstate__(self, state): self.__dict__.update(state) self.lock = threading.Lock() ``` Below that sit two heavier tools. `__reduce__` lets you specify the callable and arguments used to rebuild the object outright — the same hook that makes unpickling untrusted data dangerous, and the one to reach for when reconstruction needs a factory rather than a state dict. `copyreg.pickle` registers a reduction function for a type you do not own, without touching its class. `__getnewargs_ex__` supplies arguments to `__new__` for immutable types that need them. One more thing that surprises people: these hooks are shared with `copy.copy` and `copy.deepcopy`, which use the same reduce protocol. Adding `__getstate__` to strip a lock also changes how the object copies — usually for the better, but it is a wider blast radius than "how it pickles". ### The design rule Pickle round-trips *state*, and it does so by trusting that the same code is importable on both sides. That makes it excellent within one codebase and one interpreter version, and fragile as an archive format: your class layout becomes part of the file format, and refactoring breaks it. If the bytes must outlive the code that wrote them, define an explicit schema and serialize fields, not objects.

  • An instance comes back from pickle.loads missing an attribute the constructor always sets. What happened?
    Unpickling does not call `__init__`. The instance is allocated and the stored state is installed directly, so any attribute the constructor computes but that was not in the state dict simply is not there. It usually means a `__getstate__` that dropped the attribute without a matching `__setstate__` to recreate it, or a pickle written before the attribute existed. Restore it in `__setstate__`, which is the only hook that runs on the way back in.
  • You rename a pickled class from Record to AnnotationRecord. What breaks, and what are the options?
    Every existing pickle carries the old module and qualified name, so loading raises an import or attribute lookup failure. Options: keep an alias bound to the old name in the old module so the lookup still resolves, override `Unpickler.find_class` to map old names to new ones during a migration, or re-write the stored artefacts by loading with the old code and dumping with the new. The lesson is that class names are part of the on-disk format.
  • Why does adding __getstate__ also change how copy.deepcopy treats the object?
    `copy` and `pickle` share the same reduce protocol: `copy.copy` and `copy.deepcopy` consult `__reduce_ex__`, and therefore `__getstate__`/`__setstate__`, when a type has no dedicated copier. So stripping an attribute for pickling also strips it from copies, and rebuilding it in `__setstate__` runs there too. That is usually what you want for a lock or a connection, but it is a wider effect than the name suggests.

saying these in an interview costs you the question

  • Thinks pickle stores the class body or its source
  • Expects __init__ to run when an instance is unpickled
  • Believes any object at all can be pickled
  • Confuses __getstate__ with __getattr__
  • Says renaming a class is safe for existing pickles

context