skip to content

What do __getstate__ and __setstate__ do when Python pickles and unpickles an object?

level: juniorimportance: should knowfreq 40%

answer

  1. Two hooks on opposite sides of a pickle
  2. One decides what is saved
  3. The other rebuilds from what was saved
  4. The default is the instance dict
  5. No constructor runs on the way back

basics

~10 s

getstate returns the data written into the pickle, which by default is the instance dict. setstate receives that data on unpickling and rebuilds the object from it. Unpickling never calls init.

solid answer

~50 s

When `pickle.dumps` serialises an instance it asks the object for its **state**. If the class defines `__getstate__`, whatever that method returns is stored; otherwise the inherited `object.__getstate__` supplies the default, which for an ordinary class is the instance `__dict__`. Coming back, `pickle.loads` allocates a blank instance and **does not call `__init__`**; it then hands the stored state to `__setstate__` if the class defines one, and otherwise merges the state dict into the new instance's `__dict__`. So `__getstate__` decides what is worth saving and `__setstate__` decides how to come back to life. Anything you leave out of the state simply does not exist on the restored object unless `__setstate__` puts it back, because no constructor runs to recreate it. The state does not have to be a dict; any picklable object will do, as long as `__setstate__` knows how to read it.

code

python · 10 lines
python
import pickle

class Widget:
    def __init__(self, name):
        print("__init__ ran")
        self.name = name

w = Widget("toggle")
clone = pickle.loads(pickle.dumps(w))
print(clone.name)

go deeper

for a junior

Recall the pairing: __getstate__ produces what gets stored, __setstate__ consumes it on the way back, and the default state is the instance __dict__. Knowing that __init__ does not run on unpickling is the single fact most often checked.

for a middle

Be ready to explain the mechanics: the unpickler allocates a blank instance, then either calls __setstate__ or merges the state into __dict__. Say why a dropped attribute stays missing, and what happens if the state is not a mapping.

for a senior

Show that you treat the state as a stored format with a lifetime. Talk about restores that tolerate fields written by an older version of the class, and about re-establishing constructor invariants on the load side rather than hoping callers never touch them.

for a principal

Own the policy question of whether pickled objects should be a persistence format at all in your systems, given that the byte stream couples storage to class layout and module paths, and argue for an explicit schema where data must outlive the code.

## What "state" means to pickle The `pickle` module does not serialise a class body or any source code. It serialises two things: a way to produce a blank instance of the class, and that instance's **state** -- the data that makes this instance different from any other instance of the same class. `__getstate__` and `__setstate__` are the two hooks a class defines to take control of that state. ## On the way out `pickle.dumps(obj)` asks the object for its state. If the class defines `__getstate__`, whatever that method returns is written into the byte stream. If it defines no such method, the inherited `object.__getstate__` supplies the default: the instance `__dict__` for an ordinary class. The returned state is itself serialised with exactly the same machinery, recursively, so every value inside it must also be picklable. ## On the way back `pickle.loads(data)` allocates a blank instance of the class -- the equivalent of `object.__new__(cls)` -- and then applies the state. **It does not call `__init__`.** If the class defines `__setstate__`, the unpickler calls it with the stored state and steps back: whatever that method does *is* the restoration. If the class defines no `__setstate__`, the default restore merges the state mapping into the instance `__dict__` (and assigns slot values when the state is the two-element tuple form). Hand it something that is neither -- a string, a three-element tuple -- with no `__setstate__` in sight and the unpickler raises `pickle.UnpicklingError` with the message "state is not a dictionary". ```python import pickle class Flag: def __init__(self, name, on): print("__init__ ran") self.name = name self.on = on f = Flag("new_checkout", True) print(f.__getstate__()) # {'name': 'new_checkout', 'on': True} clone = pickle.loads(pickle.dumps(f)) # prints nothing: no __init__ print(clone.name, clone.on) ``` ## Why the constructor is skipped A pickle is a snapshot of an object *after* construction. Re-running `__init__` would require the original constructor arguments, which the stream does not carry, and would then overwrite the very state that was restored. So the unpickler deliberately bypasses it. The practical consequence is the one interviewers push on: **every invariant that `__init__` establishes and that is not part of the saved state has to be re-established by `__setstate__`**. A cached lookup table, a compiled form of a rule, a connection, a lock -- if `__getstate__` left it out and `__setstate__` does not put it back, the restored object is missing an attribute and the failure surfaces later, at first use, as an `AttributeError` far from the unpickling call. ## The asymmetries worth naming * `__getstate__` runs only on the dump side; `__setstate__` runs only on the load side. Neither one calls the other, and defining one does not oblige you to define the other -- although in practice, once you start dropping attributes, you need both. * The default state is the *instance* `__dict__`, not the class's attributes. Class attributes come back because the restored object is an instance of the same class, not because they were stored. * The state may be any picklable object: a dict, a tuple, a single integer, a nested structure. A non-dict state simply obliges you to write the matching `__setstate__`. * Old byte streams outlive class changes. State written months ago by an older version of the class will be handed to today's `__setstate__`, which is why defensive restores read fields with a mapping lookup that tolerates a missing key rather than indexing blindly. ## When you need these hooks at all Most classes need neither: an object whose attributes are plain data pickles correctly with the default state. You reach for `__getstate__` when the instance holds something that cannot be serialised or should not be, and for `__setstate__` when the restored object needs work done to be usable again. That pairing -- drop on the way out, rebuild on the way in -- is the entire idiom, and it is why the two methods are almost always introduced together even though they are independent hooks on opposite sides of the round trip. ## A worked round trip Take a small settings object that caches an expensive lookup table. `__getstate__` returns `{"source": self.source}` and leaves the table out, because it can be recomputed and is far larger than the value it stores. `__setstate__` receives that one-key mapping, assigns `self.source`, and recomputes the table. Dump it, load it in a fresh process, and the restored object is behaviourally identical to the original while the byte stream stayed small. Remove `__setstate__` and the load still succeeds -- the default restore merges `{"source": ...}` into `__dict__` -- but the table is gone, and the object fails the first time anything reads it. That asymmetry is the whole reason the two hooks are taught as a pair: the dump side controls size and picklability, the load side controls correctness. ## Where the hooks are not the answer If an object is genuinely not meant to be reconstructed from data -- a handle to something owned by the operating system, a connection whose validity is tied to a session -- writing state hooks to paper over it is a design smell rather than a fix. The honest options are to refuse to pickle the object, or to pickle a small descriptor of how to obtain an equivalent one and let the caller decide when to do that. Interviewers who ask about these hooks are usually probing whether you know the difference between data that can be serialised and a live resource that cannot.

  • If __init__ is skipped, what actually creates the instance during unpickling?
    The unpickler allocates a bare instance of the class, equivalent to `object.__new__(cls)`, and then applies the stored state to it. Nothing you wrote in `__init__` runs, which is why any invariant set up there and not carried in the state has to be re-established in `__setstate__`.
  • Does __getstate__ have to return a dictionary?
    No. The state can be any picklable object -- a tuple, a string, a number, a nested structure. But the *default* restore path only understands a mapping or the two-element slots tuple; give it anything else without defining `__setstate__` and `pickle.loads` raises `pickle.UnpicklingError` saying the state is not a dictionary.
  • Is __getstate__ ever called while unpickling?
    No. `__getstate__` runs only on the dump side and `__setstate__` only on the load side. They are independent hooks, and a class may define either one alone -- though once you start omitting attributes on the way out you almost always need the matching rebuild on the way in.

Think of moving house: getstate decides what goes in the boxes, setstate unpacks them in the new place. Nobody rebuilds the furniture from the original flat-pack instructions.

saying these in an interview costs you the question

  • Claims unpickling calls __init__ again
  • Thinks pickle stores the class source code
  • Says the default state includes class attributes
  • Believes __getstate__ also runs during unpickling
  • Assumes dropped attributes reappear automatically on load
  • Insists the state must always be a dictionary

context