skip to content

Why does `pickle.dumps` reject a lambda but accept a module-level function?

level: juniorimportance: should knowfreq 45%

answer

  1. Code is not copied into the payload
  2. Two strings are stored, not bytecode
  3. Module name plus qualified name
  4. `<lambda>` resolves to nothing
  5. Renaming breaks yesterday's files

basics

~20 s

pickle does not serialise a function's code. It stores the function's module name and qualified name and looks that pair up again on load. A lambda's qualified name is <lambda>, which resolves to nothing, so dumping raises pickle.PicklingError.

solid answer

~40 s

Functions and classes are pickled **by reference**, not by value: the payload holds `__module__` and `__qualname__`, and loading imports that module and fetches that attribute. A module-level `def` has a name the lookup can resolve; a lambda's qualified name is `<lambda>` and a nested function's qualified name points inside another function's scope, so neither can be found again — `pickle.dumps` fails immediately with `pickle.PicklingError`, and the message literally says it is "not found as" that name. The same rule governs the callable inside a `__reduce__` tuple, and it has a sharp operational edge: if you rename or move a class after writing pickles of its instances, the old files stop loading with an `AttributeError` about the missing name, because the data was never self-contained.

code

python · 11 lines
python
import pickle

def scale(x):
    return x * 2

print(pickle.dumps(scale))          # the bytes hold "__main__" and "scale"

try:
    pickle.dumps(lambda x: x * 2)
except pickle.PicklingError as exc:
    print(exc)                      # not found as __main__.<lambda>

go deeper

for a junior

Know that pickling a function or class writes down where to find it, not what it does, and that a lambda has no findable name. Recognise pickle.PicklingError on dump versus AttributeError on load.

for a middle

Explain the module-plus-qualified-name mechanism and the verification step that catches a mismatched name, and connect it to why pool tasks must live at module level.

for a senior

Show that you treat stored payloads as a compatibility surface: plan aliases or a migration before a rename, and argue for a schema-based format when data has to outlive the code that wrote it.

for a principal

Own the policy question of whether Python-specific payloads belong in durable storage or on the wire at all, given that every such payload pins your module layout and executes on load.

### Two different things happen inside one call When `pickle.dumps` walks an object graph it treats data and code very differently. Data — numbers, strings, containers, instances — is written out value by value. Functions and classes are written out as **references**: the payload records the object's `__module__` and its `__qualname__`, and nothing else. You can see this directly in the bytes of a pickled module-level function: the module name and the function name appear as plain readable text, and there is no bytecode anywhere. On load, the name is resolved the other way round: the recorded module is imported, the recorded qualified name is walked attribute by attribute, and whatever is found there becomes the function. That is the whole mechanism, and every rule around it follows from it. ### Why a lambda cannot work Every function object carries `__qualname__`. For `def scale(x)` at module level that is `"scale"`, which resolves. For a lambda it is `"<lambda>"`, which is not a legal attribute name and resolves to nothing — even if you bind the lambda to a module-level name, its qualified name is still `<lambda>`, so binding does not save it. The pickler does not just look at the name; it looks the name up and checks that the object it finds **is** the object being pickled. When the check fails you get `pickle.PicklingError` with a message of the form "it's not found as `__main__.<lambda>`". The same reasoning rules out a function defined inside another function (its qualified name contains `<locals>`), a class synthesised at runtime, and a decorated function whose wrapper was not bound at module level. The fix is always the same: give the callable a real, stable, module-level home. ### Instances are a mix of both rules An instance of your own class pickles as a *reference to the class* plus enough data to rebuild the instance. That is why pickling an instance of a class defined inside a function fails while pickling an instance of a top-level class succeeds, and why the class body itself — its methods, its source — never travels in the file. The receiving process must already have the same code importable under the same name. ### The consequence people actually get bitten by Because the reference is a name, **renaming or moving a class or function invalidates every pickle already written**. The failure appears at load time, in a completely different program run, as `AttributeError: module 'jobs' has no attribute 'Report'`. Nothing is wrong with the bytes; the name they point at no longer exists. Ordinary refactoring tools rename the definition and every call site and leave stored data silently broken. If you must keep old payloads loadable, the cheapest mitigation is to keep the old name resolvable: import the class back into its old module under its old name so the lookup still lands on something. A second option is a migration pass that loads with the compatibility alias in place and re-writes the data under the new name. The durable answer is not to use pickle as an archive format at all — a text or schema-based format survives refactoring because it never references your code. ### Why this also shows up in process pools Sending work to another process serialises the callable and its arguments through exactly this machinery, which is why "the function must be defined at module level" is a standing rule for pool tasks. With the `spawn` and `forkserver` start methods the child process re-imports your module and resolves the name there, so a callable that cannot be named cannot be dispatched. On Python 3.14 that matters more than it used to: `multiprocessing`'s default start method is now `forkserver` on Unix platforms other than macOS, where macOS and Windows already defaulted to `spawn`, and `fork` must be requested explicitly. ### Versions Protocol 4, the default from Python 3.8 through 3.13, introduced qualified-name lookup so that nested classes (`Outer.Inner`) can be referenced. `pickle.DEFAULT_PROTOCOL` is 5 on Python 3.14. None of this changes the underlying rule: code is referenced by name, never embedded.

  • A class was renamed after its instances were pickled to disk. Which error appears, and at which end?
    At load time, in whatever process reads the file: `AttributeError`, saying the module has no attribute of that name. Dumping succeeded long ago and the bytes are intact; the reference simply no longer resolves. The clue is that the message names the module and the old class name, not a corruption or a protocol problem.
  • How would you keep old payloads loadable after moving a class to a new module?
    Leave the old name resolvable: import the class into its former module under its former name so the qualified lookup still lands on it, then run a migration pass that loads everything and re-writes it. Treat the alias as temporary, and record why it exists, because it is now a compatibility surface driven by stored data rather than by code.
  • Why must a function handed to a process pool be defined at module level?
    It is serialised by module name plus qualified name like any other function reference. With the `spawn` and `forkserver` start methods the worker re-imports the module and resolves that name, so a lambda, a closure, or a locally defined function cannot be dispatched. Python 3.14 made `forkserver` the default on Unix other than macOS, so this bites on platforms where `fork` used to hide it.

A pickle is a note saying "the thing on shelf B in room 4", not a copy of the thing. Move the shelf and the note still reads fine but points at nothing.

saying these in an interview costs you the question

  • Thinks pickle stores the function's bytecode or source
  • Says binding a lambda to a module name makes it picklable
  • Believes renaming a class only affects new payloads
  • Treats a load-time AttributeError as file corruption
  • Assumes a pickle file is self-contained and readable anywhere
  • Expects the class body to travel with the instance

context