skip to content

Pickling Constraints and Payload Cost

Everything crossing a process boundary is pickled, so lambdas, closures and sockets refuse to travel and big payloads are slow. Interviewers want the workaround: module-level functions, larger chunks.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why does multiprocessing.Pool refuse a lambda as its worker function?

level: juniorimportance: must knowfreq 65%

answer

  1. The worker is a separate process
  2. Work has to travel as bytes
  3. Pickle saves a name, not code
  4. A lambda has no importable name
  5. Hoist it to module level

basics

~20 s

A pool sends the function to its worker processes by pickling it, and pickle stores a function only as a module-plus-qualified-name reference. A lambda has no importable name, so the send fails. Use a module-level function instead.

solid answer

~40 s

Worker processes have their own memory, so a pool must describe each task in bytes: the callable and its arguments are pickled in the parent, piped to a worker and rebuilt there. Pickle never serializes code for a function — it writes the function's `__module__` and `__qualname__` and re-imports that name in the child. A lambda's `__qualname__` is the literal `<lambda>`, which is not an attribute of any module, so `pickle.dumps` raises `PicklingError`; the same is true of a nested `def` or a closure, whose qualname contains `<locals>`. What does travel is a module-level `def`, a `functools.partial` wrapping one, a bound method of a picklable instance, or a callable instance of an importable class. The usual fix is to hoist the lambda to module level and pass what it captured as arguments.

code

python · 14 lines
python
import pickle


def double(x):
    return x * 2


print(double.__module__, double.__qualname__)
print(len(pickle.dumps(double)), "bytes for the reference")

try:
    pickle.dumps(lambda x: x * 2)
except (pickle.PicklingError, AttributeError) as exc:
    print(type(exc).__name__, exc)

go deeper

for a junior

Recall the rule and the fix: a worker function must be defined at module level because it is sent by name, not by code. Recognising the PicklingError message and hoisting the lambda is the whole expected answer here.

for a middle

Explain the mechanism: pickle writes __module__ plus __qualname__, the child re-imports that name, and <lambda> or <locals> can never be looked up. Name the working alternatives — functools.partial, a callable class instance, a bound method.

for a senior

Show you know the boundary is wider than the function: arguments and results pickle too, so locks, sockets and handles fail the same way. Be ready to say why multiprocessing.Process under fork behaves differently, and what 3.14's forkserver default changed.

for a principal

Own the design consequence: a codebase whose parallel work is expressed as closures cannot be moved across a process boundary later. Argue for module-level, picklable-in, picklable-out task functions as a standard that keeps thread, process and interpreter execution interchangeable.

## Workers do not share your memory A `multiprocessing.Pool` starts separate operating-system processes, each with its own heap, its own module table and its own interpreter state. Nothing you built in the parent is visible there by name. So when you hand a pool a callable and a sequence, the parent must: 1. *describe* every task in bytes, 2. push those bytes through a pipe, 3. and let a worker rebuild them. That description is produced by `pickle`, and every constraint in this question follows from how pickle treats functions. ## Pickle stores functions by reference, not by code For a plain function, pickle does not serialize the bytecode, the constants or the closure cells. It writes down two strings: the function's `__module__` and its `__qualname__`. Unpickling in the worker means "import that module, then look up that dotted name". That is why the child process needs the same source tree importable — it re-imports the function rather than receiving it. It is also why editing the function between pickling and unpickling would silently change what runs: **the reference is a name, not a snapshot.** ## Why a lambda cannot be referenced A lambda's `__qualname__` is the literal text `<lambda>`, and no module has an attribute called `<lambda>`. Assigning it to a module-level name does not rescue it, because pickle records the qualname the object carries, not the variable you happened to bind it to — `square = lambda x: x * x` at module level still fails. The same reasoning rules out: - a nested `def` (qualname `outer.<locals>.inner`), - any closure, - and a class defined inside a function: `<locals>` can never be reached by attribute lookup from a module object. The message you actually see is `PicklingError: Can't pickle <function <lambda> at 0x...>: it's not found as __main__.<lambda>`, and for locally defined objects you may instead get an `AttributeError` complaining it cannot pickle a local object. ## What can be sent A **module-level `def`** is the boring, correct answer. Beyond that: - **`functools.partial`** around a module-level function pickles as the wrapped reference plus its bound arguments, which is the standard way to "capture" a value without writing a closure. - **An instance of an importable class that defines `__call__`** pickles as a class reference plus its `__dict__`, so a callable object carries its configuration across cleanly. - **A bound method of a picklable instance** also works — pickle records the instance plus the method name. - **Builtins and functions from importable third-party modules** are fine for the same reason. ## Arguments and results are pickled too Fixing the callable is only half of it. Every argument in the sequence, and every value the worker returns, goes through the same machinery. A perfectly good module-level target still fails the moment you pass it a `threading.Lock`, a `socket.socket`, an open file object, a generator, or a database handle, because those objects deliberately refuse to be pickled: their meaning is tied to one process's kernel resources. If the worker needs such a resource, it must create its own inside the worker, not receive one. ## The fork exception, and why 3.14 matters - `multiprocessing.Process(target=...)` behaves differently from a pool **under the `fork` start method**: the child is a memory copy of the parent, so the target is inherited rather than pickled, and a lambda target genuinely works there. - **Under `spawn` and `forkserver`** the child is a fresh interpreter, so target and arguments must pickle. Since Python 3.14 the default start method on Unix platforms other than macOS is `forkserver` (macOS and Windows already used `spawn`), so code that quietly depended on fork inheritance now raises where it previously ran. A pool is unaffected by that distinction — pool tasks always travel a pipe as pickled bytes, whatever the start method, so a lambda has never worked there. ## The practical habit Write the worker function at module level, give it only picklable parameters, and bind extra configuration with `functools.partial` or by making the target a small callable class. If you find yourself wanting a closure, that is a signal the captured state should become an explicit argument — which is also what makes the function testable in-process, without a pool at all.

  • Which callables can you pass to a pool instead of a lambda?
    Anything pickle can name or rebuild: a module-level `def`; a `functools.partial` wrapping one, which pickles as the reference plus its bound arguments; an instance of an importable class defining `__call__`, which pickles as class reference plus `__dict__`; and a bound method of a picklable instance, recorded as the instance plus the method name. A lambda bound to a module-level name is still not one of them — pickle uses `__qualname__`, which is `<lambda>`.
  • Does multiprocessing.Process behave the same way as a pool here?
    Not under `fork`. There the child is a memory copy of the parent, so the target and arguments are inherited rather than pickled and a lambda target works. Under `spawn` and `forkserver` they must pickle. Pool tasks always go through a pipe as pickled bytes regardless of start method, so a lambda never works there. Since 3.14 the Unix default (outside macOS) is `forkserver`, which removes that loophole by default.
  • The target is module-level and it still raises a pickling error. What now?
    Look at the arguments and the return value, not the function. Every argument is pickled on the way out and every result on the way back, so a `threading.Lock`, a `socket.socket`, an open file object, a generator or a live connection handle in the payload fails identically. Bisect it with `pickle.dumps(arg)` on each argument in the parent, then have the worker create the offending resource itself instead of receiving one.

Pickling a function is like mailing a recipe as "page 42 of the blue cookbook" rather than copying out the steps. It only works if the other kitchen owns that book and the page has a number — and a lambda is a note scribbled on a napkin with no page number at all.

saying these in an interview costs you the question

  • Says lambdas are rejected because they are too slow
  • Thinks pickle copies the function's bytecode to the child
  • Blames the GIL for a pickling error
  • Believes assigning the lambda to a module-level name fixes it
  • Assumes workers can see the parent's globals by name
  • Forgets that arguments and return values are pickled too

context

open as a page

How do you make an object holding a threading.Lock picklable for a worker process?

level: middleimportance: should knowfreq 40%

basics

~20 s

Define getstate on the class to drop the unpicklable attribute from the state it returns, and setstate to rebuild that attribute after the object is restored in the child. Only the data needs to travel; the lock is recreated locally.

open as a page

Why did multiprocessing.Pool slow down a translation-memory updater that ships each record's full context?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Every argument and every result is pickled, piped and unpickled, so a task that ships a large context but does milliseconds of work pays more at the boundary than it saves. Send small handles and batch the records.

open as a page

Why can a custom exception raised in a multiprocessing worker fail to reach the parent?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

A worker's exception is pickled, and rebuilding it calls the class with its args tuple. If a custom init takes parameters it never forwards to the base class, reconstruction raises TypeError and the real failure is lost.

open as a page