Why does ProcessPoolExecutor.submit reject a lambda that ThreadPoolExecutor accepts?
answer
- one worker shares your memory, one does not
- the call has to become bytes
- functions are stored by name, not code
- children re-import the entry module
- a dead worker breaks the whole pool
basics
~20 sA thread pool calls the function in the same process, so any callable works. A process pool must send the call to a child over a pipe using pickle, and pickle stores a function by its module-qualified name, which a lambda does not have.
solid answer
~50 s`ThreadPoolExecutor` runs the callable in another thread of the same process: it shares memory, so closures, lambdas, bound methods and unpicklable arguments are all fine. `ProcessPoolExecutor` sends every work item to a child process through a pipe, and the wire format is `pickle`. Pickle serializes a function by reference — module plus qualified name — so the target must be importable at module level; a lambda, a nested function or a dynamically built class fails with `pickle.PicklingError`. The same applies to the arguments and to the return value, so an open socket, a `threading.Lock` or a database connection cannot cross either. Two more process-only rules follow: the submitting code must sit behind an `if __name__ == "__main__":` guard, because a `spawn` or `forkserver` child re-imports your module; and if a worker dies outright, every pending `Future` completes with a `BrokenProcessPool` error rather than a normal exception.
code
python · 13 linesfrom concurrent.futures import ProcessPoolExecutor
def double(n):
return n * 2
if __name__ == "__main__":
with ProcessPoolExecutor(max_workers=2) as pool:
print(list(pool.map(double, [1, 2, 3])))
fut = pool.submit(lambda n: n * 2, 4)
try:
fut.result()
except Exception as exc:
print("lambda rejected:", type(exc).__name__)go deeper
Recall the one-line rule: a process pool has to send the work to another process, so the function must be a plain module-level def and its arguments must be picklable. A thread pool has no such restriction.
Explain the mechanism: pickle stores a function by module plus qualified name, arguments and results are copied across a pipe, and a spawn or forkserver child re-imports __main__, which is what the guard protects against.
Show the operational consequences: transfer cost driving the batching decision, initializer for unpicklable per-worker state, max_tasks_per_child against leaking native libraries, and treating a BrokenProcessPool as pool-wide and unrecoverable.
Own the boundary as a design constraint: what data is allowed to cross it, whether workers should pull inputs themselves instead of receiving copies, and how a hard worker death maps onto the job's restart and idempotency story.
`concurrent.futures` deliberately gives `ThreadPoolExecutor` and `ProcessPoolExecutor` the same `submit`/`map`/`Future` surface, so switching between them is close to a one-word change. What is *not* the same is the boundary the work has to cross. ### The thread pool has no boundary A `ThreadPoolExecutor` worker is another thread in your process. It shares the same heap, the same module globals, the same open file objects. `submit` just puts `(fn, args, kwargs)` in a `queue.SimpleQueue` and a worker calls it. Nothing is copied, so any callable at all works: a lambda, a closure over local state, a bound method of a live object, a `functools.partial` wrapping any of those. The cost of that convenience is the GIL and shared mutable state, but the *interface* accepts everything. ### The process pool has a serialization boundary A `ProcessPoolExecutor` worker is a separate interpreter in a separate OS process with its own memory. The parent has to describe the call to it, and the description travels as bytes down a pipe. `pickle` produces those bytes, and three separate things get pickled: the callable, the arguments, and — on the way back — the return value or the exception. Pickle does not serialize a function's code. It serializes a *reference*: the module name plus `__qualname__`, so the child can `import` the module and look the name up. That is why the rule is "importable at module level": * A `lambda` has `__qualname__` ending in `<lambda>` and cannot be looked up. `pickle.PicklingError`. * A function defined inside another function is equally unreachable, for the same reason. * A class defined inside a test function, or built by `type()` at runtime, fails the same way. * An instance method works only if both the class and the instance pickle. Arguments have their own constraints. Anything holding an OS resource or a synchronization primitive is unpicklable: an open socket, a `threading.Lock`, a live database connection, a generator, a `memoryview` over a private buffer. And anything that *does* pickle gets copied — sending a large object to a worker means serializing, writing, reading and deserializing it, which is why moving a fast function to a process pool can easily be slower than leaving it in one thread. The failure surfaces on the `Future`, not at `submit`: the work item is pickled by the executor's management thread, so `submit` returns normally and the pickling error appears when you call `result()`. ### The `__main__` guard The start method decides how the child is created. `fork` clones the parent's memory. `spawn` starts a fresh interpreter and re-imports the `__main__` module to rebuild the target's namespace. `forkserver` forks from a small pre-imported server process. In Python 3.14 the default on Unix other than macOS became `forkserver`; macOS and Windows use `spawn`, and `fork` must now be requested explicitly. Under both `spawn` and `forkserver`, module-level code in `__main__` runs again in the child — so a script that creates the executor at module level recursively spawns processes. The fix is the familiar guard: ```python if __name__ == "__main__": with ProcessPoolExecutor() as pool: ... ``` Under the old `fork` default this bug was invisible on Linux and appeared only on macOS or Windows; since 3.14 it is visible almost everywhere, which is a good thing. ### Set-up state that cannot be sent When a worker needs expensive state — a loaded model file, a compiled parser, a connection — you do not pickle it per call. Both executors accept `initializer` and `initargs`, a callable run once per worker at start-up, which is the idiomatic place to build unpicklable state into a module-level global the tasks then read. `ProcessPoolExecutor` also accepts `max_tasks_per_child` to recycle a worker after a fixed number of items — it requires a non-`fork` start method — which caps damage from a leaking native library. ### When the child dies A worker process can vanish without raising: a segfault in a native extension, the OOM killer, an explicit `os._exit`. There is no exception to pickle back. `concurrent.futures` handles this by declaring the pool broken: every pending and future `Future` completes with a `BrokenProcessPool` error from `concurrent.futures.process`, and the executor is unusable from then on. That is deliberately unrecoverable — it means state you cannot see was lost — and the right response is to build a new executor rather than to retry on the old one. ### The interview answer "Threads share memory, processes do not, so the process pool pickles the callable, the arguments and the result. Pickle stores functions by qualified name, so the target must be a module-level function; the arguments must be picklable and are copied; and I need the `__main__` guard because the child re-imports my module."
- Where does the pickling failure actually appear, given that submit() returned without complaining?On the `Future`. `submit` only enqueues the work item; the executor's management thread pickles it and writes it to the call queue, and when that fails it stores the error on the corresponding `Future`. You see it when you call `result()`. This is the same store-then-re-raise rule that applies to any failure inside a submitted job.
- How do you give each worker process an expensive object that cannot be pickled?Build it inside the worker with the `initializer` and `initargs` arguments the executor accepts. The initializer runs once per worker process at start-up and typically assigns the object to a module-level global that the submitted functions read. That keeps the unpicklable object entirely inside the child and off the wire, and it amortizes the set-up cost across every task that worker runs.
- What state does a forkserver or spawn child inherit from the parent that a fork child would have had?Far less. `fork` clones the parent's whole memory image, so module globals already mutated at runtime carry over. `spawn` and `forkserver` start from a fresh import of your modules, so only what import-time code creates exists; anything set later must be passed explicitly or rebuilt in an initializer. Since 3.14 the Unix default is `forkserver` outside macOS, so code relying on inherited state now breaks there.
saying these in an interview costs you the question
- Thinks ProcessPoolExecutor accepts any callable a thread pool does
- Believes pickle serializes the function's bytecode
- Says the __main__ guard is only a Windows requirement
- Ignores the copy cost of large arguments and return values
- Expects a segfaulted worker to raise a normal exception
- Assumes the pickling error is raised by submit() itself