skip to content

Why can a custom exception raised in a multiprocessing worker fail to reach the parent?

level: seniorimportance: nice to knowfreq 18%

answer

  1. Errors have to come back somehow
  2. The exception object is pickled like anything else
  3. Rebuilding calls the class again
  4. The args tuple is what gets passed
  5. Tracebacks hold live frames and cannot travel

basics

~20 s

A worker's exception is pickled, and rebuilding it calls the class with its args tuple. If a custom init takes parameters it never forwards to the base class, reconstruction raises TypeError and the real failure is lost.

solid answer

~40 s

Exceptions cross the boundary the same way arguments do: the worker pickles the exception object and the parent unpickles it. `BaseException` reduces to its class plus its `args` tuple, with the instance `__dict__` restored afterwards as state, so unpickling effectively calls `MyError(*exc.args)`. A custom exception whose `__init__` takes, say, a record id and a reason but passes a single formatted string to `super().__init__` ends up with a one-element `args`, and reconstruction fails with `TypeError: missing 1 required positional argument`. The parent then sees a serialization error instead of the domain error. The fix is to keep the constructor compatible with `args` — forward every parameter to `super().__init__` — or implement `__reduce__` returning the class and the arguments needed to rebuild it. Tracebacks never pickle at all.

code

python · 15 lines
python
import pickle


class UpdateFailed(Exception):
    def __init__(self, record_id, reason):
        super().__init__(f"{record_id}: {reason}")
        self.record_id = record_id


error = UpdateFailed(42, "timeout")
print(error.args)
try:
    pickle.loads(pickle.dumps(error))
except TypeError as exc:
    print("round-trip failed:", exc)

go deeper

for a junior

Know that an error raised in a worker process has to be sent back to the parent, and that it arrives as a rebuilt object rather than the original. Seeing the parent report a different error than the worker hit should make you suspicious.

for a middle

Explain the mechanics: the class plus the args tuple is what gets stored, reconstruction calls the class with those args, and the instance dictionary is restored afterwards as state. Show the constructor mismatch that breaks it.

for a senior

Diagnose it in the wild: a TypeError surfacing from result handling while workers log a domain failure. Fix it with a compatible constructor or reduce, add a round-trip test, and know the traceback is lost unless the executor formats it as text.

for a principal

Treat exception classes that cross process boundaries as part of the serialization contract. Set the standard — plain data attributes, constructor parameters that map onto args, no live resources — so that failures from workers stay reportable and observable.

**Exceptions are just objects, and they travel the same road.** When a worker process raises, the pool catches the exception, pickles it, and sends it back through the result pipe so the parent can re-raise it. Nothing about being an exception exempts it from pickling rules, which means an exception can fail to travel for exactly the same reasons an argument can — and the failure is more confusing, because it replaces a meaningful domain error with a serialization error. **The default reduction.** `BaseException` defines a reduction that yields the exception's class and its `args` tuple, plus the instance `__dict__` as state when it is non-empty. Unpickling therefore *calls the class* with those args — not the allocate-and-restore path a plain object gets — and only afterwards updates the new instance's `__dict__`. Two facts fall out of that. First, attributes you set in `__init__` do survive, because they live in `__dict__` and are restored as state. Second, and fatally, the constructor must accept `args` as its positional arguments. **The common break.** The idiomatic-looking custom exception is the one that fails: ```python class UpdateFailed(Exception): def __init__(self, record_id, reason): super().__init__(f"{record_id}: {reason}") self.record_id = record_id ``` Because `super().__init__` was handed one formatted string, `args` has length one, while `__init__` demands two positional parameters. Round-tripping raises `TypeError: UpdateFailed.__init__() missing 1 required positional argument: 'reason'`. In a pool, the parent sees that `TypeError` — raised while unpickling the result — rather than the update failure the worker actually hit. Engineers chase a phantom bug in the parent for hours before noticing the exception class is the culprit. **Two fixes.** The simplest is to keep the constructor compatible with `args`: forward every parameter to `super().__init__(record_id, reason)` and derive the message from them in `__str__`, so `MyError(*exc.args)` reconstructs cleanly. The more explicit fix is `__reduce__`, returning `(UpdateFailed, (self.record_id, self.reason))`, which states the rebuild recipe directly and leaves the constructor free to look however you like. Either way the guard is a one-line unit test: `pickle.loads(pickle.dumps(exc))` for each exception type your workers can raise. That test is worth writing precisely because the failure only shows up across a process boundary. **Other ways an exception refuses.** Even with a compatible constructor, an attribute stored on the exception can sink it — an open file object, a connection, a socket, a lock, a lambda captured for a retry callback. The `__dict__` is pickled as state, so anything you attached to the exception for debugging must itself be picklable. Attaching the offending payload to the error, which feels helpful locally, is the exact move that makes it unreportable from a worker. **The traceback never travels.** A traceback object holds live frames and cannot be pickled at all. Plain `multiprocessing` re-raises the exception in the parent with the parent's own stack, so the worker's frames are simply gone. `concurrent.futures.ProcessPoolExecutor` mitigates this by formatting the worker traceback into a string and attaching it to what it re-raises, so the parent's traceback shows the remote one as text — useful, and worth knowing as a reason to prefer that executor when diagnosability matters. If you need the frames themselves, format and log them inside the worker at the point of failure, where they still exist. **The class has to be importable too.** An exception is rebuilt by looking its class up in the receiving process, exactly the way a function reference is resolved. An exception class defined inside a function, or generated dynamically at run time, cannot be found by that lookup and fails to unpickle even with a perfectly compatible constructor. The same is true of a class whose module the parent cannot import — a worker-only helper module, for instance. Keeping worker exceptions in a shared, importable module is therefore not just tidiness; it is what makes them reportable at all. **Design consequence.** Exception classes that cross a process boundary are part of your serialization contract, not just your error vocabulary. Keep them plain: constructor parameters that map onto `args`, attributes that are simple data, no live resources, no callbacks. That is enough to make worker failures arrive intact, which is the difference between an actionable error and a mysterious `TypeError` from deep inside the result-handling machinery.

  • What happens to the worker's traceback when the exception reaches the parent?
    It does not travel — a traceback holds live frame objects and cannot be pickled. Plain `multiprocessing` re-raises the exception with the parent's own stack, so the worker frames are gone. `concurrent.futures.ProcessPoolExecutor` formats the remote traceback into a string and attaches it to the re-raised exception, which is a real diagnosability advantage. If you need the frames, format and log them inside the worker.
  • Do attributes set on the exception survive the round trip?
    Yes, provided reconstruction succeeds: the default reduction includes the instance `__dict__` as state, so attributes are restored after the class is called with `args`. But every attribute value must itself be picklable, so attaching the failing payload, an open file object, a connection or a retry callback to the exception is the quickest way to make the error unreportable from a worker.
  • How would you stop this class of bug from reaching production?
    Add a unit test that round-trips every exception type a worker can raise: `pickle.loads(pickle.dumps(exc))`, then assert the rebuilt instance carries the fields you rely on. It runs in milliseconds, needs no pool, and catches the constructor-versus-args mismatch at the moment the exception class is written rather than during an incident.

saying these in an interview costs you the question

  • Assumes exceptions cross process boundaries for free
  • Thinks the worker's traceback object arrives in the parent
  • Stores an open file handle on the exception instance
  • Never round-trip tests a custom exception class
  • Blames the pool when unpickling raises TypeError
  • Believes attributes are lost because __init__ reruns

context