Why is the traceback missing when a child process's exception is re-raised in the parent?
answer
- Tracebacks hold live frames
- Frames cannot leave their interpreter
- Cause and context are dropped too
- Format where the frames still exist
- Notes ride along in the instance dict
basics
~20 sTraceback objects reference live frames, code objects and locals, so they cannot be pickled. What crosses is the exception's class and args, and the parent's copy comes back with __traceback__, __cause__ and __context__ all set to None.
solid answer
~40 sA traceback object is a linked list of frames holding live code objects and local variables; `pickle.dumps` on one raises `TypeError: cannot pickle 'traceback' object`. `BaseException.__reduce__` does not include it, and it does not include `__cause__` or `__context__` either, so a re-raised copy in the parent has all three as None and shows only the parent's own stack from the point of retrieval. The fix is to format in the child, where the frames still exist: capture `traceback.format_exc()` and carry the text -- attaching it with `add_note` works well, since notes live in the instance dictionary and do pickle. Some machinery already does a version of this, chaining a text-only synthesized cause onto the parent-side exception, which is why such tracebacks sometimes show two stacks.
code
python · 13 linesimport pickle
import traceback
try:
try:
1 / 0
except ZeroDivisionError as inner:
raise RuntimeError("scrape failed") from inner
except RuntimeError as exc:
exc.add_note(traceback.format_exc())
back = pickle.loads(pickle.dumps(exc))
print(back.__traceback__, back.__cause__) # None None
print("".join(traceback.format_exception(back))) # the note still printsgo deeper
Remember that the stack does not travel: a traceback points at live frames, so a copy of the exception in another process has none. Diagnostics must be turned into text first.
Explain why traceback objects are unpicklable and that __cause__ and __context__ are dropped as well, and show the fix: traceback.format_exc() in the child, carried across as a string or an added note.
Design the worker so it produces its own diagnostics -- formatted trace, worker identity, target identity -- and read a doubled parent-side traceback correctly, knowing the upper block is text from a process that no longer exists.
Decide the observability contract for worker failures across the fleet: what a failed unit of work must emit and where, so that debugging never depends on an exception copy that structurally cannot carry a stack.
**What a traceback actually is.** `exc.__traceback__` is not text. It is a chain of traceback objects, each pointing at a frame object, which in turn points at a code object, the local and global namespaces of that frame, and the instruction offset. Those references reach into the running interpreter: compiled bytecode, live objects, possibly open files and sockets sitting in locals. None of that is serializable, and serializing it would be meaningless anyway -- the frames belong to an interpreter that is about to exit. `pickle.dumps(exc.__traceback__)` says so plainly: `TypeError: cannot pickle 'traceback' object`. **So what does cross?** Only what `BaseException.__reduce__` yields: the class reference, `args`, and the instance dictionary as state. Three things people expect to survive do not: * `__traceback__` -- gone, for the reason above. * `__cause__` -- the exception attached by `raise ... from`. It is a C-level slot, not an entry in the instance dictionary, so it is not part of the pickled state. * `__context__` -- the implicit chaining set when one exception is raised while another is being handled. Same reason. All three come back as `None`. That is the real sting: not only is the child's stack missing, the *chain* is missing too, so a carefully wrapped `raise DomainError(...) from driver_error` arrives in the parent as a bare `DomainError` with no trace of what caused it. The stack the parent then prints starts wherever it re-raised the copy -- the line that asked for the result -- which is exactly the least informative place in the program. **Notes are the exception to the rule.** A note attached with `BaseException.add_note` is stored on the instance, so it lands in the instance dictionary and *does* pickle. That makes notes the cheapest reliable channel for diagnostic text across a process boundary. Format the traceback where the frames still exist and attach the string: ```python try: scrape(target) except Exception as exc: exc.add_note(traceback.format_exc()) raise ``` When the parent prints the reconstructed exception, the note prints with it, and you see the child's stack as text. **Do not try to ship a smarter object.** The obvious idea -- serialize `traceback.TracebackException`, which is designed as a lightweight capture -- does not work on 3.14: its frame summaries hold code objects, and pickling one raises `TypeError: cannot pickle code objects`. Format it to strings first, then send the strings. The rule generalises: cross-process diagnostics are text or plain data, never live objects. **Why some parent-side tracebacks look doubled.** Machinery that returns results already applies this idea. It formats the child's traceback into a string, wraps that string in a placeholder exception, and chains it onto the copy it re-raises in the parent. The output then shows the child's stack as text, a chaining line, and then the parent's real stack. Recognising that shape is a practical skill: the upper block is a rendering of something that happened in another process and its frames no longer exist, so a debugger cannot step into them. **Designing for it.** Treat the child as the only place where good diagnostics can be produced. Log the full failure there, with the target identity in the message, before anything crosses. Attach the formatted trace as a note or as a plain-data field on a purpose-built exception. Include the worker's identity, since the parent otherwise cannot tell which of many workers failed. Concretely: a metrics scraper fanning out across a 17-service dependency graph should have every worker log its own failure with the service name attached; a parent that only re-raises copies will tell you that something failed 17 times and nothing about where. **And remember the case with no exception at all.** If the child dies during startup -- an import cycle in a fresh interpreter under `spawn` or `forkserver` -- there is no exception object to format, no note to attach and no copy to re-raise. The child's stderr and its exit status are the only evidence, which is one more reason to capture worker stderr rather than discard it. **One more thing worth carrying.** Since the copy has no stack, it also has nothing that identifies which unit of work produced it. Put that identity in the exception's own data -- the target, the batch, a correlation id -- rather than trusting a log line the parent will have to join up by timestamp. It costs one field and it is the difference between a report you can act on and one that says only that something, somewhere, failed.
- Besides `__traceback__`, what else does a pickled exception lose?`__cause__` and `__context__`. Both are C-level slots on the exception rather than entries in its instance dictionary, and `BaseException.__reduce__` carries only the class, `args` and that dictionary. So a copy re-raised in the parent has all three as `None`, and a chain built with `raise ... from` arrives with no record of what it was raised from.
- Why does a note survive the trip when the traceback does not?A note added with `BaseException.add_note` is stored on the instance, so it lives in the instance dictionary that `__reduce__` carries as state, and it is a list of plain strings -- trivially picklable. A traceback is a chain of frame and code objects tied to a running interpreter and cannot be serialized at all. That asymmetry makes notes the practical channel for cross-process diagnostics.
- Why does a parent-side traceback from a worker sometimes show two separate stacks?Because the machinery formats the child's traceback into text, wraps it in a placeholder exception and chains that onto the copy it re-raises. The upper block is a rendering of frames that no longer exist in any process, and the lower block is the parent's real stack from the retrieval point. A debugger can step into the second but never the first.
saying these in an interview costs you the question
- Expects the child's stack frames to arrive intact
- Thinks __cause__ survives a pickle round trip
- Tries to pickle a traceback or a frame object
- Formats the traceback in the parent, where frames are gone
- Ignores the worker's own logs when diagnosing
- Assumes a re-raised copy can be debugged step by step