How do you walk a Python exception's __cause__ and __context__ to the root failure when only the wrapper is reported?
answer
- It is a linked list with two pointers
- One pointer wins over the other
- A flag can end the walk early
- Guard against a hand-made loop
- Emit root type as its own field
basics
~10 sFollow cause if it is set, otherwise context unless suppress_context is True, and repeat until the link is None. Carry a set of seen object ids so a hand-assigned cause cannot loop forever.
solid answer
~50 sReproduce the renderer's own rule in code: from the reported exception, step to `__cause__` when it is not `None`; otherwise step to `__context__` when `BaseException.__suppress_context__` is `False`; otherwise stop. The last exception you reach is the root failure, and every exception you passed through is a layer that chose to wrap. Guard the loop with a set of `id()` values, because `__cause__` can be assigned by hand and is not protected against cycles the way implicit context is. In practice the fix is usually upstream of the walk: the wrapper is all you see because the error path logged `str(exc)` or a formatted message instead of handing the exception object to a reporter that renders chains. Add the walk to the reporter, emit the root type and message as their own fields, and the chain becomes searchable rather than something a human reads out of a text blob.
code
python · 19 linesdef chain(exc):
seen = set()
while exc is not None and id(exc) not in seen:
seen.add(id(exc))
yield exc
if exc.__cause__ is not None:
exc = exc.__cause__
elif not exc.__suppress_context__:
exc = exc.__context__
else:
exc = None
try:
try:
raise ImportError("partially initialized module")
except ImportError as err:
raise RuntimeError("loader registry failed") from err
except RuntimeError as top:
print([type(e).__name__ for e in chain(top)])go deeper
Know that the earlier failures are attributes on the exception object, not just text in a traceback, so code can reach them once you have caught the exception.
Reproduce the precedence rule accurately: cause first, then context unless suppress_context is True, stopping at None — and know why a hand-assigned cause can create a loop.
Show where the chain was really lost: an error path that logged a formatted string instead of the exception object, and a reporter that never walked the chain or emitted the root as a field.
Own the error-reporting contract across services — one walker, one documented stance on suppressed contexts, and root type and message as structured fields so incidents can be aggregated rather than read.
### The rule you are reimplementing A chained exception is a linked list with two possible next-pointers, and the default traceback picks between them with a fixed precedence: `__cause__` if it is not `None`; otherwise `__context__` if `__suppress_context__` is `False`; otherwise nothing. Any code that wants the root cause — an error reporter, a log enricher, a test asserting on a wrap — follows the same rule. ```python def chain(exc): seen = set() while exc is not None and id(exc) not in seen: seen.add(id(exc)) yield exc if exc.__cause__ is not None: exc = exc.__cause__ elif not exc.__suppress_context__: exc = exc.__context__ else: exc = None try: try: raise ImportError("partially initialized module") except ImportError as err: raise RuntimeError("loader registry failed") from err except RuntimeError as top: print([type(e).__name__ for e in chain(top)]) ``` The `seen` set is not decoration. CPython refuses to build a `__context__` loop when it sets the attribute implicitly, but `__cause__` is assignable, and two exceptions can be made to point at each other in a few lines of ordinary code. A reporter that hangs while formatting an error is a bad failure to debug. A deliberate variant is to ignore `__suppress_context__` when walking. The display rule exists to keep tracebacks readable for humans; a machine reporter can reasonably follow `__context__` even when it was suppressed, because the link is still on the object and the root cause is what it is after. Decide that once, document it, and keep it consistent — otherwise two dashboards disagree about the same incident. ### The scenario this actually shows up in A clinical-lab result loader imports analyser plugins from a registry at startup. One release introduces a circular import between the registry and a plugin, so importing it raises `ImportError`. The registry catches `ImportError`, raises its own configuration error, and the startup path logs a formatted message. What the on-call engineer sees is one line naming the plugin as unavailable, with no file, no cycle and no import chain. With a 4-person team and no one who has seen this failure before, that line burns an afternoon. Two things are wrong, and only one of them is the walk. The visible symptom is that the reporter lost the chain; the deeper cause is that the error path threw the exception object away and kept a string. Fixing the reporter to accept the exception, walk the chain and emit the root type and message as separate fields makes the same incident a two-minute read — and makes it searchable, because "root_type: ImportError" is a field you can aggregate over, while a formatted message is not. ### What to emit A chain walk is most useful when it produces structure rather than more prose: * the reported exception's type and message — what the layer called it; * the root exception's type and message — what actually happened; * the depth of the chain, and whether each hop was a cause or a context. The cause/context distinction is diagnostic in itself. A hop through `__cause__` was a deliberate translation by a developer; a hop through `__context__` was the interpreter noticing a second failure inside a handler. A chain that is all causes is a well-layered system reporting itself; a chain with unexpected context hops usually means a handler broke while handling. ### Retention, and where the walk should live Exceptions keep their `__traceback__`, tracebacks keep frames, frames keep locals — and a chain multiplies that by its length. Walking a chain is cheap, but *storing* the exceptions you walked is not: a cache of "recent errors" holding exception objects can pin a large amount of memory, including whatever those frames had loaded. Extract the fields you need, then drop the reference. The same reasoning explains why CPython deletes the `as` name at the end of an `except` block. Put the walk in one place — the error reporter or logging integration — not at every call site. Scattered ad-hoc `while exc.__cause__` loops drift apart, disagree about `__suppress_context__`, and are exactly the sort of code that acquires an infinite loop under a hand-assigned cause. ### Version notes The attributes and the precedence rule are stable across the whole Python 3 line and unchanged in 3.14, so a walker written against them needs no version guards.
- Why does the walk need a set of seen ids at all?Because `__cause__` is an assignable attribute with no cycle protection. Two exceptions can be made to reference each other, directly or through a longer ring, and a naive `while` loop then never terminates — inside an error reporter, which is the worst place to hang. The interpreter only guards the implicit `__context__` link it sets itself.
- Should a machine walker honour `__suppress_context__`?It is a deliberate choice. Honouring it reproduces exactly what a human sees in the traceback; ignoring it finds root causes that a `from None` deliberately hid, which is often what an incident investigation wants. Pick one, document it, and apply it everywhere — the real hazard is two tools disagreeing about the same error.
- What does the mix of cause and context hops tell you about a chain?A hop through `__cause__` was a developer deciding to translate one failure into another, so an all-cause chain is a layered system reporting itself as designed. A `__context__` hop was the interpreter recording a second exception raised while the first was being handled, which often means a handler itself failed — worth flagging separately in a report.
- Why not cache the exception objects you walked for later inspection?Each exception holds its `__traceback__`, which holds frames, which hold their locals — and a chain multiplies that. A cache of recent exception objects can pin large amounts of memory that would otherwise be freed. Extract the fields you need, such as root type, message and hop kinds, then drop the references.
saying these in an interview costs you the question
- Follows only __cause__ and misses implicit context links
- Writes an unbounded loop with no cycle guard
- Believes a suppressed context is gone from the object
- Logs str(exc) and expects the chain to survive
- Stores exception objects long-term in an error cache