Why can a stack dumper built on sys._current_frames() miss the very hang it was written to diagnose?
answer
- It is ordinary Python code
- Your dumper still has to be scheduled
- One thread can hold the interpreter lock throughout
- The admin endpoint hangs with everything else
- Frames returned are live and pin their locals
basics
~20 sA dumper built on sys._current_frames() is ordinary Python: its thread must be scheduled and must run bytecode. If another thread is wedged inside a long native call holding the interpreter lock, the dumper never runs at all.
solid answer
~50 s`sys._current_frames()` returns a dict mapping each live thread's `ident` to the frame object it is currently executing, which you format with `traceback.print_stack` or `traceback.format_stack` and join to names via `threading.enumerate()`. It is attractive because it is pure Python: no signals, works on Windows, output goes through your own logging with your own formatting. The catch is that it can only be reached *by running Python*. A dumper thread — or an admin endpoint — has to be scheduled and hold the global interpreter lock, so if a thread is stuck inside a long call in a compiled extension that never releases the lock, your dumper never gets to run at all and the endpoint hangs alongside everything else. That is the case a signal-armed `faulthandler` dump handles and this one cannot. There is also a lifetime trap: the frames you get back are live objects that keep their locals alive.
code
python · 18 linesimport sys
import threading
import time
import traceback
def render_page():
time.sleep(5)
threading.Thread(target=render_page, name="render-0", daemon=True).start()
time.sleep(0.2)
names = {t.ident: t.name for t in threading.enumerate()}
frames = sys._current_frames()
for ident, frame in frames.items():
print(f"--- {names.get(ident, ident)} ---")
traceback.print_stack(frame)
del frames, framego deeper
Know that Python can be asked for every thread's current frame from inside the process, that the result is keyed by thread identifier, and that the traceback module turns a frame into readable text.
Explain the mechanics: what the keys and values are, how to join them to thread names, and why the frames must be formatted and released immediately rather than stored.
Show that you know its blind spot — a dumper that needs to execute Python cannot run while a thread holds the interpreter lock inside a long native call — and that the signal-armed dump is the fallback for exactly that case.
Decide the diagnostic strategy: which mechanism each service ships, how dumps reach the logging pipeline with request context, and the trade between an ergonomic in-process dumper and one that survives an interpreter that has stopped running Python.
## What the call gives you `sys._current_frames()` returns a `dict` mapping the `ident` of each thread in the current interpreter to the frame object that thread is executing right now — the innermost frame, from which you can walk outwards. Combined with `threading.enumerate()` for names and `traceback` for formatting, twelve lines of Python give you a labelled dump of every thread: ```python names = {t.ident: t.name for t in threading.enumerate()} for ident, frame in sys._current_frames().items(): print(f"--- {names.get(ident, ident)} ---") traceback.print_stack(frame) ``` The leading underscore is honest signposting: this is a CPython implementation detail. It is documented, it has been there for many releases, and plenty of production code uses it — but it carries no cross-implementation guarantee, and you should treat it as a diagnostic rather than as something a feature depends on. ## Why it is genuinely attractive Against a signal-based dump, this approach wins on several axes at once. It works on Windows, where signal registration is unavailable. Output goes wherever your logging goes — as structured records, one event per thread, with your request ids attached — instead of raw text on a file descriptor captured at startup. You can filter: dump only the worker pool, only threads whose innermost frame is in your own package, only when a queue has been non-empty for a minute. And you can trigger it from inside the process, on a condition the process itself notices, rather than needing a human with a shell. ## Why it can be blind exactly when you need it All of that depends on your dumper being **able to run Python**, and that is the assumption a hang breaks. CPython executes bytecode under a global interpreter lock. A thread holds it while running Python and releases it at intervals and around most blocking calls. But a thread that enters a long-running function in a compiled extension which does *not* release the lock holds it for the entire duration. During that window no other thread executes a single bytecode. Your dumper thread is not blocked on anything of yours; it is simply never scheduled to run Python. If you exposed the dump through an administrative endpoint, that endpoint stops responding too — and the person debugging concludes the process is completely dead, when in fact one thread is very much alive and holding everything else hostage. The same reasoning covers a wedged main thread in a single-threaded program: there is nobody left to call the function. A signal-armed `faulthandler` dump does not have this problem, because its handler is C code that runs in the signal handler and walks thread state directly without needing the interpreter lock or the eval loop. The honest summary is: build the Python-level dumper for its ergonomics, and arm the signal-level one for the case where the Python-level one cannot run. ## The frame-lifetime trap The values in the returned dict are **live frame objects**, and a frame references its locals. If you stash frames in a module-level variable to look at later, you pin every local variable of every thread — including, in a renderer, a page buffer that accounts for most of a 2.4 GB working set. Frames also participate readily in reference cycles, so the memory is not necessarily reclaimed the moment you think it is. The discipline is to format immediately into strings and drop the references: ```python frames = sys._current_frames() text = {i: "".join(traceback.format_stack(f)) for i, f in frames.items()} del frames ``` Holding a frame is also a correctness hazard in a sampling loop: repeated samples that each retain frames turn a diagnostic into a leak. ## Scope The dict covers threads of the **current interpreter** only. It is not a window into another process, and a sub-interpreter's threads are not yours to see. And like any Python-level snapshot it is racy: threads move on between the call returning and your formatting the frames, which is fine for diagnosis and wrong for anything that needs a consistent picture. A related call, `sys._current_exceptions()`, gives the exception each thread is currently handling — since Python 3.12 it returns exception instances rather than the older three-tuples. It answers a different question, but it is the natural companion when you are dumping a pool that seems to be stuck inside error handling.
- How do you turn what sys._current_frames() returns into readable, labelled tracebacks?The values are frame objects, so `traceback.print_stack(frame)` or `traceback.format_stack(frame)` renders each one outermost-to-innermost. The keys are thread identifiers, so build `{t.ident: t.name for t in threading.enumerate()}` first and label each block with the worker's name. Format straight into strings and drop the frame references rather than carrying frames around.
- What is the memory hazard of keeping the frames it returns?A frame holds its locals alive. Stashing frames for later inspection pins every local of every thread — in a renderer, that can be a multi-gigabyte page buffer that should have been freed — and frames enter reference cycles easily, so collection is not immediate. Format to text immediately and `del` the dict; a sampling loop that retains frames turns a diagnostic into a leak.
- When would you deliberately prefer this over a signal-armed dump?On Windows, where signal registration is not available at all; when the output must go through structured logging with request context attached rather than raw text to a descriptor; when you want to filter to a subset of threads; and when the trigger is a condition the process notices about itself. It is the better ergonomic choice whenever the interpreter is healthy enough to run it.
saying these in an interview costs you the question
- Assumes it works when the interpreter cannot run Python
- Thinks it returns Thread objects rather than frame objects
- Stashes the returned frames and pins their locals
- Believes it can see another process or interpreter's threads
- Treats the underscore-prefixed name as a stable public API
- Claims it can dump a thread stuck inside a C call in detail