skip to content

Why does a tracer installed with sys.settrace see nothing from a webhook receiver's worker threads?

level: seniorimportance: should knowfreq 24%

answer

  1. Ask which thread armed the hook
  2. Interpreter state that is not shared
  3. Threads started later start empty
  4. The threading module has its own setter
  5. A raising callback disarms itself silently

basics

~20 s

The trace function is per-thread state. sys.settrace instruments only the thread that calls it, so worker threads started later begin untraced. Use threading.settrace for threads started afterwards, or call sys.settrace at the top of each worker.

solid answer

~50 s

`sys.settrace` sets a hook on **the calling thread only**; there is no process-wide trace state. A webhook receiver that installs a tracer at startup on the main thread and then dispatches each request to a pool worker will record the main thread's accept loop and nothing else. Two fixes: call `threading.settrace(func)` before the workers are created, which makes every thread started through the `threading` module install that hook as it starts - and that covers a thread pool executor's workers, since they are ordinary threads - or call `sys.settrace` as the first statement inside the worker function. Since 3.12 `threading.settrace_all_threads` additionally reaches threads that are already running. Check the second failure mode too: if the callback raised inside a worker and the handler swallowed the exception, the interpreter unset tracing for that thread and `sys.gettrace()` there now returns `None`.

code

python · 24 lines
python
import sys
import threading

traced = set()

def tracer(frame, event, arg):
    traced.add(frame.f_code.co_name)
    return tracer

def worker():
    x = 1

sys.settrace(tracer)                 # this thread only
t = threading.Thread(target=worker)
t.start(); t.join()
sys.settrace(None)
print("worker seen:", "worker" in traced)   # False

traced.clear()
threading.settrace(tracer)           # every thread started from now on
t = threading.Thread(target=worker)
t.start(); t.join()
threading.settrace(None)
print("worker seen:", "worker" in traced)   # True

go deeper

for a junior

Take away the core fact: the trace hook belongs to one thread, so instrumentation set up on the main thread does not follow work handed to another thread. That alone explains most empty tracing output.

for a middle

Explain the mechanism and the fix: per-thread hook state, threading.settrace for threads started later, sys.settrace inside the worker as the explicit alternative, and sys.gettrace() as the way to check what is armed where.

for a senior

Diagnose like an operator: check thread scope, check whether another tool holds the single slot, and check whether the callback raised and silently disarmed itself behind a broad exception handler. Then justify the overhead before arming anything on a request path.

for a principal

Own the standing rule for live instrumentation: what may be armed in production, on how many workers, for how long, who arbitrates the single per-thread hook, and whether deep tracing belongs in the service at all rather than in a reproduction environment.

## The state is per-thread, and nothing warns you The trace hook lives in thread state, not in the interpreter's global state. `sys.settrace(func)` arms the thread that executes it. Every other thread - already running or created a microsecond later - has its own empty slot. There is no exception, no warning, and no log line; the tracer just reports a suspiciously small part of the program. A webhook receiver makes this vivid. The process installs instrumentation during startup, on the main thread, then hands each inbound delivery to a worker from a pool. All the interesting code - signature verification, payload parsing, the handler that occasionally swallows an exception - runs on threads the tracer never touched. What you get back is a trace of the accept loop. ## The fixes, in order of preference **1. `threading.settrace(func)` before the workers exist.** The `threading` module keeps a hook that every thread it starts installs for itself, via `sys.settrace`, before running its target. Set it before the pool is created and every worker arrives instrumented. This covers a thread pool executor's workers too, because those are ordinary `threading.Thread` objects. The matching `threading.setprofile` does the same for the profiling hook. **2. `threading.settrace_all_threads(func)`, added in 3.12.** Same idea, but it also installs into Python threads that are already running - the right tool when the pool was created before you decided to instrument. `threading.setprofile_all_threads` is its profiling counterpart. **3. Install inside the worker.** Calling `sys.settrace(func)` as the first statement of the worker function always works and is explicit. It is the only option for threads that were not created through the `threading` module at all - a thread spawned by a C extension, or by the low-level thread module, never consults the `threading` hook. ```python def worker(job): sys.settrace(tracer) try: handle(job) finally: sys.settrace(None) ``` ## The second cause: the tracer uninstalled itself If the callback raises, the exception propagates at the traced location **and** the interpreter unsets the hook for that thread, exactly as if `sys.settrace(None)` had been called. Now combine that with a handler that catches broadly - the classic `except Exception: log_and_continue()` around a webhook payload. The tracer's error is swallowed as if it were a bad payload, tracing stops for that worker, and the worker keeps serving traffic. The symptom is identical to "tracing was never installed", which is why the diagnostic is to call `sys.gettrace()` *inside a worker* and log what comes back: `None` means either never armed or armed and then killed, and only a guarded callback tells the two apart. Wrap the callback body in `try`/`except`, report to a side channel, and always return the callback itself. ## Deciding whether to do this at all Even when it is wired correctly, per-line tracing on request-handling threads is a production hazard. Every executed line becomes a Python-level call, and instrumented code loses the specialized fast paths the adaptive interpreter would otherwise take. Against a 92nd-percentile latency budget of a few tens of milliseconds, a per-line hook on the handler path can consume the budget by itself. Sensible production use looks like: instrument a narrow slice of frames by filtering at the `'call'` event; arm one worker rather than the pool; arm it for a bounded window and disarm in a `finally`; or move measurement to a coarser hook entirely. ## The other thing to check first Before blaming thread scope, confirm nobody else took the slot. There is exactly one trace function per thread, and the last caller wins silently. A debugger attached to the process, or a coverage tool started by the test harness, will have displaced yours - or yours theirs, which is the more embarrassing outcome in a live process. `sys.gettrace()` at the moment of installation tells you what you are about to overwrite. ## What good looks like in the answer A senior answer names three things without prompting: the hook is per-thread state; `threading.settrace` (and, since 3.12, `settrace_all_threads`) is how you extend it to workers; and a callback that raises silently disarms itself. The candidate who only says "you have to set it in each thread" has the mechanism but not the failure modes. ## A checklist for the empty-output case When instrumentation reports less than you expected, work through it in this order rather than rewriting the tracer. First, *which thread armed it* - log the thread name at install time and again from inside a worker. Second, *is it still armed* - `sys.gettrace()` from the worker, since a callback that raised has already removed itself. Third, *did someone else take the slot* - a debugger or a coverage run installs into the same single per-thread hook and the last writer wins silently. Fourth, *is the global handler filtering too aggressively* - returning `None` at the `'call'` event for a frame you actually wanted produces exactly the same symptom as never being armed. Only after those four does it make sense to suspect the events themselves. ## Cleaning up Whatever the mechanism, disarm deliberately. `threading.settrace(None)` stops future threads picking the hook up, and `sys.settrace(None)` clears the current thread, but neither reaches back into workers that are already running and armed; those keep tracing until code on those threads clears their own hook. In a long-lived receiver that means a debugging session left half-disarmed can go on costing latency for hours, which is the argument for a bounded window and a `finally` around every arming.

  • Does threading.settrace affect threads that are already running?
    No. `threading.settrace` records a hook that each thread started from the `threading` module afterwards installs for itself; threads already running keep whatever they had, which is usually nothing. Since Python 3.12, `threading.settrace_all_threads` installs into currently running Python threads as well, which is what you want when the pool was created before you decided to instrument it.
  • How would you confirm from inside a worker that tracing is actually active?
    Call `sys.gettrace()` on that thread and log the result. `None` means the thread was never armed, or was armed and then disarmed because the callback raised - the interpreter unsets the hook on an exception in the callback. A guarded callback that logs its own errors to a side channel is what lets you distinguish the two cases after the fact.
  • Which threads will threading.settrace never reach?
    Any thread not created through the `threading` module: threads started by the low-level thread API, and threads created inside a C extension that later call into Python. Those never consult the module's hook, so the only way to instrument them is for code running on that thread to call `sys.settrace` itself.

saying these in an interview costs you the question

  • Believes sys.settrace applies process-wide once installed
  • Expects an error when tracing a thread that is not armed
  • Thinks threading.settrace retro-fits already-running threads
  • Never considers that the callback raised and disarmed itself
  • Leaves per-line tracing armed on request-handling threads
  • Forgets that another tool may hold the single per-thread slot

context