skip to content

Why does sys.setprofile installed in the main thread miss work done in worker threads?

level: seniorimportance: should knowfreq 30%

answer

  1. Where does the hook actually live?
  2. It is per-thread interpreter state
  3. Threads started later versus already running
  4. A sampler walks every thread's frame
  5. Async needs attribution, not coverage

basics

~20 s

The profiling hook set by sys.setprofile is per-thread interpreter state, so it only fires for the thread that installed it. Use threading.setprofile for threads started later, threading.setprofile_all_threads for existing ones, or a profiler that samples every thread's stack.

solid answer

~40 s

`sys.setprofile()` stores the hook in the calling thread's state, so a profiler started in the main thread instruments only the main thread; worker threads run entirely uninstrumented and their work simply never appears. `threading.setprofile()` registers a hook that each *newly started* thread installs during its bootstrap, and `threading.setprofile_all_threads()`, added in Python 3.12, applies to already-running threads as well. The alternative is to change instrument: a sampling profiler periodically walks every thread's current frame, so thread coverage is automatic. `sys.monitoring`, added in 3.12 by PEP 669, is the modern low-overhead hook API and registers per interpreter rather than per thread. Async is a different problem again: all tasks on one event loop share one thread, so a thread-level profile lumps them together and per-task attribution needs instrumentation that understands coroutine frames.

code

python · 19 lines
python
import threading
from collections import Counter

calls = Counter()

def profiler(frame, event, arg):
    if event == "call":
        calls[frame.f_code.co_name] += 1

threading.setprofile(profiler)   # applies to threads started after this

def work():
    return sorted(str(i) for i in range(500))

t = threading.Thread(target=work)
t.start()
t.join()
threading.setprofile(None)
print(calls["work"], calls["<genexpr>"])

go deeper

for a junior

Remember that profiling hooks belong to the thread that installed them, so starting a profiler in the main thread does not measure worker threads. The threading module provides the hook that new threads pick up.

for a middle

Explain the mechanics: sys.setprofile writes per-thread state, threading.setprofile is a default installed during each new thread's bootstrap, and threading.setprofile_all_threads, added in 3.12, also reaches threads that already exist.

for a senior

Show judgement about instrument choice in a running service — hook cost per event versus a sampler's statistical picture — and be able to say why an async service needs per-task attribution rather than more thread coverage.

for a principal

Own the strategy for observability inside long-lived processes: what instrumentation is always on, why per-interpreter monitoring with self-disabling callbacks changed what is affordable in production, and how much accuracy you trade for a sampling budget.

Two facts explain almost every "my profiler saw nothing" report in a concurrent Python program: **profiling hooks are per-thread**, and **asyncio tasks are not threads**. ## The per-thread hook `sys.setprofile(callback)` and `sys.settrace(callback)` write into the *current thread's* interpreter state. The interpreter consults that thread's own hook when it needs to fire a profile or trace event, so a hook installed on the main thread has no effect anywhere else. This is deliberate — a trace function is how debuggers and coverage tools work, and making it global would surprise every thread in the process — but it means a profiler wrapped around `main()` measures the main thread's calls and nothing more. If the real work runs in a pool of workers, the report can be nearly empty while the job is saturating cores. `threading` provides the fix for the common case. `threading.setprofile(fn)` records a process-wide default that the threading bootstrap installs in **each thread started after the call**; the same exists for tracing with `threading.settrace(fn)`. Because it only touches threads created afterwards, threads that already exist stay uninstrumented — which is why Python 3.12 added `threading.setprofile_all_threads()` and `threading.settrace_all_threads()`, which set the default *and* apply it to running threads. ```python import threading from collections import Counter calls = Counter() def profiler(frame, event, arg): if event == "call": calls[frame.f_code.co_name] += 1 threading.setprofile(profiler) # applies to threads started after this t = threading.Thread(target=lambda: sorted(str(i) for i in range(500))) t.start(); t.join() threading.setprofile(None) ``` Note also that a callback installed this way runs *inside* every worker, so it must be cheap and thread-safe: it is called on every call and return in every instrumented thread, and a naive dict-of-lists accumulator is both a bottleneck and a source of contention. ## The modern hook layer `sys.monitoring`, added in Python 3.12 by PEP 669, is the interpreter's low-overhead monitoring API and the one new tools build on. Its shape differs from `sys.setprofile` in ways that matter here: * A tool claims an id with `sys.monitoring.use_tool_id()`, registers per-event callbacks with `sys.monitoring.register_callback()`, and enables the events it wants with `sys.monitoring.set_events()` — for example `sys.monitoring.events.PY_START`. * Events can be enabled **per code object** with `sys.monitoring.set_local_events()`, and a callback can return `sys.monitoring.DISABLE` to permanently switch itself off for that location, so instrumentation you no longer need costs nothing. * Registration is per interpreter, not per thread, so it does not have `sys.setprofile`'s coverage gap. The design goal is that unused monitoring is close to free, which is what makes always-available instrumentation plausible rather than a special profiling run. ## Sampling sidesteps the problem Changing instrument avoids the coverage question entirely. A sampling profiler wakes on a timer and records the current stack of *every* thread — the in-process mechanism is a periodic walk of `sys._current_frames()`, which returns a mapping of thread id to that thread's topmost frame. Coverage is automatic, cost is set by the sample rate rather than by call volume, and a thread parked in a blocking call keeps accumulating samples so the wait is visible. What you give up is exactness: call counts are statistical, and short-lived work between samples can be missed entirely. ## Async is a separate blind spot An event loop runs every task on one thread, so: * **Thread coverage is not the problem** — a main-thread profiler *does* see coroutine code, because it all executes in that thread. * **Attribution is the problem.** Every task's frames land in one thread-level report. A coroutine's cost is smeared across the loop's callbacks, and time an `await` spends suspended belongs to no frame at all: the coroutine is not on the stack while it waits, so a stack sample taken during the await shows the event loop, not the coroutine that is waiting. * Per-task numbers therefore require task-aware instrumentation — timing the synchronous stretches between awaits, or tooling that follows coroutine frames and task identity rather than thread identity. * Anything the loop's thread runs synchronously — a blocking call slipped into a coroutine — *is* visible in a thread profile, and shows up as a long stretch of samples in the loop's thread, which is usually how it is caught. ## What to say in an interview The compact version: hooks are per-thread state, so `sys.setprofile` covers one thread; `threading.setprofile` covers threads started later, `threading.setprofile_all_threads` since 3.12 covers existing ones; `sys.monitoring` since 3.12 registers per interpreter; and a sampler that walks all thread stacks bypasses the question at the cost of statistical rather than exact numbers. For async, the issue is not coverage but per-task attribution, because tasks share a thread and suspended awaits are not on any stack.

  • What does sys.monitoring give an instrumentation tool that sys.setprofile does not?
    Granularity and a way to switch itself off. A tool enables only the event kinds it needs, can enable them per code object with `sys.monitoring.set_local_events()`, and can return `sys.monitoring.DISABLE` from a callback so that location stops firing altogether. Registration is per interpreter rather than per thread, and multiple tool ids can coexist. The result is that instrumentation you are not using costs close to nothing, whereas a profile hook fires on every call in its thread.
  • Why can't a per-thread profile tell you which asyncio task consumed the time?
    Because every task on a loop runs in the same thread, so all their frames are attributed to that one thread, and a suspended `await` is not on the stack at all — a sample taken while a task waits shows the event loop. You get a truthful picture of what the loop's thread executed, but per-task cost and per-task waiting need instrumentation that keys on task identity rather than thread identity.
  • What must be true of a callback passed to threading.setprofile?
    It has to be cheap and safe to run concurrently, because it is installed into every thread started afterwards and fires on every call and return in each of them. Shared accumulators must tolerate concurrent updates, allocation per event should be avoided, and anything that itself calls Python functions risks recursion and enormous overhead. This cost is exactly why sampling or `sys.monitoring` is preferred for long-running processes.

saying these in an interview costs you the question

  • Believing sys.setprofile installs a process-wide hook
  • Expecting threading.setprofile to cover already-running threads
  • Assuming asyncio tasks need per-thread hooks to be seen
  • Treating a suspended await as time on the stack
  • Ignoring the per-event cost of a hook in every thread
  • Thinking a sampler gives exact call counts

context