skip to content

Instrumentation Hooks

The event streams the interpreter exposes so tools can watch a program run: the legacy trace and profile callbacks, the modern low-overhead event API, statistical sampling, and audit events.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

15

What does returning sys.monitoring.DISABLE from an event callback do?

level: middleimportance: must knowfreq 30%

answer

  1. The sentinel a callback can return
  2. Stops paying for a spot you already saw
  3. Keyed by instruction offset, not by line
  4. Only local events; restart_events undoes it

basics

~20 s

It retires that one code location for that event and that tool: CPython de-instruments the instruction, so the callback never fires from that spot again and later executions run at almost full speed. sys.monitoring.restart_events re-arms every retired location.

solid answer

~40 s

`sys.monitoring.DISABLE` is a sentinel object. Returning it from a callback tells the interpreter never to call that tool back for that event at that exact instruction again, and CPython removes the instrumentation there. Because instrumentation is per code object and per instruction, a line inside a hot loop costs one callback rather than one per iteration — that is what makes line coverage cheap on 3.12 and later. Only the local events can be retired this way: `PY_START`, `PY_RESUME`, `PY_RETURN`, `PY_YIELD`, `CALL`, `LINE`, `INSTRUCTION`, `JUMP`, `BRANCH_LEFT`, `BRANCH_RIGHT` and `STOP_ITERATION`. There is no per-location undo; `sys.monitoring.restart_events()` re-arms everything for every tool at once. Returning `None` leaves the location armed, which is what a counting profiler wants.

code

python · 30 lines
python
import sys

TOOL_ID = sys.monitoring.COVERAGE_ID
hits = []


def on_line(code, line_number):
    hits.append(line_number)
    return sys.monitoring.DISABLE


def loop():
    total = 0
    for i in range(1000):
        total += i
    return total


sys.monitoring.use_tool_id(TOOL_ID, "line-coverage")
sys.monitoring.register_callback(TOOL_ID, sys.monitoring.events.LINE, on_line)
sys.monitoring.set_local_events(TOOL_ID, loop.__code__, sys.monitoring.events.LINE)
loop()
print("callbacks after 1000 iterations:", len(hits))
sys.monitoring.restart_events()
loop()
print("callbacks after restart_events:", len(hits))
sys.monitoring.set_local_events(
    TOOL_ID, loop.__code__, sys.monitoring.events.NO_EVENTS
)
sys.monitoring.free_tool_id(TOOL_ID)

go deeper

for a junior

Recall that a callback can return a special sentinel meaning 'never call me here again', and that this is why modern coverage tooling is cheap. Naming sys.monitoring.DISABLE and its effect is enough at this level.

for a middle

Explain the mechanics: retirement is keyed to one instruction offset in one code object for one tool, only local events qualify, and restart_events is the only way back. Be able to say why a for header may fire twice.

for a senior

Demonstrate the judgement call — reachability tools should retire locations, counting and timeline tools must not — and be able to reason about the residual cost of instrumented-but-disabled code left armed in a long-running service.

for a principal

Own the tradeoff between always-on instrumentation and on-demand diagnostics: what overhead budget the platform accepts, whether a shared restart_events call is safe when several tools coexist, and how sampling policy is agreed across teams.

`sys.monitoring.DISABLE` is a unique sentinel object exposed by the `sys.monitoring` module. Returning it from an event callback is the API's central performance idea, and it is the reason line-level coverage on Python 3.12 and later can run at a fraction of the cost it used to. ### What actually gets retired When a tool subscribes to an event, CPython rewrites the affected code objects so that the relevant bytecode instructions dispatch into the monitoring machinery. The unit of instrumentation is therefore **one instruction offset inside one code object**, not a function and not a source line. When your callback returns `sys.monitoring.DISABLE`, the interpreter de-instruments that exact location for that exact event and that exact tool. The next time control reaches that instruction, there is no dispatch at all: no callback frame, no argument boxing, no return-value check. The location is simply back to being ordinary bytecode for you, while remaining instrumented for any other tool that subscribed to it. That is a fundamentally different economy from "call me every time and I will return early". A flag check inside the callback still pays a full Python call on every execution. `DISABLE` pays for the first execution and nothing thereafter. ```python def on_line(code, line_number): hits.append(line_number) return sys.monitoring.DISABLE ``` A loop body that executes a million times fires that callback once per line, then goes quiet. For a coverage recorder — which only needs to know *whether* a line ran, never how often — this is exactly the right trade. ### Per instruction, not per line A subtle consequence: a single source line can compile to more than one instrumented offset. A `for` header is the common case — the loop setup and the iteration step are different instructions attributed to the same line — so a `LINE` callback that returns `DISABLE` may still see that line twice before it goes quiet. This is not a bug and it is worth being able to explain: the disable state is keyed by instruction offset, and the line number is just what the callback is handed. ### Which events can be retired this way `sys.monitoring` divides its events into **local** and **global-only**. The local events are the ones attached to a specific instruction in a specific code object, and they are exactly the set `sys.monitoring.set_local_events` will accept: `PY_START`, `PY_RESUME`, `PY_RETURN`, `PY_YIELD`, `CALL`, `LINE`, `INSTRUCTION`, `JUMP`, `BRANCH_LEFT`, `BRANCH_RIGHT` (and the older `BRANCH`), and `STOP_ITERATION`. The rest — `RAISE`, `RERAISE`, `PY_UNWIND`, `PY_THROW`, `EXCEPTION_HANDLED`, `C_RETURN`, `C_RAISE` — are global-only. `sys.monitoring.set_local_events` rejects them with `ValueError: invalid local event set`, and per-location retirement is not the tool to reach for with them. If you need to sample exception events, filter in your own callback and accept the dispatch cost. ### Undoing it There is **no per-location undo**. `sys.monitoring.restart_events()` clears the disable state for every location and every tool at once, re-arming everything that is still subscribed. A coverage tool calls it at the start of each measured run; a long-lived diagnostic that wants a fresh sample calls it on an interval. Because it is process-wide across tools, a library should not call it casually — it will re-arm someone else's carefully disabled locations too. Disabling is otherwise sticky for the life of the process. Toggling the event off with `sys.monitoring.set_events(tool_id, sys.monitoring.events.NO_EVENTS)` and back on does not by itself restore a disabled location. ### When not to use it `DISABLE` is wrong whenever you need every occurrence: * A counting profiler that reports call totals. Retire the location and the count stops at one. * A tracer reconstructing an execution order or a timeline. * Anything sampling values rather than reachability — if you disable after the first observation, you never see the value that actually differed. The heuristic is simple: `DISABLE` answers "did this ever run?", never "how often" or "with what". Returning `None` — or anything that is not the sentinel — leaves the location armed, which is the default and the right choice for those cases. ### The residual cost Retiring every location does not return the code object to its pre-instrumentation form; the code stays rewritten and pays a small, constant overhead even with everything disabled. It is close to baseline and far below the un-disabled cost, but "free" overstates it. That distinction matters when someone proposes leaving a fully-disabled coverage tool armed in production.

  • How do you re-arm locations that callbacks retired earlier in the same process?
    `sys.monitoring.restart_events()` clears the disable state for every location and every tool at once. There is no per-location undo, and toggling the event set off and on again does not restore it. A coverage tool calls `restart_events` at the start of each measured run. Because it is process-wide across tools, a library should not call it casually — it re-arms other tools' deliberately retired locations too.
  • Does DISABLE retire the whole code object or a single instruction?
    A single instruction offset. A source line that compiles to more than one instrumented offset — a `for` header is the usual case — can therefore fire the callback twice for the same line number before it goes quiet. The disable state is keyed by offset; the line number is only what the callback is handed.
  • Why is returning DISABLE cheaper than checking a flag inside the callback?
    A flag check still pays a full Python call on every execution: frame setup, argument boxing, return-value handling. DISABLE removes the dispatch itself, so subsequent executions never enter your code at all. The instrumented code object does keep a small residual overhead, so it is close to baseline rather than exactly free.
  • When should a tool not return DISABLE?
    Whenever it needs every occurrence: a profiler counting calls, a tracer reconstructing execution order, or anything sampling values rather than reachability. Retire the location and the count stops at one, and you never see the execution whose value actually differed. DISABLE answers 'did this ever run?', never 'how often' or 'with what'.

A tripwire that cuts itself after the first trigger. You learn that someone walked through the door, you stop paying to watch that door, and only a global reset restrings every wire in the building.

saying these in an interview costs you the question

  • Thinks DISABLE turns the event off for the whole program
  • Returns False or None expecting the location to be retired
  • Believes retirement is per source line rather than per instruction
  • Assumes retired locations re-arm by themselves over time
  • Uses DISABLE in a profiler that must count every call
  • Claims a fully disabled tool costs exactly zero

context

open as a page

How does a sampling profiler estimate where a Python program spends its time?

level: middleimportance: must knowfreq 52%

basics

~20 s

A sampling profiler interrupts the program at a fixed interval, records the current call stack, and counts how often each stack appears. A frame present in 30% of samples held the interpreter for roughly 30% of the profiled period.

open as a page

What does returning a function from a sys.settrace 'call' event do?

level: middleimportance: must knowfreq 34%

basics

~20 s

The global callback set by sys.settrace fires only on 'call' events, and whatever it returns becomes that frame's local trace function, receiving the frame's 'line', 'return' and 'exception' events. Returning None means that frame is not instrumented further.

open as a page

What three sys.monitoring calls must you make before an event callback fires?

level: juniorimportance: should knowfreq 12%

basics

~10 s

Claim a slot with sys.monitoring.use_tool_id, attach a function with sys.monitoring.register_callback, then turn the event on with sys.monitoring.set_events. Skip any one of the three and nothing is delivered; using an unclaimed tool id raises ValueError.

open as a page

What does sys.addaudithook install in CPython, and which operations raise audit events?

level: middleimportance: should knowfreq 20%

basics

~20 s

sys.addaudithook registers a callable that CPython invokes for every audit event, passing the event name and an argument tuple. The runtime raises events at sensitive points: compiling and executing code, importing, opening files, and starting a subprocess.

open as a page

In a Python profile, what separates wall-clock sampling from CPU-time sampling?

level: middleimportance: should knowfreq 40%

basics

~10 s

Wall-clock sampling fires on elapsed time, so blocked threads keep accumulating samples and waiting shows up. CPU-time sampling fires in proportion to processor time consumed, so a thread parked on I/O contributes nothing.

open as a page

How do sys.settrace and sys.setprofile differ in the events they deliver?

level: middleimportance: should knowfreq 30%

basics

~20 s

A trace function set by sys.settrace fires per executed line as well as on calls, returns and exceptions, and uses per-frame local callbacks. A profile function set by sys.setprofile fires only on call and return, plus calls into C functions, and its return value is ignored.

open as a page

A hook added with sys.addaudithook slowed a 6-hour nightly ETL export and then aborted it -- where do the cost and the failure come from?

level: seniorimportance: should knowfreq 14%

basics

~20 s

Installing any hook turns every audit point into a Python call plus a tuple, and the hook sees all events, not just its own. The abort is worse: an exception inside the hook propagates out of the audited operation.

open as a page

How would you use sys.monitoring.set_local_events to trace only the two functions suspected of a floating-point rounding drift in a feature-flag service's 340-case regression pack?

level: seniorimportance: should knowfreq 18%

basics

~20 s

Claim a tool id, register callbacks, then call sys.monitoring.set_local_events on each suspect function's code instead of the global sys.monitoring.set_events. Only those code objects are instrumented, so the rest of the service and the passing cases run uninstrumented.

open as a page

How do you profile a live Python image-thumbnail worker without restarting it?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Attach a sampling profiler from outside the process. An out-of-process sampler reads the target's memory, walks its interpreter frames and aggregates stacks without redeploying, restarting, or adding per-call instrumentation to the hot path.

open as a page

Why does a tracer installed with sys.settrace see nothing from a webhook receiver's worker threads?

level: seniorimportance: should knowfreq 24%

basics

~20 s

The trace function is per-thread state. sys.settrace instruments only the thread that calls it, so worker threads started later begin untraced. Use threading.settrace for threads started afterwards, or call sys.settrace at the top of each worker.

open as a page

When is Python's sys.audit event stream the right basis for fleet-wide runtime security monitoring, and what does it miss?

level: principalimportance: should knowfreq 10%

basics

~20 s

Use the sys.audit stream when you control the base image and want semantic, pre-execution visibility of what Python code does. It is not an isolation boundary, and it sees nothing that bypasses CPython's own entry points.

open as a page

How do you choose the sample rate for continuous Python stack sampling in production?

level: principalimportance: should knowfreq 24%

basics

~20 s

Work backwards from the decisions the profile must support. Overhead is sample rate times stack-walk cost, so pick the lowest rate that still lands enough samples in each interesting code path, then measure that cost rather than guessing it.

open as a page

What does sys.settrace() install in CPython, and which tools rely on it?

level: juniorimportance: nice to knowfreq 18%

basics

~20 s

sys.settrace installs a Python callback that the interpreter invokes on execution events in the calling thread: a Python function being entered, a line executed, a return, an exception. Debuggers built on bdb.Bdb and coverage tools are built on it.

open as a page

Why can an audit hook added with sys.addaudithook never be removed from a running interpreter?

level: middleimportance: nice to knowfreq 12%

basics

~20 s

CPython exposes no hook-removal API at all: once sys.addaudithook accepts a callable, it stays for the life of that interpreter. The irrevocability is deliberate -- an audit facility the audited code could switch off would be worth nothing.

open as a page