skip to content

Audit Events and sys.audit

The security-facing event stream: install a hook that sees every sensitive operation the runtime reports, raise your own events, and accept that a hook can never be uninstalled once added.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What does sys.addaudithook install in CPython, and which operations raise audit events?

level: middleimportance: should knowfreq 20%

answer

  1. A runtime-wide stream of sensitive operations
  2. PEP 578, shipped in Python 3.8
  3. One call raises an event, one subscribes
  4. Hook gets an event name and args tuple
  5. sys.audit and sys.addaudithook

basics

~20 s

sys.addaudithook registers a callable that CPython invokes for every audit event, passing the event name and an argument tuple. The runtime raises events at sensitive points: compiling and executing code, importing, opening files, and starting a subprocess.

solid answer

~40 s

PEP 578 gave the interpreter a runtime audit stream with two halves: `sys.audit(name, *args)` raises an event, and `sys.addaudithook(hook)` subscribes to every event raised afterwards. The hook is called as `hook(event: str, args: tuple)`, synchronously, on the thread that raised the event, in installation order; its return value is ignored, and an exception it raises propagates out of the audited operation and aborts it. CPython raises events from C at sensitive points -- `"compile"` and `"exec"`, `"import"` with the module name and filename, `"open"` with path, mode and flags (covering both the builtin `open` and `os.open`), `"subprocess.Popen"` with the executable and argument list, `"os.system"` -- and stdlib modules and your own code raise their own by calling `sys.audit`. There is no per-event subscription: every hook sees every event and filters itself.

code

python · 8 lines
python
import sys

def audit_hook(event, args):
    if event in {"compile", "exec", "subprocess.Popen"}:
        print("audit:", event, args[0])

sys.addaudithook(audit_hook)
exec("total = 6 * 60")

go deeper

for a junior

Recall that Python has a built-in event stream for sensitive operations, that sys.addaudithook subscribes to it, and that the hook is a plain function taking an event name and a tuple of arguments.

for a middle

Be ready to explain the hook contract precisely: called synchronously on the raising thread, in installation order, return value ignored, exceptions abort the audited call, and every hook sees every event and filters itself.

for a senior

Show that you know where the events actually come from and where the blind spots are -- the interpreter's own entry points are instrumented, arbitrary native code is not -- and that a hook must be installed before startup work to see it.

for a principal

Own the framing that this stream is detection and defence in depth rather than an isolation boundary, and be able to say what a team gets from it that operating-system-level tracing cannot give them.

## The two halves of the mechanism PEP 578, "Runtime Audit Hooks", shipped in Python 3.8 and is unchanged in shape on 3.14. It adds a *raise* side and a *listen* side. The raise side is `sys.audit(event, *args)`. It takes an event name -- by convention a dotted string namespaced to whoever raises it, such as `"subprocess.Popen"` or `"etl.export.start"` -- and any number of positional arguments describing the operation. It returns `None`; there is no result to inspect. The listen side is `sys.addaudithook(hook)`. The hook is any callable taking two positional parameters: the event name as a `str`, and a `tuple` of the arguments. It is invoked synchronously, in the thread that raised the event, before the audited operation completes, and hooks are invoked in the order they were installed. The return value is discarded, so a hook cannot approve or veto by returning; the only way it can influence the operation is by raising, and an exception raised inside a hook propagates out of the audited call and aborts it. ## Where the events come from Most interesting events are raised by CPython itself, from C, at points chosen for security relevance rather than for performance instrumentation: * `"compile"` with the source and filename, and `"exec"` with the code object -- so `exec`, `eval` and `compile` are all visible, including code assembled at runtime. * `"import"` with the module name, the resolved filename, and the import machinery's own search state (`sys.path`, `sys.meta_path`, `sys.path_hooks`). * `"open"` with the path, the mode and the OS flags, raised by both the builtin `open` and `os.open`. * `"subprocess.Popen"` with the executable, the fully resolved argument list, the working directory and the environment, plus `"os.system"` for the shell entry point. * Networking, `ctypes` and other sensitive stdlib surfaces raise their own events, and on 3.14 the new remote-execution entry point `sys.remote_exec` raises `"sys.remote_exec"`. Stdlib modules written in Python raise events by calling `sys.audit` directly, and application code can do the same. Because the namespace is a flat string space, prefix your own events with something you own so that a hook can filter them cheaply with `str.startswith`. ## What a hook can and cannot do A hook is an observer with one weapon. It sees the operation's *semantic* arguments -- an argv list before the process is spawned, a module name before it is loaded, a path before the file is opened -- which is strictly more informative than what an operating-system-level tracer sees after the fact. Raising from the hook is how a "deny subprocess launches in this service" policy is written, and it genuinely aborts the specific call. What it is not is a sandbox, and PEP 578 says so explicitly. The audit points exist where CPython chose to put them; native code that calls the C library directly, rather than going through the interpreter's own entry points, raises nothing. Treat the stream as high-quality detection and defence in depth, not as an isolation boundary -- isolation belongs to the process, container or operating-system layer. ## Cost, filtering and discovery There is deliberately no way to subscribe to a subset of events. Every installed hook is called for every event, and filtering happens in Python inside your hook. That is the single most important consequence for the shape of hook code: the first line should be the cheapest possible test -- membership of the event name in a `set`, or a prefix test -- with an immediate return for everything else. When no hook is installed at all, the audit points are close to free: the C-level audit call checks a flag and skips building the argument tuple. Installing one hook turns every audit point in the process into a Python-level call plus tuple construction. Discovering the event names is done empirically. A hook that counts event names into a `collections.Counter` and prints the totals at the end tells you exactly which events your workload actually raises and how often -- which is the right first step before writing any policy. Be careful what the hook itself does: writing to a file from inside a hook raises `"open"` again, so any I/O in a hook needs a re-entrancy guard or it will amplify. ## Installing it early enough A hook only sees events raised after it is installed, so the imports and file opens that happen during your own startup are invisible to a hook installed in `main()`. Installing as early as possible means a path-configuration file or a startup module picked up by the `site` machinery in the base image; an embedder can install a hook from C before the interpreter is initialised, which is the only way to catch everything.

  • Can a hook stop the operation it was told about, and how?
    Only by raising. The hook's return value is ignored, so there is no "deny" result; an exception raised inside the hook propagates out of the audited call and aborts it. That is how a policy such as "this service never launches subprocesses" is implemented, by raising on the `"subprocess.Popen"` event. It is not an isolation boundary, though: it only covers operations that pass through CPython's own audit points.
  • How would you find out which audit events a given workload actually raises?
    Install a hook that does nothing but tally event names into a counter, run the workload, and print the totals at exit. That gives you both the vocabulary and the frequency, which is what you need before writing any filtering policy. Keep the hook free of I/O while doing it -- writing to a file from inside a hook raises the `"open"` event again and feeds itself.
  • Why does a hook installed at the top of main() still miss things?
    Hooks only see events raised after installation, and a real application raises a great many `"import"` and `"open"` events before its own entry point runs. To cover startup you must install from the interpreter's own startup path -- a path-configuration file or a startup module executed by the `site` machinery -- or, for an embedded interpreter, from C before initialisation.

It is the interpreter's own CCTV feed: every sensitive doorway reports who walked through it, with the details, before the door opens. The camera can scream loudly enough to stop that one person, but it was never built to be the lock.

saying these in an interview costs you the question

  • Thinks an audit hook is a sandbox that contains untrusted code
  • Believes a hook returns True or False to allow or deny
  • Expects to subscribe to a single event name
  • Thinks hooks run asynchronously or on a background thread
  • Confuses audit events with a profiler or tracing callback
  • Assumes native code paths always raise an event

context

open as a page

A hook added with sys.addaudithook slowed a 6-hour nightly ETL export and then aborted it -- where do the cost and the failure come from?

level: seniorimportance: should knowfreq 14%

basics

~20 s

Installing any hook turns every audit point into a Python call plus a tuple, and the hook sees all events, not just its own. The abort is worse: an exception inside the hook propagates out of the audited operation.

open as a page

When is Python's sys.audit event stream the right basis for fleet-wide runtime security monitoring, and what does it miss?

level: principalimportance: should knowfreq 10%

basics

~20 s

Use the sys.audit stream when you control the base image and want semantic, pre-execution visibility of what Python code does. It is not an isolation boundary, and it sees nothing that bypasses CPython's own entry points.

open as a page

Why can an audit hook added with sys.addaudithook never be removed from a running interpreter?

level: middleimportance: nice to knowfreq 12%

basics

~20 s

CPython exposes no hook-removal API at all: once sys.addaudithook accepts a callable, it stays for the life of that interpreter. The irrevocability is deliberate -- an audit facility the audited code could switch off would be worth nothing.

open as a page