skip to content

What does Py_BEGIN_ALLOW_THREADS do inside a CPython C extension?

level: middleimportance: should knowfreq 40%

answer

  1. It comes as a matched pair
  2. Something is saved before the lock goes
  3. The macros open and close a brace
  4. A local variable holds the thread state
  5. Balance them or the thread runs unlocked

basics

~10 s

It saves the calling thread's interpreter state into a local variable and drops the global interpreter lock, so other threads can run. The matching Py_END_ALLOW_THREADS re-acquires the lock and restores that saved state.

solid answer

~40 s

The two macros form a matched pair that brackets a span of native work. `Py_BEGIN_ALLOW_THREADS` opens a block, declares a hidden local that holds the thread's saved interpreter state, detaches the thread and releases the lock; `Py_END_ALLOW_THREADS` re-acquires the lock, restores the state from that local and closes the block. Because they open and close braces, they must be balanced in the same scope -- you cannot `return` or jump out of the middle, and every error path inside must reach the END macro. The span between them must be pure C over memory that stays valid: no Python objects, no reference counting, no raising. Extensions wrap blocking calls and long byte-level loops this way, and skip it for short work, where the release-and-reacquire round trip costs more than it saves.

code

python · 21 lines
python
import time
import zlib
from concurrent.futures import ThreadPoolExecutor

blob = bytes(8 * 1024 * 1024)

def one_long_span(_):
    return len(zlib.compress(blob, 6))

def many_short_spans(_):
    comp = zlib.compressobj(6)
    size = 0
    for i in range(0, len(blob), 1024):
        size += len(comp.compress(blob[i:i + 1024]))
    return size + len(comp.flush())

for fn in (one_long_span, many_short_spans):
    start = time.perf_counter()
    with ThreadPoolExecutor(max_workers=4) as pool:
        list(pool.map(fn, range(4)))
    print(fn.__name__, f"{time.perf_counter() - start:.2f}s")

go deeper

for a junior

Recall that a C extension gives the interpreter lock back around long native work using a matched macro pair, and takes it back before returning any Python object. The names and the pairing are enough at this level.

for a middle

Explain the expansion: save the thread state into a local, release, then restore and re-acquire, with braces that force the pair to balance in one scope. Say why an early return between them is fatal.

for a senior

Demonstrate judgement about granularity -- one long released span beats many short ones, and a per-record call can be correct yet slower than never releasing. Know the callback case and which API covers it.

for a principal

Own the boundary design: decide the chunk size at which work crosses into native code so that released spans are long enough to matter, and treat that contract as part of the interface you review, not an implementation detail.

## The macro pair Every thread that runs Python code owns a thread state -- the structure holding its frame stack, its current exception, and its link to the interpreter. When a native function is about to do work that needs none of that, it detaches: it stores its thread state somewhere it can find again, releases the global interpreter lock, and lets other threads in. When the work is done, it takes the lock back and reattaches the saved state so the C-API becomes usable again. That whole dance is the macro pair: ```c Py_BEGIN_ALLOW_THREADS result = crunch(buffer, length); /* pure C, no Python objects */ Py_END_ALLOW_THREADS ``` The macros expand to something like `{ PyThreadState *_save = PyEval_SaveThread();` and `PyEval_RestoreThread(_save); }`. Two consequences follow directly from that expansion and are worth stating in an interview: * They **open and close a brace**, so they must be balanced within one scope. You cannot release in one function and re-acquire in another, and you cannot leave the block by `return` or `goto` without re-acquiring first. * The saved state lives in a **local variable**, so it is per call, not per thread-global. Nesting is possible but rarely what you want. ## What may run in the released span Only code that touches no interpreter state: arithmetic, memory copies, compression, hashing, an operating-system call, a library entry point that itself knows nothing about Python. Anything else is undefined behaviour rather than a caught error, which is what makes this a discipline question rather than an API question. The practical pattern in a well-written extension is to do all Python-side work up front while still holding the lock: parse arguments, obtain a pointer to the input memory, take a reference so the owner cannot be collected, and copy out any small values you need. Then release, crunch, re-acquire, and only then allocate the result object. ## Granularity is the whole design question Releasing is not free. Each round trip is a lock handoff and a scheduling opportunity, and on a contended lock re-acquiring may mean waiting behind other threads. That gives a simple rule: release around **one long span**, never around **many short ones**. This is where real code goes wrong. Two extensions can both release the lock and behave completely differently under threads, purely because one processes a whole buffer per call while the other is called once per small record. In the second case each call pays parse, release, tiny crunch, re-acquire, allocate -- and the lock spends most of the wall clock being handed around rather than being released for anything useful. The fix is on the calling side: feed the extension bigger chunks so the released spans are long. The standard library shows both sides of this. A single `zlib.compress` over a large buffer is one long released span. The same data pushed through `zlib.compressobj` a kilobyte at a time is hundreds of short spans plus hundreds of Python-level calls, and the threading behaviour collapses accordingly. ## Re-entering Python from the released span Sometimes native code genuinely must call back into Python while detached -- a callback, a progress hook, an error translator. The correct move is not to reach for the macros again, but the `PyGILState_Ensure` / `PyGILState_Release` pair, which acquires the lock and attaches a thread state whether or not the current thread ever had one. That is what a thread created by a C library, which the interpreter has never seen, has to use before touching any Python object. ## Failure modes * **Unbalanced macros**: an early `return` between them leaves the thread running Python code without the lock. The result is corruption, usually far from the cause. * **A Python object touched inside**: reference counts are not atomic in the default build, so a concurrent increment and decrement can be lost, freeing a live object or leaking a dead one. * **A dangling pointer**: the object owning the buffer is collected or resized during the released window because no reference was held. * **Releasing too finely**: correct, but slower than not releasing at all. ## Free-threaded builds In the 3.14 free-threaded build, officially supported under PEP 779, there is no process-wide lock to release, but the macro pair still exists and still matters: it detaches the thread from the interpreter, which is what lets a stop-the-world pause proceed while a long native call is in flight. Extensions keep using it, and the rule about not touching Python objects in the released span is unchanged.

  • What goes wrong if a C function returns from between the two macros?
    The macros are brace-balanced, so the compiler usually rejects it outright; where the code is restructured to escape anyway, the thread carries on without the lock and with no attached thread state. Every subsequent C-API call is then undefined -- typically a corrupted reference count or a crash somewhere unrelated. Extension code handles errors inside the released span by setting a flag, reaching the END macro normally, and raising afterwards.
  • When would you use PyGILState_Ensure instead of these macros?
    When a thread needs the lock and has no attached state to restore -- typically a thread created by a C library that the interpreter has never seen, calling back into Python. `PyGILState_Ensure` acquires the lock and attaches a state, creating one if needed, and `PyGILState_Release` undoes exactly that. The allow-threads macros are the mirror image: they assume you already hold the lock and want to give it up temporarily.
  • Is releasing the lock always the right call for a native function?
    No. For short work it is a net loss, because the release and re-acquire round trip costs more than the computation and adds a scheduling point on a contended lock. Extensions typically apply a size threshold, releasing only for inputs large enough that the native span dominates. Releasing around a call that is invoked once per small record is the classic mistake -- correct, and slower than holding the lock throughout.

saying these in an interview costs you the question

  • Thinks the macros can be split across different functions
  • Believes releasing the lock is always a performance win
  • Says the macros make the surrounded code thread-safe
  • Confuses them with PyGILState_Ensure for foreign threads
  • Assumes the saved thread state is global rather than a local

context