What happens to the GIL while zlib.compress runs on a large buffer?
answer
- The lock is not held throughout
- C code can hand the lock back
- Only raw bytes are touched meanwhile
- Hashing and compression scale, loops do not
- Re-acquired before the result object is built
basics
~20 sThe C code inside zlib drops the global interpreter lock before it starts deflating and takes it back when the call finishes. Other Python threads run real bytecode during that window, so threaded compression genuinely overlaps.
solid answer
~50 sThe GIL is not held for the whole call. `zlib.compress` validates its arguments while holding the lock, then releases it around the deflate loop, which is pure C touching only raw bytes, and re-acquires it before building the result object. During that released window another thread can be scheduled and execute Python bytecode, so N threads compressing large buffers really do use N cores. The same is true of `hashlib` digests over large inputs and of blocking I/O calls such as socket and file reads. What never overlaps is Python-level work: a loop, a comprehension or arithmetic in your own code holds the lock the whole time. The practical rule is that threads help when the hot work is inside a C call that releases the lock, and help nothing when the hot work is bytecode.
code
python · 20 linesimport hashlib
import time
from concurrent.futures import ThreadPoolExecutor
blob = bytes(32 * 1024 * 1024)
def native_work(_):
return hashlib.sha256(blob).hexdigest()
def python_work(_):
total = 0
for i in range(2_000_000):
total += i
return total
for fn in (native_work, python_work):
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as pool:
list(pool.map(fn, range(4)))
print(fn.__name__, f"{time.perf_counter() - start:.2f}s")go deeper
Be ready to say plainly that the lock is dropped around long native work such as compression, hashing and blocking I/O, and taken back before a Python object is built. Have one concrete example ready in each direction.
Explain the three phases of such a call: parse and pin under the lock, crunch raw bytes released, build the result under the lock again. Mention the size threshold that decides whether releasing is worth it.
Show you reason about where the time is actually spent rather than about the CPU-bound label. Be able to point at a threaded workload and say which spans overlap and which serialize, and to measure it rather than guess.
Own the framing for the team: thread-friendliness is a property of the libraries in the hot path, not of the workload description, and it should be a stated criterion when a dependency is chosen. Be clear about when this pushes you to processes or a separate build instead.
## What the lock actually protects CPython's global interpreter lock protects interpreter state: the object reference counts, the memory allocator's internal free lists, the type cache, and the bookkeeping the evaluation loop touches on nearly every bytecode. Any thread that wants to execute bytecode, or to call almost any C-API function, must hold it. That is why two threads running Python code never execute simultaneously in the default 3.14 build. The lock does not protect memory that is not a Python object. A block of bytes that has already been located, whose owning object is guaranteed to stay alive, and which nothing else is going to mutate, can be read and transformed by C code with no interpreter involvement at all. That is exactly the shape of a compression pass, a hash update, a decode loop, or a blocking read into a fixed buffer -- and it is the shape that makes releasing the lock legal. ## What a call like zlib.compress does A well-written extension function follows the same three phases: 1. **Holding the lock**, it parses arguments, gets a pointer to the input bytes, and makes sure the input object stays alive for the duration of the call. 2. **With the lock released**, it runs the long native work over that raw memory. Between release and re-acquire it must not touch any Python object at all -- no attribute access, no reference counting, no raising. 3. **Holding the lock again**, it wraps the produced bytes in a new Python object and returns it. The C source expresses phase 2 with a matched macro pair, `Py_BEGIN_ALLOW_THREADS` and `Py_END_ALLOW_THREADS`, which save the thread's interpreter state and drop the lock, then take it back and restore the state. ## Which standard-library calls do this Broadly three families: * **Blocking I/O** -- file reads and writes, socket operations, `subprocess` waits, `time.sleep`. The lock is released for as long as the operating system call blocks, which is why threads have always been a reasonable answer for I/O concurrency. * **Compute-heavy pure-C work over bytes** -- `hashlib` digests, `zlib` and the other compression modules, encoding and decoding paths. These scale across cores for large enough inputs. * **Third-party compiled extensions** that follow the same discipline. A compiled array or numeric library that spends its time inside C loops over a contiguous block typically releases the lock around them; that is why some libraries appear to defeat the GIL and pure Python never does. A useful negative example is a small call. `hashlib.sha256` on a 40-byte string never releases anything worth releasing: CPython only bothers past a size threshold, because the release-and-reacquire round trip costs more than the work itself. Granularity, not the module name, decides whether you see parallelism. ## Why pure Python cannot do the same A pure-Python loop executes bytecode, and bytecode requires the lock by definition. There is no point in the loop where the interpreter could safely hand the lock away, because every step manipulates objects and reference counts. This is the whole reason the advice `threads for I/O, processes for CPU` exists -- and the reason the advice is not quite right: threads are also fine for CPU work whose CPU is burned inside a native call that releases the lock. ## The 3.14 free-threaded build Since 3.13 CPython has shipped an experimental build with no GIL, and in 3.14 that free-threaded build is officially supported (PEP 779), at a roughly 5-10% single-threaded cost. It is a separate build, not the default interpreter, and it changes the framing rather than the mechanics: there is no process-wide lock to release, so pure-Python threads can also run in parallel. Extensions still detach from the interpreter around long native work, and the rule that a detached thread must not touch Python objects is unchanged. For the default 3.14 build that almost everyone runs, the answer above stands exactly as stated. ## What to say in an interview Name the boundary: the lock covers interpreter state, native code that has stopped touching interpreter state can give it back, and the standard library does so around blocking calls and long byte-crunching. Then give one concrete pair -- `hashlib` and `zlib` scale across threads, a Python `for` loop does not.
- Does hashlib release the lock for every call, however small the input?No. Releasing and re-acquiring the lock costs a system-level handoff, so CPython only bothers when the input is large enough for the native work to dominate. Hashing a short string runs entirely under the lock, which is the right trade: the release round trip would cost more than the digest. This is why threaded hashing of many tiny values shows no speedup while threaded hashing of large blobs does.
- If threads can overlap inside zlib, why is multiprocessing still recommended for CPU-bound Python?Because most CPU-bound Python is bytecode, not native calls. A loop, a comprehension, attribute access and arithmetic all execute in the evaluation loop and hold the lock throughout, so threads serialize them. Threads win only when the hot span is inside a C call that has released the lock. The honest test is where the time is actually spent, not whether the workload is labelled CPU-bound.
- Can pure Python code release the lock explicitly?No. There is no Python-level API to drop it, and there could not be: bytecode execution requires the lock by definition. `time.sleep` releases it for you while it waits, and any call into a cooperating native function releases it for the duration of the native work, but a Python statement cannot hand it away. Moving work into a compiled extension, a subprocess or a subinterpreter is how you get past it.
The lock is like the single key to a shared workshop. A worker who has already carried their materials out to the yard can hang the key back on the hook while they cut wood outside, and pick it up again only when they need to come back in and put the finished piece on the shelf.
saying these in an interview costs you the question
- Claims the GIL is held for the entire duration of any call
- Says threads are useless for every CPU-bound workload
- Believes threads never run in parallel in CPython, ever
- Thinks a pure-Python loop can somehow release the lock
- Assumes every hashlib or zlib call releases it regardless of size
- Confuses releasing the lock with running in a free-threaded build