What does CPython's GIL actually protect inside the interpreter?
answer
- It guards the interpreter, not your data
- Every object carries a counter
- Read-modify-write on shared memory
- Coarse lock keeps the common case cheap
- Allocator freelists and caches too
basics
~20 sIt protects CPython's own mutable state: above all every object's reference count, and alongside it interpreter-wide structures such as allocator freelists and internal caches. Serialising bytecode execution keeps those updates correct without an atomic operation on each one.
solid answer
~50 sEvery CPython object carries a reference count that is incremented and decremented constantly, on nearly every operation, including ones that look read-only. If two threads updated the same count without synchronisation, the counts would drift: a lost increment frees an object still in use, a lost decrement leaks it. Locking each object individually would mean an atomic operation on almost every bytecode instruction, plus lock-ordering hazards, so CPython took the coarse route — one interpreter-wide mutex, held while bytecode runs. That also covers other shared interpreter state such as allocator freelists and internal caches. The payoff is that single-threaded code, which is most code, pays nothing: an uncontended refcount update is a plain non-atomic integer bump. The cost is that CPU-bound threads cannot overlap. The free-threaded build, supported since 3.14, buys that back with atomic and deferred refcounting at a 5 to 10 percent single-threaded penalty.
code
pycon · 8 lines>>> import sys
>>> data = [1, 2, 3]
>>> alias = data
>>> sys.getrefcount(data)
3
>>> del alias
>>> sys.getrefcount(data)
2go deeper
You mainly need the headline: CPython counts references on every object, those counters are updated constantly, and one lock keeps concurrent updates from corrupting them.
Explain the read-modify-write hazard concretely, why a coarse lock keeps the single-threaded case free of atomics, and that the lock covers other interpreter internals such as allocator freelists as well.
Draw the line between interpreter consistency and application consistency, and explain why the accumulated assumption of serialised execution across the codebase, not the mutex itself, is what makes removal a multi-year project.
Frame it as the design tradeoff it is: the interpreter optimised for the dominant single-threaded case and charged multi-threaded compute for it. Be able to argue when paying the free-threaded build's single-threaded penalty is the right call for a fleet.
## The invariant that needs protecting Ask why the GIL exists and the answer is not "because threading is hard" but something much more specific: **CPython mutates shared interpreter state on almost every bytecode instruction, and the most frequently mutated piece of that state is the per-object reference count.** Every object CPython allocates carries a counter of how many references currently point at it. Binding a name, pushing a value on the evaluation stack, passing an argument, returning a result, storing into a list — each of these adjusts counters, and so does simply reading a value, because the reader takes a reference while it works. A trivial loop performs millions of these updates a second. A counter update is read-modify-write. Two threads doing it on the same object with no synchronisation can interleave so that one update is lost. Lose an increment and the count reaches zero while a live reference still exists, so the object is freed underneath a thread that is using it, which is a use-after-free crash or silent corruption. Lose a decrement and the object is never freed, which is a leak. Neither failure is reproducible, and both surface far from the code that caused them. ## Why one big lock rather than many small ones The obvious alternative is to make refcount updates atomic, or to lock each object. Both were tried, historically and recently, and both are expensive for the common case: * An atomic increment is far more costly than a plain one, and it forces cache-line ownership to bounce between cores whenever two threads touch the same hot object — and in Python the hottest objects are shared by construction: `None`, `True`, small integers, interned strings, every module and class object. * Fine-grained per-object locks multiply that cost and introduce lock ordering, which means deadlock risk in the interpreter itself. One coarse lock, held for the duration of bytecode execution, makes the whole question disappear. Inside the lock, a refcount update is an ordinary non-atomic integer operation on memory only one thread can be touching. Single-threaded programs — still the overwhelming majority — pay essentially nothing. ## What else it covers Reference counts are the headline, but the GIL is a blanket over CPython's shared mutable internals generally: the object allocator's freelists and arenas, various internal caches, and assorted interpreter bookkeeping. Because bytecode execution is serialised, none of these needed their own synchronisation, and decades of interpreter code was written on that assumption. That accumulated assumption, not the lock itself, is the reason removing the GIL has been such a long project. ## What it does not protect This is the part interviews probe, so be precise. The GIL guarantees that the interpreter's *own* structures stay consistent. It guarantees nothing about the consistency of *your* multi-step operations, because it is released and reacquired between bytecode instructions. An operation that reads a value, computes a new one and stores it back is several instructions, and a thread switch can land between any two of them. Application-level invariants still need application-level synchronisation. ## The naming subtlety "Global" is now slightly historical. Since 3.12 (PEP 684) each interpreter in a process can own its own GIL, so the lock is interpreter-wide rather than process-wide; 3.14 exposes multiple interpreters from the standard library as `concurrent.interpreters` and an interpreter-backed executor (PEP 734). Two interpreters in one process can therefore execute bytecode simultaneously — because they do not share the object state the lock protects. ## What changes when the lock goes away The free-threaded build, experimental in 3.13 and officially supported in 3.14 (PEP 779), does not simply delete the lock. It has to replace what the lock was doing: biased and deferred reference counting so that the hottest objects avoid contention, immortal objects whose counts are never touched, atomic operations where they are unavoidable, and per-object locking for containers. The measured price is a 5 to 10 percent single-threaded slowdown, which is exactly the cost the GIL was originally buying away. ## How to say it in an interview A strong answer is three beats. One: CPython refcounts every object, and refcount updates are read-modify-write on shared memory. Two: making each of those safe individually would tax the single-threaded case that dominates, so CPython serialises bytecode execution with a single lock instead. Three: the lock keeps the *interpreter* consistent, not your data structures. If you can add that removing it means replacing the mechanism rather than deleting it, you have said everything the question is looking for.
- Why not just make every reference-count update atomic instead of holding one lock?Because it taxes the case that dominates. An atomic increment is much dearer than a plain one, and Python's hottest objects are shared by design — `None`, small integers, interned strings, module and class objects — so the cache line holding their count would ping-pong between cores continuously. Single-threaded programs would pay that cost for a benefit only multi-threaded compute workloads collect. The free-threaded build pays it anyway, softened by deferred counting and immortal objects, and still measures a 5 to 10 percent single-threaded penalty.
- If the GIL serialises bytecode, why does application code still need its own synchronisation?Because the lock is released between bytecode instructions, not around your logical operations. Anything that reads a value, computes from it and writes it back spans several instructions, and a thread switch can occur between any two of them. The GIL guarantees the interpreter's internal structures stay consistent; it makes no promise about the consistency of the invariants your code maintains across multiple steps.
- Is the lock really global to the process?Not since 3.12. PEP 684 gave each interpreter its own GIL, so the correct description is interpreter-wide. Two interpreters running in one process can execute bytecode at the same time precisely because they do not share the object state the lock protects. Python 3.14 exposes this from the standard library as `concurrent.interpreters` together with an interpreter-backed executor, under PEP 734.
It is a single shop-wide till key rather than a separate lock on every drawer: one key is clumsy when several assistants are serving at once, but it costs nothing when there is only one assistant, and nobody has to remember which order to unlock the drawers in.
saying these in an interview costs you the question
- Says the GIL exists because Python is interpreted
- Cannot connect the lock to reference counting at all
- Claims the GIL protects application data structures
- Thinks removing the lock is a one-line deletion
- Believes fine-grained locking would be strictly better
- Says the lock is still process-wide in every version