skip to content

Thread Safety and Races in Python

The GIL serializes bytecode but does not make code thread-safe: a read-modify-write spanning bytecodes still races. Interviewers use the counter example and expect queues and confinement as defaults.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why can two Python threads running `counter += 1` lose increments despite the GIL?

level: juniorimportance: must knowfreq 75%

answer

  1. One lock does not equal safe code
  2. Count the steps in one statement
  3. Read, add, store — a gap between them
  4. The demo printing the total proves nothing
  5. A call in the window loses increments

basics

~20 s

counter += 1 is three steps — read the value, add one, store it back — and another thread can run between them. Both threads then read the same value and one increment is silently lost.

solid answer

~50 s

The GIL guarantees that only one thread executes Python bytecode at a time; it does **not** make a statement atomic. `counter += 1` compiles to a load, an add and a store, so if thread A loads 10 and is switched out before storing 11, thread B also loads 10 and stores 11, and A then stores 11 too — one increment vanished. Nothing is corrupted, so you get a wrong number rather than a crash. Worth knowing on CPython 3.14: the interpreter checks for a pending thread switch only at a few points, chiefly a loop's back edge and function entry, so the textbook tight-loop demo often prints the *exact* total. That is an accident of the loop's shape, not safety — put any call between the read and the write and the increments disappear. Fix it with a `threading.Lock` around the whole read-modify-write, or per-thread counters summed after `join()`.

code

python · 17 lines
python
import threading

counter = 0

def bump(n):
    global counter
    for _ in range(n):
        counter += 1

threads = [threading.Thread(target=bump, args=(200_000,)) for _ in range(4)]
for t in threads:
    t.start()
for t in threads:
    t.join()

# Usually exact: no switch point falls between the load and the store here.
print(counter, "of", 800_000)

go deeper

for a junior

Be ready to say plainly that the GIL serializes bytecode, not statements, and to name the three steps hiding inside +=. Knowing the symptom is a wrong total rather than a crash is most of the answer.

for a middle

Explain the interleaving concretely with values, and show the fix you would ship: a lock around the whole read-modify-write, or per-thread counters summed after join(). Be able to point at the load and the store in a disassembly.

for a senior

Expect the follow-up about a demo that prints the exact total, and be able to say why — switch checks land at loop back edges and function entry, not inside the statement — and why that is not safety. Argue for designs with no shared mutable counter at all.

for a principal

Own the policy angle: correctness that depends on where the interpreter happens to check for a thread switch is not a property a codebase can maintain. Push conventions — confine and aggregate, or hand off through a queue — instead of reviewing every increment for a missing lock.

### What the GIL actually promises The Global Interpreter Lock is a single interpreter-wide mutex that a thread must hold to run Python bytecode in CPython's default build. It exists to protect the interpreter's own data — reference counts, container internals, the allocator — not the invariants of your program. The common mis-reading is "only one thread runs at a time, therefore my code is thread-safe". The unit of serialization is a *bytecode instruction*, not a *line of source*, and almost every interesting update spans several instructions. ### Decomposing the increment On CPython 3.14, a module-level `counter += 1` inside a loop disassembles to this: ``` LOAD_GLOBAL 2 (counter) LOAD_SMALL_INT 1 BINARY_OP 13 (+=) STORE_GLOBAL 1 (counter) JUMP_BACKWARD 18 ``` The value is read by `LOAD_GLOBAL` and published by `STORE_GLOBAL`, and everything between those two instructions is a window. If another thread runs in that window, the interleaving ``` thread A: load counter -> 10 thread B: load counter -> 10 thread B: add, store -> 11 thread A: add, store -> 11 ``` is legal, and two increments advance the counter by one. That is the **lost update**: a read-modify-write where the value read is stale by the time it is written. ### Why the result is a wrong number, not a crash Under the GIL you cannot get a torn integer, corrupt the dict holding the global, or break a reference count — those are exactly what the GIL protects. That is what makes this bug dangerous: it produces a plausible but wrong number, silently, only under load. A counter of processed items reads 998,412 instead of 1,000,000, and nobody notices until a reconciliation job disagrees. ### Why the textbook demo may print the exact total on 3.14 Run the classic demo — four threads, a few hundred thousand `counter += 1` iterations each — on CPython 3.14 and you will very likely see the exact expected total, even with sixteen threads and sixteen million increments. This surprises people who learned the example a decade ago, and the explanation matters more than the folklore. CPython does not test for a pending thread switch after every instruction; it checks at a few designated points, chiefly the back edge of a loop and the `RESUME` at the top of a called function. In the disassembly above there is no such check between `LOAD_GLOBAL` and `STORE_GLOBAL`: the only check in the loop body sits at `JUMP_BACKWARD`, *after* the store. So the switch lands between iterations, not inside the statement, and no update is lost. The moment anything runs Python code between the read and the write, the check reappears inside the window and the updates vanish: ```python def one(): return 1 counter = counter + one() # the call is a switch point, mid-window ``` Four threads doing that lose roughly half their increments. This is the crucial lesson to carry away: **the demo printing an exact total proves nothing about safety.** The window's exposure depends on the exact shape of the code — a function call, a property, a `__add__` on a user class, a logging call, a comparison against an object with a Python-level `__eq__`. A refactor that adds any of those turns a correct-looking loop into a racy one, with no change to the line you are staring at. ### Reproducing it deliberately ```python import sys, threading sys.setswitchinterval(1e-6) # switch more often, to widen the demo counter = 0 def one(): return 1 def bump(n): global counter for _ in range(n): counter = counter + one() threads = [threading.Thread(target=bump, args=(300_000,)) for _ in range(4)] for t in threads: t.start() for t in threads: t.join() print(counter, "of", 1_200_000) ``` `sys.setswitchinterval` lowers the interval at which a waiting thread requests the GIL, so the switch points are hit far more often; it does not create the race, it only makes it obvious. ### A nuance worth knowing Not every `+=` is equally exposed. `shared_list += [item]` calls the list's in-place concatenation, which mutates the existing list and returns *the same object*; the store that follows merely rebinds the name to the object it already held, so nothing is lost. `n += 1` on an `int` is different because integers are immutable: the addition produces a *new* object and the store is the only thing that publishes it, so a stale read becomes a stale write. The rule is not "avoid `+=`" but "any update whose new value depends on a previously read value must be serialized". ### The three real fixes 1. **A lock around the whole read-modify-write.** Wrap the increment in `with counter_lock:` so load, add and store form one critical section. Correct and obvious; the cost is contention when the counter is hot. 2. **Message passing.** Workers push results into a `queue.Queue` and one consumer owns the total, so there is no shared mutable state left to race on. 3. **Confinement.** Give each thread a plain local counter and sum the per-thread values after `join()`. Usually the fastest and the easiest to reason about: the shared write happens once per thread, after the work is done. ### What "thread-safe" means here A workable definition: an operation is thread-safe if no interleaving of other threads can produce a result that no sequential execution could produce. Under that definition the GIL buys almost nothing at the application level. Interviewers ask this question because a candidate who believes the GIL confers thread safety will write racy counters, racy caches and racy check-then-act code — and, on 3.14, will point at a demo that printed the right number as proof they were fine. ### Version note This is CPython 3.14's default build. The free-threaded build became officially supported in 3.14 (PEP 779); with no GIL the threads genuinely run at once, so the same read-modify-write loses updates without needing any switch point in the window at all.

  • On CPython 3.14 a four-thread `counter += 1` demo prints the exact total. What do you conclude?
    Only that no switch point happened to fall between the load and the store in that loop. CPython checks for a pending thread switch at a few places — a loop's back edge, function entry — and in a tight increment loop the check sits after the store, so the threads interleave between iterations rather than inside the statement. Add any call, property or user-defined `__add__` between the read and the write and the increments start disappearing. The code was never safe.
  • Why is `shared_list += [item]` far less dangerous than `n += 1` on a shared integer?
    In-place concatenation mutates the existing list object in one C-level call and returns that same object, so the store that follows rebinds the name to the object it already held. An integer is immutable: the addition builds a new object, and the store is the only step that publishes it, so a stale read becomes a stale write. The hazard is read-modify-write on an immutable value, not the operator itself.
  • If you must keep a shared counter, what is the cheapest correct design?
    Confine it: each worker increments a plain local counter, and you sum the per-thread totals after joining the threads. The shared write then happens once per thread instead of once per item, so there is nothing hot to contend on. Use a lock around the increment only when other threads must read a live running total, and a `queue.Queue` when one thread should own the aggregate.

Two clerks each photocopy the ledger's current total, add their own sale to the copy, and write the copy back. The second write erases the first sale, even though only one clerk was ever at the desk at a time.

saying these in an interview costs you the question

  • Claims the GIL makes Python code thread-safe
  • Says one line of Python is one atomic operation
  • Concludes the loop is safe because a demo printed the exact total
  • Thinks the risk is memory corruption rather than a lost update
  • Cannot name the read-modify-write steps behind `+=`
  • Believes removing the GIL would make the increment atomic

context

open as a page

Which CPython list and dict operations are effectively atomic across threads?

level: middleimportance: must knowfreq 55%

basics

~20 s

Single C-level calls that never run Python bytecode are effectively atomic: list.append, list.pop, d[k] = v, x = d[k], d.setdefault, d.update. Anything built from two of them — d[k] += 1, or a membership test followed by an assignment — is not.

open as a page

Python worker threads all rebuild the same cache on a 45-second cold start — what race is this and how do you fix it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It is a check-then-act race: every thread finds the cache empty during the 45-second build window and starts its own build. Serialize the build behind a lock and re-check inside it, or build once eagerly before the pool starts accepting work.

open as a page

How do you confirm a live Python process is deadlocked on two threading locks?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Dump every thread's stack from inside the process — faulthandler.dump_traceback_later(30, repeat=True), or faulthandler.register on a signal — and look for two threads parked in acquire() at different call sites. Zero CPU with no progress and no exception is the giveaway.

open as a page