skip to content

True Parallel Execution

The two routes by which CPython finally runs bytecode on several cores at once — the free-threaded build and multiple interpreters — and what each one costs. Both became answerable questions in 3.14.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

11

Is `list.append` on a shared list still safe without a lock in free-threaded CPython 3.14?

level: juniorimportance: must knowfreq 55%

answer

  1. One operation versus several
  2. The GIL's critical section had a size
  3. Per-object locking replaced global locking
  4. Container consistent, program invariant not
  5. queue.Queue and deque are documented safe

basics

~10 s

Yes. The free-threaded build locks each container object internally, so a single list.append completes without corrupting the list or losing the item. Anything whose correctness spans two operations still races and still needs threading.Lock.

solid answer

~40 s

Yes, and for a reason worth stating precisely. Under the GIL, a call into a C-implemented container method ran to completion because a thread could only be switched out between bytecodes — atomicity came free at bytecode size. The free-threaded build removes that global serialization, so CPython rebuilt the guarantee per object: mutable built-ins take an internal per-object lock for the duration of one operation. So concurrent `list.append` calls never corrupt the list's storage, never lose an element and never crash the interpreter; four threads appending 1000 items each end with 4000 items in some interleaved order. What is *not* covered is any invariant spanning two operations — read then write, test then insert. For those, take a `threading.Lock` yourself.

code

python · 15 lines
python
import threading

items = []

def worker(n):
    for i in range(1000):
        items.append((n, i))  # one container operation

threads = [threading.Thread(target=worker, args=(n,)) for n in range(4)]
for t in threads:
    t.start()
for t in threads:
    t.join()

print(len(items))  # 4000, on the GIL build and the free-threaded build alike

go deeper

for a junior

Recall the one-line rule: a single built-in container operation is safe on its own, several of them in a row are not. Being able to say that cleanly is the whole expectation here.

for a middle

Explain the mechanism: the GIL gave a bytecode-sized critical section for free, and the free-threaded build reproduces it with per-object locking inside the container. Name where the guarantee stops.

for a senior

Show that you audit real code by invariant, not by feel — which sequences need an explicit threading.Lock, where a snapshot beats a lock, and why a lock per element destroys the parallelism you migrated for.

for a principal

Own the policy: what your codebase is allowed to assume about container atomicity, whether shared mutable state is permitted at all versus per-worker state merged once, and how you keep that rule reviewable as the team adopts free-threaded builds.

### What the GIL was quietly guaranteeing Before free-threading existed, a thread had to hold the global interpreter lock to execute bytecode, and it could only be descheduled between two bytecode instructions or when a C function voluntarily released the lock (typically around blocking I/O). That produced an unwritten form of atomicity: any operation implemented entirely in C that never called back into Python ran start to finish with no other thread executing Python. `list.append`, `dict.__setitem__`, `set.add` and `collections.deque.popleft` are all that shape, so each behaved as one indivisible step. The critical section was bytecode-sized and cost nothing to obtain. ### What the free-threaded build had to rebuild The free-threaded build (PEP 703; experimental in 3.13, officially supported in 3.14 by PEP 779, carrying a 5-10% single-threaded penalty) removes that free serialization. The guarantee therefore had to be reconstructed object by object. Mutable built-in containers are protected by internal per-object locking that the interpreter takes around a container operation, paired with a fast, mostly lock-free read path and a reference-counting scheme that keeps shared immutable objects from turning into a cache-line war. None of that is API you call; it is machinery underneath the container. For a Python author the practical consequence is a single sentence: **one operation on a built-in list, dict or set stays internally consistent.** Concurrent appends do not corrupt the list's internal array, do not lose an element, do not leave a half-written slot and do not segfault the interpreter. `len()` observes a coherent size. Four threads appending a thousand items each finish with four thousand items, in an interleaving you cannot predict. ### What the guarantee is not It is not a guarantee about anything spanning two operations. `d[k] = d[k] + 1` is a read and a write with a gap between them, and the gap is where the lost update lives. A membership test followed by an insert is the same defect wearing a different hat. It is not a guarantee about user-defined types. A container built out of Python-level attributes, or a `list` subclass whose `__setitem__` is written in Python, executes bytecode and can be interleaved anywhere. The internal locking protects the C implementation of the built-in, not your class that wraps it. It is not a licence to mutate a container while another thread iterates it. Changing a dict's size during iteration still raises `RuntimeError`, and iterating a list that another thread is appending to yields an unspecified — though non-crashing — view. If you need a stable view, take a snapshot (`list(shared)`) under whatever discipline protects the container. And it is emphatically not "free-threading made my threaded code correct". It made the *interpreter* safe under true parallelism. A program whose correctness leaned on the GIL's incidental serialization is not newly broken; it is now simply wrong far more often, because the window it was losing races in got wide. ### Which operations you can lean on Two things in the standard library are documented as safe rather than merely observed to be: `queue.Queue` is designed for producer/consumer hand-off between threads, and `collections.deque` documents thread-safe appends and pops at both ends. Beyond those, CPython does not publish a per-method atomicity table, and "it looked atomic in `dis` output" is not a specification. The professional rule is the boring one: if you have to squint at a sequence to decide whether it is one operation, it is not — reach for `threading.Lock`. ### The cost side, since an interviewer will push there A lock is cheaper than a bug, but a lock taken per element in a hot loop serializes exactly the parallelism the free-threaded build was bought for. The usual shape is to keep shared mutation rare: give each worker private state, do the bulk of the work unshared, and merge once at the end under one lock — or hand results to a single consumer through a `queue.Queue`. That way the per-object guarantee on built-ins covers the cheap operations, and your own lock covers the one place where an invariant genuinely spans several of them.

  • Does the same guarantee cover a class of your own that wraps a list?
    No. The internal locking protects the built-in's C implementation. A wrapper whose methods are written in Python executes bytecode, so another thread can be scheduled between your membership test and your append. Wrap the whole method body in a `threading.Lock` you own, or keep the wrapper's state private to one thread.
  • Which standard-library containers are documented thread-safe rather than merely observed to be?
    `queue.Queue` is built for cross-thread hand-off and documents that role, and `collections.deque` documents thread-safe appends and pops at both ends. Everything else you should treat as internally consistent per operation but with no published per-method atomicity contract, which is why compound logic gets an explicit lock.
  • Is it safe to iterate a shared dict while another thread writes to it?
    No. Changing a dict's size during iteration raises `RuntimeError`, and even where you get away with it the view is unspecified. Take a snapshot under a lock, or copy the container and iterate the copy, rather than relying on the per-operation guarantee to cover a whole loop.

The free-threaded build guarantees each drawer of the filing cabinet closes cleanly even when two clerks pull at once. It says nothing about two clerks who each read the same folder, edit it, and file their own version back.

saying these in an interview costs you the question

  • Says the free-threaded build made existing threaded code thread-safe
  • Claims concurrent appends can corrupt the list or crash the interpreter
  • Extends single-operation safety to read-modify-write sequences
  • Assumes a Python-level list subclass inherits the same guarantee
  • Thinks a lock is never needed once the GIL is gone
  • Iterates a shared dict while another thread mutates it

context

open as a page

Why does `totals[key] = totals[key] + 1` race across threads when each dict operation is indivisible?

level: middleimportance: must knowfreq 65%

basics

~20 s

That line is two dict operations, not one: a read, then a write, with a gap between them. Two threads can both read the same old value and both store it plus one, so an increment is lost. Lock both.

open as a page

How does `concurrent.futures.InterpreterPoolExecutor` differ from `ThreadPoolExecutor` for CPU-bound work?

level: middleimportance: must knowfreq 38%

basics

~20 s

Both run workers as OS threads in one process, but ThreadPoolExecutor workers share one interpreter and one GIL, so pure-Python CPU work serializes. Each InterpreterPoolExecutor worker gets its own interpreter and its own GIL, so CPU work runs genuinely in parallel.

open as a page

Your service runs python3.14t but sys._is_gil_enabled() returns True — why?

level: seniorimportance: must knowfreq 42%

basics

~20 s

Either the run asked for the lock with -X gil=1 or PYTHON_GIL=1, or a C extension without a free-threading declaration was imported and the interpreter turned the GIL back on for the whole process, warning as it did. Check the requested options first, then bisect the imports.

open as a page

How do you check whether CPython is running the free-threaded build?

level: juniorimportance: should knowfreq 28%

basics

~10 s

Call sys._is_gil_enabled(): it returns False only while the GIL is actually off. For the build itself, sysconfig.get_config_var('Py_GIL_DISABLED') is 1 on a free-threaded interpreter, and the binary is named python3.14t.

open as a page

What do the -X gil switch and PYTHON_GIL do on a free-threaded CPython build?

level: middleimportance: should knowfreq 33%

basics

~20 s

On an interpreter built without the GIL, -X gil=1 (or PYTHON_GIL=1) forces the lock back on for that run, and -X gil=0 pins it off even when an imported extension would otherwise re-enable it. The command-line switch wins over the environment variable.

open as a page

Which objects can cross a `concurrent.interpreters` queue, and which cannot?

level: middleimportance: should knowfreq 26%

basics

~20 s

Immutable scalars and tuples of them cross efficiently, and a memoryview genuinely shares its buffer. Dicts, lists and ordinary instances are copied, so the receiver gets an equivalent new object. Locks, sockets, generators, modules and closures raise NotShareableError.

open as a page

Why can `if rec_id not in seen: seen.add(rec_id)` let two threads emit the same catalogue record twice?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Because the membership test and the insert are two separate set operations. Both threads can test before either adds, both find the id absent, and both go on to emit the record. The set is fine; the claim was never exclusive.

open as a page

Why does every `concurrent.futures.InterpreterPoolExecutor` worker re-import your modules?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Each worker runs its own interpreter with its own sys.modules, so no import from the parent is visible and every worker re-executes the import graph. Caches and configuration are per worker, and some native extensions refuse to load at all.

open as a page

When is the free-threaded build's 5-10% single-thread cost worth paying?

level: principalimportance: should knowfreq 32%

basics

~20 s

Only when a real workload is CPU-bound in Python inside one process and cannot be split across processes cheaply, and when every C extension in the graph ships free-threading-capable builds. Otherwise the roughly 5-10% tax and the second wheel supply buy nothing.

open as a page

What does `concurrent.interpreters.create()` return, and where does that interpreter run?

level: juniorimportance: nice to knowfreq 18%

basics

~20 s

It returns an Interpreter object: an additional Python interpreter living inside the same OS process, not a new process. It has its own sys.modules, its own module globals and its own GIL, and you drive it with Interpreter.exec().

open as a page