skip to content

Why does CPython 3.14's free-threaded build run single-threaded work slower than the default build?

level: seniorimportance: should knowfreq 26%

answer

  1. Removing a lock is not free
  2. Something must still protect each object
  3. Counts and containers now guard themselves
  4. A single-digit percent tax, quoted in the PEP
  5. Pays off only with parallel CPU-bound Python

basics

~20 s

Without the GIL, every reference count update and every container access has to be safe on its own, so the interpreter pays for synchronisation the default build got for free. PEP 779 puts the single-threaded overhead in 3.14 at roughly five to ten percent.

solid answer

~50 s

The GIL is not only a scalability limit; it is also a very cheap correctness guarantee. Because only one thread ran bytecode at a time, reference counts could be plain non-atomic increments and built-in containers needed no locking. A free-threaded build has to replace all of that: atomic or deferred reference counting, immortal objects for things like `None` and small ints, per-object locking on mutable containers, and a different allocator strategy so reads can proceed without locks. That work shows up as single-threaded overhead — around **5–10%** in **3.14**, down from roughly a third in the first experimental build in **3.13**. 3.14 also re-enabled bytecode specialization in free-threaded builds, which recovered much of the gap. It is a separate build and a separate ABI, so extensions must be rebuilt, and one that does not declare support makes the runtime switch the GIL back on.

code

python · 5 lines
python
import sys
import sysconfig

print("free-threaded build:", bool(sysconfig.get_config_var("Py_GIL_DISABLED")))
print("GIL currently enabled:", sys._is_gil_enabled())

go deeper

for a junior

Be ready to say that the free-threaded build is a separate interpreter that lets threads run Python code at the same time, and that this costs a little speed when only one thread is running.

for a middle

Explain the mechanism: reference counts and container mutations were safe for free under one global lock and now need atomics, immortal objects and per-object locking, which is where the single-threaded percentage goes.

for a senior

Show you would characterise the workload first, check every compiled dependency for free-threading support, watch for the runtime silently re-enabling the GIL, and compare against a plain multi-process deployment before adopting.

for a principal

Own the adoption question: what evidence justifies a second interpreter build in your estimate, how you would stage it across services, the ecosystem risk of extensions that lag, and when more processes remain the cheaper answer.

## What the GIL was buying It is easy to describe the global interpreter lock only as a limitation: one thread executing bytecode at a time, so CPU-bound threads do not scale. The half that matters for this question is what it *bought*. Because only one thread ran interpreter code, an enormous amount of internal machinery could be written as if it were single-threaded: - **Reference counts** could be ordinary increments and decrements. No atomics, no memory fences, no cache-line ping-pong between cores. - **Built-in containers** needed no locks. Appending to a list or updating a dict was safe because nobody else was running. - **Interpreter-internal caches** — type version tags, method caches, the inline caches behind bytecode specialization — could be read and written without coordination. - **The allocator and the cycle collector** could assume exclusive access to their structures. Removing the GIL means paying, explicitly, for every one of those guarantees. ## What the free-threaded build has to do instead Reference counting is the big one, because it touches essentially every operation. Making counts atomic on every object would be ruinous, so the design mixes strategies: **immortal objects** whose counts are never touched at all (small ints, `None`, `True`, `False`, interned strings, type objects), **deferred reference counting** for objects the interpreter manipulates constantly, and biased or thread-local schemes so that the common case — an object touched only by the thread that created it — avoids contention. Objects genuinely shared between threads fall back to atomic operations, which are cheap in isolation and expensive when several cores fight over the same cache line. Mutable containers get **per-object locks**, so a list append is a lock acquisition rather than nothing. The allocator is restructured so lookups can proceed without taking a lock. The cycle collector needs a way to stop the world across real threads. Add it all up and a single-threaded program does strictly more work than it would under the GIL. That is the answer to the question in one sentence: **the free-threaded build has not removed work, it has moved the cost of thread safety from one big lock onto every object that needs protecting.** ## The numbers, and the version story The first free-threaded build shipped as an **experimental** option in **3.13**, and its single-threaded overhead was large — reported in the region of a third — largely because bytecode specialization had to be disabled in that build; the inline caches were not yet safe to mutate concurrently. In **3.14**, free-threading became **officially supported** (PEP 779) rather than experimental, and the overhead fell to roughly **5–10%**. A major part of that recovery was re-enabling the specializing adaptive interpreter in free-threaded builds, so hot code once again benefits from type-specialized instructions. The number is still a real cost, and it is still per release: quote it as "single digits to low double digits on 3.14" and say you would measure your own workload, not as a fixed constant. ## Everything else you pay Single-threaded slowdown is only the headline cost. **It is a separate build with a separate ABI.** You install a distinct interpreter, and every compiled extension in your dependency set must have been rebuilt for it. An extension that does not declare free-threading support causes the runtime to switch the GIL back on at import (with a warning) unless you explicitly override that — at which point you have a free-threaded build running with a GIL, paying the overhead and getting none of the benefit. **Thread-safety bugs that the GIL hid become real.** The GIL never made Python code race-free, but it made many races so unlikely they never fired. Removing it changes those odds. **Tooling lags.** Profilers, debuggers and native libraries need their own support. ## Making the call The deciding question is never "is the free-threaded build faster?" — it is "does this workload have parallel CPU-bound work in pure Python that the GIL is currently serialising?" Take a document-conversion queue as the concrete case. If its workers spend their time waiting on object storage and on an external converter process, the GIL is not the bottleneck: threads already overlap fine across blocking calls, and switching builds buys a slowdown and an ecosystem risk for nothing. If instead the conversion itself is pure-Python CPU work — parsing, transforming, laying out — running in many threads inside one process, then the GIL is exactly what is capping throughput, and the free-threaded build trades that 5–10% single-thread tax for scaling across cores. Weigh it against the alternative that needs no new build at all: separate processes, which cost memory and serialisation but keep every extension working. Whichever way you lean, decide it with a measurement of your own workload on both builds, at your real concurrency, over a long enough run to include steady-state behaviour — not with a benchmark someone published.

  • Your free-threaded interpreter reports that the GIL is enabled anyway. What happened?
    Something imported a compiled extension that does not declare support for running without the GIL, so the runtime switched it back on and warned. You are now paying free-threading's per-object overhead and getting none of its parallelism. The fix is to find the offending extension, upgrade or replace it, and re-measure — not to force the override and hope.
  • What changed between 3.13 and 3.14 for free-threading?
    3.13's build was experimental and ran single-threaded code roughly a third slower, partly because bytecode specialization was turned off in it. 3.14 makes free-threading officially supported under PEP 779, re-enables the specializing interpreter, and brings single-threaded overhead down to around 5–10%. It is still a separate build, not the default interpreter.
  • How would you decide between the free-threaded build and simply running more processes?
    By what the work is. Processes need no new interpreter, keep every extension working and isolate crashes, at the cost of memory and of serialising data between them. The free-threaded build wins when threads share a large in-memory structure that would be expensive to copy per process, and when the parallel work really is pure-Python CPU time.

saying these in an interview costs you the question

  • Says removing the GIL is free performance
  • Expects a free-threaded build to speed up I/O-bound workers
  • Believes the free-threaded build is the default in 3.14
  • Assumes existing compiled extensions work unchanged
  • Still calls free-threading experimental in 3.14
  • Thinks the GIL was the only thing protecting Python-level data structures

context