Python worker threads all rebuild the same cache on a 45-second cold start — what race is this and how do you fix it?
answer
- Every worker sees an empty cache
- The window is as wide as the build
- Check-then-act with a very slow act
- Lock, then check the same condition again
- Publish a finished object, never a partial one
basics
~20 sIt is a check-then-act race: every thread finds the cache empty during the 45-second build window and starts its own build. Serialize the build behind a lock and re-check inside it, or build once eagerly before the pool starts accepting work.
solid answer
~50 sThe code almost certainly reads `if _cache is None: _cache = build()`, and the gap between the test and the assignment is 45 seconds wide, so every thread that arrives in that window passes the test and builds. That is a check-then-act race producing a **cache stampede**: four workers, four builds, four times the memory and CPU. There is usually a second bug alongside it — filling a shared dict in place lets readers see a half-populated cache, and can raise `RuntimeError` if one iterates while another inserts. Fix both: build under a `threading.Lock`, re-check the cache after acquiring it, and publish by rebinding the module-level name to a fully built new object so readers see either the old cache or the complete new one. If callers cannot wait 45 seconds, build eagerly at startup or let one builder thread work while the rest serve stale data and wait on a `threading.Event`.
code
python · 28 linesimport threading
import time
_index = None
_index_lock = threading.Lock()
builds = 0
def build_index():
global builds
builds += 1
time.sleep(0.2) # stands in for a slow build
return {"ticket": "triage"}
def get_index():
global _index
if _index is None:
with _index_lock:
if _index is None: # re-check inside the lock
_index = build_index() # publish one finished object
return _index
threads = [threading.Thread(target=get_index) for _ in range(8)]
for t in threads:
t.start()
for t in threads:
t.join()
print("builds:", builds)go deeper
Recognise the shape: a test for an empty cache followed by slow work is a check-then-act race, and the GIL does not prevent it. Knowing that a lock belongs around the whole build is enough at this level.
Explain the double-checked pattern and why the inner re-check matters, and describe publishing a finished object by rebinding the name rather than filling a shared dict in place while readers can see it.
Show the diagnosis path from logs, latency and memory, and argue the tradeoff: block callers, serve stale, or build eagerly behind a readiness check. Mention caller deadlines and Lock.acquire(timeout=...) rather than unbounded waiting.
Own the cold-start policy across services: where warm-up belongs in the deploy lifecycle, whether stale responses are contractually acceptable, and how much memory headroom the fleet needs if a stampede ever recurs under a partial outage.
### The symptom and the shape of the code A triage service starts, four worker threads begin pulling tickets, and the first four tickets each take 45 seconds instead of one. Memory peaks at four times the expected size. Occasionally a worker classifies a ticket against a cache that is missing half its entries, and once in a while a thread dies with `RuntimeError: dictionary changed size during iteration`. The code behind this is almost always some variant of: ```python _index = None def get_index(): global _index if _index is None: # check _index = build_index() # ...45 seconds... then act return _index ``` ### Why it races The test and the assignment are separate operations, and everything between them is a window. Here the window is the entire build: 45 seconds. Every thread that calls `get_index()` during those 45 seconds sees `None`, because nothing was published yet, and starts its own build. The GIL does not help — it serializes bytecode, and the threads are not fighting over a single instruction, they are all sitting inside a long-running function. This is the same family as `if key not in d: d[key] = ...`, just with a window wide enough that the race is not occasional but guaranteed. The variant that fills a shared structure in place is worse: ```python _index = {} def get_index(): if not _index: for k, v in load_rows(): _index[k] = v # readers can see this half-done return _index ``` Now a second thread finds `_index` truthy after the very first insert and returns a cache with one entry. Correctness bugs of that kind are far harder to diagnose than the duplicated work, because the service does not slow down — it just produces wrong answers for a few seconds after every restart. ### Diagnosing it Log the build, with the thread name, at both ends: ```python import logging, threading logging.info("building index on %s", threading.current_thread().name) ``` If the same start line appears four times per process launch, you have your answer. Timing the first N requests per restart shows the same thing from the outside: a burst of very slow requests at exactly the width of the build, on every worker, once per deploy. Memory tells the third part of the story — four concurrent builds hold four intermediate copies. ### Fix 1: serialize and re-check ```python import threading _index = None _index_lock = threading.Lock() def get_index(): global _index if _index is None: # fast path, no lock with _index_lock: if _index is None: # re-check: someone may have built it _index = build_index() return _index ``` The inner re-check is the whole point: a thread that waited on the lock must not build again. The outer check is only an optimisation for the common case where the cache already exists. Note what the callers pay — the second, third and fourth threads block for the remainder of the 45 seconds. That is usually right, since they had no useful work without the cache, but it is a decision, not a default. ### Fix 2: publish by rebinding, never by mutating Build into a fresh local dict and assign the finished object to the module-level name in one store. Readers hold a reference to whichever complete object existed when they looked, and no reader ever observes a partial structure: ```python def refresh(): global _index new_index = {k: v for k, v in load_rows()} _index = new_index # single publication step ``` Readers should copy the reference into a local variable once (`index = _index`) and use the local, so a refresh partway through their work cannot swap the object under them. ### Fix 3: build eagerly, or serve stale while one thread builds If a 45-second stall on the first request is unacceptable, move the cost: build the cache during startup, before the pool starts accepting work, so the process is either not ready or fully warm. Readiness checks exist precisely for this. Where stale data is acceptable, keep the previous cache in place, let a single builder thread produce the replacement, and have readers continue on the old object until the new one is published — a `threading.Event` (or simply a "build in progress" flag guarded by the lock) prevents a second builder from starting. Where callers have their own deadline — say five seconds — do not let them block for 45. Use `lock.acquire(timeout=...)` and, on timeout, serve stale data or fail fast with a clear message. Blocking a caller past its own deadline converts a throughput problem into a queue of doomed requests. ### What about `functools.lru_cache`? It is a reasonable memoisation tool and its internal bookkeeping is thread-safe, but it does **not** collapse concurrent misses: several threads calling it for the same missing key can each run the wrapped function. It is therefore not a stampede fix. If you want dedupe, you still need the lock-and-re-check pattern, or a single owner thread fed through a `queue.Queue` that returns the built object to everyone who asked. ### The general lesson Check-then-act races are ranked by the width of their window. A two-bytecode window makes a flaky test; a 45-second window makes a guaranteed production incident. When you see a test whose corresponding action is expensive, assume every concurrent caller will take the branch, and design for that.
- Why is the inner re-check inside the lock not redundant?Without it, every thread that queued on the lock rebuilds as soon as it acquires — you have serialized the stampede instead of removing it. The outer check is only a lock-free fast path for the warm case; the inner check is the one that guarantees a single build. Skipping the outer check is safe but costs an uncontended lock acquisition on every call.
- How would you confirm from production evidence that more than one thread built the cache?Log the start of the build with `threading.current_thread().name` and count the lines per process start; four starts per restart is conclusive. Corroborate with the latency profile — a burst of requests slow by exactly the build duration after each deploy — and with peak memory, which multiplies by the number of concurrent builders because each holds its own intermediate structure.
- When is blocking every caller for the full build the wrong choice?When callers have their own deadline shorter than the build, or when stale data is acceptable. Blocking then converts one slow start into a backlog of requests that will time out anyway. Use `Lock.acquire(timeout=...)` and serve the previous cache or fail fast, and prefer building during startup behind a readiness check so the process never accepts traffic it cannot serve.
Four cooks each look in the empty pot, and each starts a fresh 45-minute stew. One pot, one cook, and a note on the lid saying who is already cooking is all that was missing.
saying these in an interview costs you the question
- Says the GIL prevents two threads from building at once
- Adds a lock but omits the re-check inside it
- Fills the shared cache in place while readers read it
- Believes `functools.lru_cache` deduplicates concurrent misses
- Holds the lock for 45 seconds when callers time out in five
- Calls it a deadlock rather than a check-then-act race