Why does a value set on a threading.local() inside one ThreadPoolExecutor task reappear in a later task?
answer
- The pool exists to not recreate threads
- Two lifetimes that do not line up
- Keyed by the worker, scoped to the job
- Nothing runs between two work items
- initializer fires per worker, not per task
basics
~20 sBecause concurrent.futures.ThreadPoolExecutor reuses its worker threads. The slot is keyed by the thread, not by the task, and nothing clears it between work items, so the next task on that worker inherits whatever the previous one left.
solid answer
~50 sA pool exists to avoid recreating threads, so its workers pull work item after work item off one queue and stay alive until the executor is shut down. A `threading.local` slot is keyed by the thread, so a value written during one task is still there when the next task runs on that same worker — and the code usually does not crash, it just proceeds with the previous task's identity. The executor's `initializer` argument does **not** fix this: it runs once per worker at startup, not per task. Real fixes are to clear the slot in a `finally` (never only on the happy path), to set it unconditionally at task entry, or — better — to keep per-task state in a `contextvars.ContextVar` and submit `contextvars.copy_context()`'s `Context.run` so each task runs in its own copy.
code
python · 13 linesimport threading
from concurrent.futures import ThreadPoolExecutor
slot = threading.local()
def handle(device_id):
previous = getattr(slot, "device_id", None)
slot.device_id = device_id
return f"{device_id}: worker still held {previous!r}"
with ThreadPoolExecutor(max_workers=1) as pool:
for line in pool.map(handle, ["sensor-1", "sensor-2", "sensor-3"]):
print(line)go deeper
Recall that a pool reuses its threads instead of creating one per job, and that thread-local storage is tied to the thread. Those two facts together are the whole answer at this level.
Explain the mechanics precisely: workers loop over a queue and outlive individual work items, nothing clears storage between them, and cleanup must sit in a finally so an exception path cannot skip it.
Show the diagnosis, not just the cause: the symptom is plausible wrong output rather than a crash, one worker reproduces it deterministically while many workers hide it, and an entry assertion converts it into a loud failure.
Own the policy: decide whether per-request identity may live in ambient storage at all, and if it does, make the hand-off mechanism standard across services rather than leaving each team to remember its own cleanup.
This is the single most common way `threading.local` goes wrong in a real service, and the cause is a mismatch between two lifetimes: the slot is keyed by **thread**, and the work is scoped to a **task**. ## What ThreadPoolExecutor actually does with threads `concurrent.futures.ThreadPoolExecutor` is a queue plus a set of long-lived worker threads. When you call `submit`, the callable and its arguments are wrapped in a work item and pushed onto an internal queue; the executor starts a *new* worker thread only when there is no idle worker and the count is still below `max_workers`. Once created, a worker loops forever: pull an item, run it, pull the next item. It does not exit between items, and it is not recreated per task — that reuse is the entire point of a pool, because thread creation is the cost the pool exists to amortize. Workers stay alive until the executor is shut down (its `shutdown` method, or leaving its `with` block). So the object you set an attribute on lives across tasks, and so does the thread that keys the storage. Nothing in the executor clears anything between work items — it has no idea your callable stashed something on a thread-local. Task 2 on that worker therefore starts with everything task 1 left behind. ## Why this is worse than an ordinary bug The value that leaks is usually *identity-shaped*: which device, tenant, request or user this unit of work is for. When the leak happens, the work does not crash — it succeeds with the wrong context. A telemetry collector that stores the current device's locale-dependent number format in a thread-local will happily format the next device's readings with the previous device's format, and the output is plausible. That class of bug has a habit of surviving several release trains before anyone believes the numbers are wrong. It is also intermittent in exactly the wrong way. With `max_workers=1` it reproduces on every second task; with a large pool it appears only when a worker happens to be reused for a task that forgets to set the slot. Raising `max_workers` therefore makes it *look* fixed while making it harder to reproduce. ## The fix that does not work The most common wrong fix is the executor's `initializer` argument. It runs once per worker thread, at that worker's startup — not once per task. It is the right hook for building a per-thread resource (a connection, a buffer) and the wrong hook for resetting per-task state, because by the time task 2 runs, the initializer has long since finished. ## The fixes that do work * **Clear it in a `finally`.** Whatever sets the slot must unset it on the way out — `finally: slot.__dict__.pop("device_id", None)` or `del` guarded against the exception path. Doing the cleanup at the end of the happy path only is the classic half-fix: the first exception leaves the value behind. * **Set it unconditionally at task entry.** If every task writes the slot before reading it, staleness cannot be observed. This is fragile — one code path that forgets brings the bug back — but it is cheap and pairs well with an assertion that the slot is unset on entry, which turns the silent bug into a loud one in tests. * **Wrap the callable.** Submit a wrapper that sets up and tears down around the real function, so no individual task author has to remember. * **Stop using a thread-local for task state.** This is the real answer. Per-task ambient state belongs in `contextvars.ContextVar`, and the way to carry it across the pool boundary is to snapshot the caller's context with `contextvars.copy_context()` and submit `Context.run`, which runs the callable in a *copy* — mutations do not escape back, and there is no residue on the worker for the next task. * **Or pass it explicitly** as an argument to the submitted callable, which is the option that makes the dependency visible in the signature and needs no cleanup at all. ## How to talk about it in an interview Do not describe it as a race, and do not blame the GIL. Nothing here is concurrent: the two tasks are strictly sequential on that one worker. The bug is that a per-thread store was used for per-task data, and a pool deliberately makes threads outlive tasks. Say that, then name the cleanup and the `contextvars` alternative, and mention that you would add an entry assertion so the failure is loud.
- Does the initializer argument of concurrent.futures.ThreadPoolExecutor solve this?No. It runs once per worker thread when that worker starts, not once per submitted callable, so it is the right hook for building a per-thread resource and the wrong one for resetting per-task state. By the time the second task runs on a worker, the initializer finished long ago.
- Would raising max_workers make the problem go away?No — it makes it intermittent, which is worse. With one worker every second task sees the residue and the bug is obvious; with many workers it only surfaces when a task lands on a worker that a previous task dirtied, so it reproduces rarely and looks like corrupted data rather than a lifetime bug.
- How would you catch this in a test rather than in production?Assert at task entry that the slot is unset, so a leaked value raises instead of being used. Beyond that, submit a probe task after a batch that fails if any slot is still populated, and run the suite with max_workers=1 so a leak reproduces deterministically.
The desk is cleared when a worker goes home, not between the jobs on it — and a pool worker never goes home.
saying these in an interview costs you the question
- Blames the GIL or a race between the two tasks
- Thinks each submit runs on a freshly created thread
- Uses initializer expecting it to run per task
- Clears the slot only on the happy path, not in finally
- Says raising max_workers fixes it
- Believes the executor resets state between work items