Why can a value parked in per-thread storage leak when the thread comes from a pool rather than ending with the request?
answer
- scoped to the thread, not the task
- a pooled worker never ends
- the last task's value stays attached
- ceiling is pool size x value size
- removal belongs at the task boundary
basics
~20 sPer-thread storage is scoped to the thread, not to the task. A pooled thread never ends, so a value set during one request stays attached to it - and reachable - until something explicitly removes it or a later task overwrites it.
solid answer
~50 sPer-thread storage attaches a value to the thread object, and the thread is the retainer. When a thread is created per request and ends with it, the attachment is effectively request-scoped and cleanup is automatic. A pooled worker inverts that: the thread outlives every task it runs, so whatever the last task left behind stays reachable from a live thread for as long as the pool exists. The leak is bounded - roughly pool size multiplied by the retained size of one value - but that bound can be large when the value is a request payload, an accumulated buffer or a context object that transitively holds a session. It is also a correctness hazard, because the next unrelated task on that thread can read the previous task's value and treat it as its own. The rule is that whatever sets a per-thread value owns removing it, on an unconditional cleanup path at the task boundary.
code
pseudocode · 11 linescontext = thread_local_slot()
function worker_loop(queue):
while true:
task = queue.take() // same thread, next task
context.set(task.payload) // large, per-request
handle(task)
// missing: context.remove()
// after the queue drains, every idle worker still holds
// the payload of the last task it ran: pool_size x payloadgo deeper
Recall that per-thread storage belongs to the thread, so it is only request-scoped when the thread itself is. A pooled thread is reused, so anything left in it stays.
Explain the ceiling - pool size times slot count times value size - and why removal must sit at the task boundary rather than inside each handler.
Diagnose it from the signature: memory that steps up as the pool warms and then plateaus without receding, plus one retained object per live worker, and name the cross-task leakage risk.
Decide the policy: ambient per-thread context on pooled executors needs a framework-level set-and-clear contract, or the context is passed explicitly, because correctness here cannot depend on every handler author remembering.
## The scope is the thread, and the thread is the retainer Per-thread storage gives each thread its own slot for a value, so code deep in a call chain can reach request context without threading a parameter through every frame. Mechanically, the value hangs off the thread object itself. That makes the **thread** the thing holding your data alive, and the lifetime of the data is the lifetime of the thread, not the lifetime of the work. When a thread is created for one request and ends when the request ends, those two lifetimes coincide, so nothing goes wrong and nobody thinks about it. Pooling breaks the coincidence deliberately: the whole point of a pool is that thread creation is expensive, so workers are created once and reused for the life of the process. | | thread per request | pooled worker | |---|---|---| | thread lifetime | one task | the life of the pool | | value lifetime if never removed | ends with the task | ends with the process | | who cleans up | thread termination | only an explicit removal | | failure if you forget | none visible | retention plus cross-task leakage | ## How big is the leak, really? This shape does not grow without bound, and saying otherwise in an interview is a mistake. Each thread holds at most one value per slot, so the ceiling is roughly: > pool size x number of slots x retained size of one value With a pool of 200 workers and a context object that transitively retains a 2 MB request payload, that ceiling is about 400 MB of memory that is never used again and never returns. It climbs as traffic touches idle workers for the first time, then flattens - which is precisely what makes it confusing to diagnose. It does not look like a classic runaway leak; it looks like a service whose baseline is inexplicably high and never comes back down after a traffic spike. The other half of the cost is invisible in a memory graph: - **Cross-task leakage of data.** The next task on that worker sees the previous task's value if it reads before it writes. When the value is a user, tenant or authorisation context, this is a security bug, not an efficiency one. - **Stale correlation.** Diagnostic context carried this way attaches the previous request's identifiers to the current one, which quietly corrupts the trail you would use to debug anything else. - **Growth that outlives the work item.** A value that accumulates - a list appended to per operation - keeps accumulating across unrelated tasks on the same worker, because nothing resets it. ## The discipline 1. **Whoever sets it removes it,** at the same boundary, on a cleanup path that runs whether the task succeeded or failed. Removing only on success guarantees the failure path leaks, and failures cluster. 2. **Put the removal in the framework layer, not in business code.** The task boundary - the loop that takes work from the queue and runs it - is one place; every handler is hundreds of places, each of which can forget. 3. **Set on entry as well as clearing on exit.** Defending both ends means a missed removal degrades into wasted memory rather than into one request reading another's context. 4. **Prefer passing the context explicitly** where the call depth makes it practical. Per-thread storage is an ambient-state mechanism, and ambient state and pooled execution are a known-bad combination; an explicit parameter has no lifetime question at all. 5. **Be careful with work that hops threads.** If a task starts on one worker and continues on another - because it suspended, or handed off to a different pool - the value stays attached to the first thread and is absent on the second. That both leaks and silently loses context, which is why frameworks that migrate work usually provide their own context propagation instead. ## Why it is hard to see There is no single growing container to point at, which is what makes this shape distinctive among retention bugs. The retained values are spread one per worker, each reachable from a different live thread, and each perfectly legitimate-looking. The tell is the **shape of the number**: memory that steps up as the pool warms, plateaus at a level proportional to pool size, and never falls afterwards - plus a per-worker distribution where every live thread holds one object of the same type. Shrinking the pool shrinks the plateau proportionally, which is a cheap confirmation and a terrible fix.
- Is this leak unbounded?No, and claiming so is a common error. Each thread holds one value per slot, so the ceiling is pool size times slot count times the retained size of one value. It climbs as workers are used for the first time and then plateaus. The danger is the size of that plateau and the fact that it never recedes, not unbounded growth.
- Beyond memory, what else goes wrong if the value is never removed?The next unrelated task on that worker can read the previous task's value before writing its own. If the value carries user, tenant or authorisation context, one request acts on another's identity - a correctness and security failure that usually matters far more than the retained bytes.
- Why is removing the value only after a successful task insufficient?Because failures are exactly when the value is most likely to be stale or oversized, and they cluster during incidents. Cleanup must run on an unconditional path at the task boundary, so that an exception on a handler leaves the worker as clean as a success does.
- What happens if a task migrates to a different thread part-way through?The value stays attached to the original thread and is simply absent on the new one. That leaks the old value and loses the context at once, which is why runtimes that migrate work provide explicit context propagation rather than relying on per-thread storage.
saying these in an interview costs you the question
- Says clearing is pointless because the thread ends anyway
- Claims the leak grows without bound with traffic
- Removes the value only on the successful path
- Assumes each task gets a fresh empty slot automatically
- Treats reading a leftover value as harmless staleness, not cross-request leakage
- Sets the value in the handler but removes it in the caller