A service stores per-request data such as the caller's identity in thread-local storage while handling requests on a pool of reusable worker threads. What can go wrong across requests, and how do you prevent it?
answer
- pool workers are immortal → entries outlive the task
- set without remove = next task inherits it
- bleed: unauthenticated request reads previous caller's identity
- remove(), not set(null); wrap at the task boundary in a finally
- fail closed on missing context; assert empty slots in tests
basics
~20 sPool threads never die, so a value left in a thread-local slot survives into the next request on that worker. That leaks memory and, worse, leaks data: a later request that forgets to set the identity silently inherits the previous caller's. Always clear in a finally block.
solid answer
~60 sThread-local entries live as long as the **thread**, and pooled workers are immortal by design. So anything a request sets and does not remove is still there when the worker picks up an unrelated request. Two distinct failures: - **Memory leak.** Entries accumulate — one per key, per worker — and each pins whatever it references, potentially a large object graph. With hundreds of workers this is unbounded growth that no request path explains. - **Cross-request data bleed.** This is the serious one. If a later request does not set the value (unauthenticated path, internal task, a code path that forgot), a read returns the *previous* request's value. Now you are authorizing, filtering, or logging under someone else's identity. It is a security bug, it is intermittent, and it depends on which worker the scheduler picked. Prevent it by never setting without a paired removal: a scoped wrapper that sets on entry and **removes** (not sets-to-null) in a finally, applied at the framework boundary so no handler can forget. Then verify: assert slots are empty at task boundaries in tests, and fail closed — treat missing context as an error, never as "reuse whatever is there".
code
text · 11 linesworker-3 t=0 request A (alice, authenticated)
CTX.set(user=alice, tenant=acme)
...handle...
return <-- no remove()
worker-3 t=1 request B (internal job, no auth)
CTX.get() -> user=alice, tenant=acme
query rows for CTX.tenant ==> acme's data returned to job B
Nothing throws. B's logs name alice. Reproduces only when B lands on
the same worker as A, i.e. rarely, and never on an idle test box.go deeper
Say that pool threads are reused so a value left behind is still there for the next task, and that you must clear it in a finally block.
Separate the two failures — memory retention and cross-request bleed — and explain why removal beats nulling and why the finally must wrap the whole task.
Own it as a security defect: boundary decorator, fail-closed reads, task-id stamping, heap-dump and log-correlation diagnosis, and tests that assert clean slots.
Argue the policy — identity and tenant are explicit parameters at trust boundaries, ambient context is for diagnostics only — and prefer scoped bindings or platform mechanisms that make the leak structurally impossible rather than relying on discipline.
## Why pooling turns a convenience into a hazard Thread-local storage keys state to the thread. That is fine when the thread's life and the work's life coincide — a thread created for one task, ending with it, taking its slots with it. Pools break that coincidence deliberately: a worker exists for the life of the process and executes an unbounded sequence of unrelated tasks. Its per-thread map persists across all of them. So every entry a task writes is, by default, part of the *next* task's initial state. Nothing clears it. The pool does not know your keys, and there is no per-task lifecycle to hook into unless you add one. ## Failure one: unbounded memory growth Each worker holds one entry per key ever set on it, and each entry pins its value — which may be a small string or may be a request object holding a payload, a connection, a whole parsed document, or a reference chain into a large cache. Multiply by workers. The result is heap growth that correlates with pool size and time rather than with concurrent load, so it looks like a slow leak with no obvious owner: heap dumps show live objects retained by *thread* references, not by anything in your request path. A related, nastier variant: an entry pins a value whose class was loaded by a module or plugin classloader. As long as the worker lives, that classloader cannot be collected, so an entire module's classes stay resident after unload/reload. That is the classic redeploy leak, and its root cause is exactly this lifetime mismatch. ## Failure two: cross-request data bleed (the serious one) The correctness failure matters more than the memory one. Consider the sequence: ``` worker-3: request A (authenticated as alice) -> set USER = alice ; ... ; return [no clear] worker-3: request B (health check, no auth) -> code reads USER -> gets alice ``` Request B never authenticated, but a read of the ambient identity returns a real, valid-looking user. Depending on what consumes it, the consequences range from mislabeled logs and metrics attributed to the wrong tenant, through audit records naming the wrong actor, to authorization decisions and data queries executed as someone else. The last is a genuine data-disclosure vulnerability. What makes it brutal to find: - **It is scheduler-dependent.** It only manifests when a request lands on a worker whose previous task left a value, so it reproduces at low rates and never on an idle test machine. - **It looks legitimate.** The value is well-formed, present, and of the right type — nothing validates that it belongs to *this* request. - **The setter and the victim are far apart.** The code that forgot to clear is in one module; the damage appears in another that merely reads ambient context. The same shape applies to every ambient value: tenant, locale, feature-flag overrides, trace/span ids (which corrupt traces by attaching new work to a finished span), transaction or database-session handles (which can join work to a stale transaction), and security-role sets. ## Prevention: pair every set with a removal The rule is mechanical: **never call set without a paired remove in a finally block**, and never leave it to individual handlers to remember. ``` withContext(ctx, body): CONTEXT.set(ctx) try: body() finally: CONTEXT.remove() // remove, not set(null) ``` Two details matter. **Remove, do not null out.** Setting a null/empty value still leaves an entry in the map, so it still consumes a slot and, in some implementations, still participates in the structures that keep the map alive. Removal is what releases the reference. Nulling also makes "present but empty" indistinguishable from "deliberately cleared" in diagnostics. **Apply it at the boundary.** Wrap the point where a task begins — the server's request filter, the pool's task decorator, the message consumer's loop — not in each handler. A convention that every handler must remember is a convention that will be violated, and the violation is silent. One decorator around task execution guarantees the invariant regardless of what the handler does, including when it throws. ## Defence in depth **Fail closed on missing context.** If code needs an identity and the slot is empty, that must be a loud error, not a fallback. This flips the failure from silent-and-wrong to obvious-and-safe, and it removes the incentive to "leave the last value around just in case". **Stamp and validate.** Store a task/request identifier alongside the context and have readers assert it matches the task they are running. A mismatch means stale context and should throw. Cheap, and it converts an invisible bleed into an immediate exception naming both requests. **Assert cleanliness at boundaries in tests.** After each task in an integration test, assert the worker's known slots are empty. This catches the missing-finally the moment it is introduced rather than months later in production. **Keep values small and immutable.** A leaked immutable string is a bounded, harmless leak; a leaked mutable request object retains a graph and can be mutated by whoever finds it. **Prefer scoped bindings where the platform offers them.** Mechanisms that bind a value only for the dynamic extent of a call and unbind automatically on exit make the leak structurally impossible instead of relying on discipline. Where such a mechanism exists, it is strictly better than manual set/remove. ## Diagnosing an existing case If you suspect bleed: take a heap dump and look for objects retained via thread references from pool workers — that names both the leaking key and its value type. For the correctness symptom, correlate logs where the identity or tenant on a request disagrees with its authenticated principal, and check whether affected requests share a worker thread name with an earlier request from the other party. Thread names are the thread you pull.
- Why is removing the entry better than setting it to null or an empty value?Setting null leaves the entry in the thread's map, so the slot is still occupied and the map's own structures still hold on to the key — you have cleared the payload but not the entry. Removal is what actually releases the reference and lets the value be collected. It also preserves a useful distinction in diagnostics: absent means 'no context for this task', while present-but-null is ambiguous and invites readers to treat it as a legitimate empty value.
- You suspect a service is leaking identity between requests. How do you confirm it?Two lines of evidence. From a heap dump, look for values retained through thread references from pool workers — that names both the leaking key and the retained object graph. From logs, look for requests whose recorded identity or tenant disagrees with the authenticated principal on that request, then check whether each such request ran on the same worker thread name as an earlier request from the party whose identity appeared. Thread names are the correlating key, which is why naming pool threads meaningfully pays off here.
- Does the problem disappear on a runtime where each task gets its own cheap thread?The cross-task bleed does, because the thread dies with the task and takes its slots with it, so there is no reuse to inherit from. But two things replace it: memory now scales with the number of concurrent threads rather than pool size, so large per-thread values become expensive, and propagation gets harder because context still does not follow work across thread boundaries. It removes one failure mode rather than making ambient context safe in general.
- How do you stop individual handlers from forgetting the cleanup?Take the decision away from them. Wrap the point where a task begins — request filter, pool task decorator, consumer loop — so the set/finally-remove pair happens exactly once, outside any handler's control, and runs even when the handler throws. Then enforce it: forbid raw set calls in review or lint, expose only the scoped wrapper, and assert in integration tests that known slots are empty after each task.
A hot-desk office where nobody clears the desk between shifts. The next person sits down, finds a signed authorization form from whoever sat there this morning, and — because it looks perfectly official — uses it.
saying these in an interview costs you the question
- Assuming the pool clears thread-local state between tasks
- Clearing by setting null or an empty object rather than removing the entry
- Putting the cleanup in the handler's own code instead of at the task boundary, so an early return or a throw skips it
- Treating this purely as a memory leak and missing that it is a data-disclosure bug
- Falling back to whatever value is present when context is missing, instead of failing closed
- Believing the bug is rare because it does not reproduce locally — it is scheduler-dependent, not rare