skip to content

OpenTelemetry keeps a "current context" in an implicit, scope-based store within a process. What exactly breaks when work is handed to a thread pool, a callback, or a reactive/coroutine pipeline, and how is it fixed?

level: seniorimportance: must knowfreq 58%

answer

  1. Context immutable; "current" is thread-local/async-local
  2. Scope token: close on same thread, reverse order
  3. Loss → orphan root spans; leak → wrong parent, plausible-looking
  4. Capture at submit, wrap the task, or pass explicitly
  5. Reactive/coroutines need the framework's context, not thread-locals

basics

~20 s

The implicit store is per-thread (or per-async-task), so work that moves to another thread sees an empty or stale context. Spans created there become new roots or attach to the wrong parent. Fix by capturing the Context at hand-off and re-attaching it inside the task, or by passing context explicitly.

solid answer

~60 s

OpenTelemetry's Context is immutable and passed explicitly in the API; the *current* context is a convenience stored in a thread-local (async-local in Node, contextvars in Python, a coroutine context element in Kotlin). Attaching returns a scope token that must be closed on the same thread, in reverse order. Two failure modes follow. **Loss**: submitting a task to an executor runs it on a thread whose local store is empty, so a span started there has no parent and becomes a new trace root — or worse, has the *previous* request's context if a pooled thread never closed its scope, silently attaching spans to an unrelated trace. **Leak**: a scope that is not closed poisons every later task on that thread. The fixes are to capture the current Context at submission time and wrap the runnable so it attaches and closes around the body, use context-propagating executor wrappers, or pass the Context explicitly and start spans with an explicit parent. Reactive and coroutine pipelines hop schedulers between operators, so context must ride the framework's own context, not a thread-local.

code

text · 11 lines
text
LOSS
  T1: [span A current] submit(task) --------------
  T2:                                  [current = empty] start(B) -> B is a NEW ROOT

LEAK
  T2: attach(ctx_req1) ... handler returns without close()
  T2: (thread reused) start(C) for req2 -> parent = req1's span, wrong trace

FIX
  T1: ctx = current(); submit(wrap(ctx, task))
  T2: scope = ctx.makeCurrent(); try { start(B) } finally { scope.close() }

go deeper

for a junior

Say the current context lives per thread, so work moved to another thread loses it and spans there become new roots; the fix is to carry the context to the task.

for a middle

Distinguish loss from leak, describe capture-and-wrap and the scope-closing discipline, and know which store your language uses.

for a senior

Diagnose from symptoms (orphan roots versus impossibly long roots with mixed tenants), cover reactive and coroutine propagation, and know the limits of agent instrumentation.

for a principal

Set the codebase convention — explicit context at asynchronous boundaries, wrapped executors provided centrally rather than per call site — and treat context correctness as a testable property of the concurrency layer.

## Explicit context and the implicit convenience OpenTelemetry's `Context` is an immutable map. Every API that needs it can take it explicitly — `tracer.start(name, parentContext)` — and in a purely explicit style nothing can be lost, because context is just an argument threaded through your code. That is verbose, so the API also offers an implicit "current context", held in a per-execution-unit store: - JVM and .NET: a thread-local (`ThreadLocal`, `AsyncLocal`) - Node.js: `AsyncLocalStorage`, which follows async continuations - Python: `contextvars`, which follow `await` boundaries - Go: no implicit store at all — `context.Context` is an explicit first parameter by convention - Kotlin coroutines: a coroutine context element that travels with the coroutine Attaching looks like `scope = context.makeCurrent()` and the scope must be closed. The scope is not the span: closing the scope restores the previous current context, while ending the span records it. They are independent lifetimes and conflating them causes half of all context bugs. ## Failure mode 1: loss ``` request thread: ctx = current (span A is current) executor.submit(task) pool thread: current = <empty> (different thread-local) span B = tracer.start("work") -> parent = none ``` Span B has no parent, so the SDK treats it as a trace root: a brand-new trace id. Symptom: the backend shows many short one-span traces that ought to have been children, and the parent trace looks truncated exactly where the asynchronous hand-off occurs. With a parent-based sampler this also re-rolls the sampling decision, so the orphan may be sampled when the parent was not, or vice versa. The same loss happens with raw thread creation, timer/scheduler callbacks, custom queue-and-worker designs, and any library callback invoked from an internal pool. ## Failure mode 2: leak (the worse one) ``` pool thread handling request 1: attach(ctx1) ... forgot to close ... pool thread reused for request 2: current is STILL ctx1 span for request 2 -> parent = request 1's span ``` Now spans are attached to a completely unrelated trace. This is worse than loss because the data looks plausible: the trace exists, is well formed, and is wrong. It also spreads — every subsequent task on that thread inherits the stale context until something overwrites it. Symptoms are traces with impossible durations (a root that appears to last hours), spans from different tenants in one trace, and parent/child pairs whose timestamps do not nest. The discipline that prevents it is mechanical: always close the scope, on the same thread that opened it, in reverse order of opening, with the language's scoped-resource construct so an exception cannot skip it. Never attach in one method and close in another. ## The fixes **Capture and wrap.** At the hand-off point, capture the current Context (an immutable value, cheap to hold) and wrap the task so it attaches and closes around the body: ``` ctx = current_context() executor.submit(() -> { scope = ctx.makeCurrent(); try { work(); } finally { scope.close(); } }) ``` SDKs ship helpers that do this: a context-propagating wrapper around an executor, or a `Context.wrap(runnable)` helper. Wrapping the *executor* once at construction is more reliable than remembering at every call site. **Pass explicitly.** For fan-out and pipeline code, carrying the Context as a parameter and starting spans with an explicit parent removes the whole class of bug. This is the only option in Go and is often the cleanest in asynchronous code generally. **Use the framework's context.** Reactive streams and similar pipelines schedule operators on different threads deliberately; a thread-local cannot follow them. The integration must put the Context into the framework's own context object (the subscriber context in reactive-streams implementations, the coroutine context element for coroutines) so it travels with the logical flow, not the thread. Instrumentation libraries for these frameworks exist precisely because the naive thread-local approach cannot work. **Automatic instrumentation.** Agent-based auto-instrumentation patches the common executor and scheduler types so submitted tasks carry context without code changes. This is why the same code can be correct under an agent and broken without one — and why hand-rolled thread creation, custom pools and unusual concurrency libraries remain broken even with an agent installed. ## Diagnosing it If orphan traces cluster around one code path, look for a hand-off there. Log the trace id at the start of the task and at the caller: differing ids means loss; the *previous* request's id means a leaked scope. A quick reproduction is to run the path with a single-threaded executor — loss disappears, leaks get worse and more obvious. Also check that the span and the scope are closed in the right order in any code where a span is started in one method and ended in another.

  • You see traces whose root span lasts several hours and contains spans from unrelated tenants. What do you suspect?
    A leaked scope on a pooled thread: some handler attached a context and never closed it, so every later task on that thread inherits the stale context and parents its spans into that old trace. The long root duration is the giveaway — the root cannot end while children keep arriving, and the tenant mixing shows the context is shared across requests. Find the attach without a matching close, and enforce closure with a scoped-resource construct so exceptions cannot skip it.
  • Why is closing the scope not the same as ending the span?
    They control different things. Ending the span fixes its duration and hands it to the processor for export. Closing the scope restores the previously current context in the thread-local store. A span can legitimately outlive its scope — for example an asynchronous operation whose span ends in a callback — and a scope can be opened around a context that owns no span at all. Confusing them produces either spans that never end or contexts that never restore.
  • Does an auto-instrumentation agent make this problem go away?
    Only for the concurrency primitives it knows how to patch, typically the standard executor and scheduler types. Hand-rolled threads, custom pools, third-party concurrency libraries and manual callback registries are untouched, so context is still lost there. Agents also cannot fix code that starts a span in one method and closes the scope in another; the discipline still has to exist in application code.

The implicit context is a note pinned to the desk, not to the job folder. Move the job to another desk and the note stays behind — and if the previous worker never took their note down, the next job gets filed under the wrong case.

saying these in an interview costs you the question

  • Believing Context is mutable or globally shared rather than immutable and per-execution-unit.
  • Closing a scope on a different thread, or out of order, and expecting the store to recover.
  • Treating scope closure and span ending as the same operation.
  • Assuming an auto-instrumentation agent covers hand-rolled threads and custom pools.
  • Fixing orphan spans by storing the span in a static field — which converts a loss bug into a cross-request leak.
  • Expecting thread-locals to follow a reactive pipeline that deliberately hops schedulers.

context