skip to content

When ambient context is carried across a thread hand-off, you can either copy its value at hand-off time or have the receiving side read the current value when it eventually runs. Compare the two semantics and say when each is correct.

level: seniorimportance: should knowfreq 35%

answer

  1. capture = bound to the work item; re-read = bound to the executing thread
  2. context identifies the operation → capture by default
  3. re-read on a pool: empty or another task's leftovers
  4. capture a frozen snapshot, never a mutable reference
  5. deadline: capture the absolute instant, derive the remainder at run time

basics

~20 s

Capturing takes a snapshot at hand-off, so the task sees the context of the operation that created it, however late it runs. Re-reading sees whatever is current on the executing thread — usually empty or another task's leftovers. Capture is almost always what you want; re-read only suits genuinely global, dynamic settings.

solid answer

~60 s

**Capture** binds the value at hand-off: the snapshot travels with the work, so the task always sees the context of the operation that *caused* it, even if it runs a minute later. This is what correlation ids, tenants, and trace parents need — the context describes the logical operation, and re-reading later would attach the work to whatever unrelated request happens to be on that thread. **Re-read** looks up the value on the executing thread when the body runs. On a pool that means empty, or a stale value another task left — plausible-looking and wrong. It is only correct when the value is genuinely *ambient and global* rather than per-operation: a dynamic log level, a kill switch, a config generation you want read as late as possible. Two practical rules. Capture *immutable snapshots*: capturing a reference to a mutable context object gives you neither semantics, since the value can change under you between hand-off and execution. And decide per key — most keys capture, a small explicitly-listed set re-reads.

code

text · 10 lines
text
submit at t=0 (request R1 in flight, CTX = cid:R1)

CAPTURE:  task carries {cid:R1}
          runs at t=8s on worker-2  -> logs cid:R1        (correct)

RE-READ:  task carries nothing
          runs at t=8s on worker-2
          worker-2 currently holds leftovers {cid:R7}
                                     -> logs cid:R7       (wrong, and
                                        looks entirely legitimate)

go deeper

for a junior

Say that capture takes a copy when the work is handed off while re-read looks at whatever the running thread has, and that copying is normally what you want.

for a middle

Explain why per-operation identity must be captured and what re-read actually returns on a pooled worker, with the stale-context failure spelled out.

for a senior

Show the edge cases: immutable snapshots rather than mutable references, absolute deadlines rather than durations, and derived child spans for tracing.

for a principal

Frame it as per-key propagation policy owned by the platform layer, with explicit defaults, an allow-list for re-read keys, and a rule that anything needing mid-flight mutation should not be ambient context at all.

## The two semantics stated precisely A context value is read at *some* moment. Propagation designs differ only in which moment. **Capture (snapshot-at-hand-off).** At the instant work is handed to another thread, the current values are copied into the task. When the task runs — one millisecond or ten minutes later, on any worker — it installs that copy. The value is bound to the *work item*. **Re-read (lookup-at-execution).** The task carries nothing. When the body runs, it reads the ambient slot on whichever thread is executing it. The value is bound to the *thread at execution time*. Almost every real bug in this area is a case of one being used where the other was meant. ## Why capture is the default Most ambient context answers the question "which logical operation is this?" — correlation id, tenant, user, trace parent, request deadline. Those properties belong to the operation that *created* the work. Under re-read semantics on a pool, the worker either has nothing (context lost) or holds a leftover from an unrelated task (context wrong). Wrong is worse than lost: your logs, metrics, and traces will confidently attribute background work to a request that had nothing to do with it, and an on-call engineer will chase the wrong request for an hour. Capture also has the property you want under delay. A task queued behind a long backlog, a retry firing thirty seconds later, a scheduled follow-up — all still carry the originating operation's identity. Nothing about their lateness changes which operation caused them. ## When re-read is actually right Re-read is correct when the value describes *the environment at the time of execution*, not the operation: - **Dynamic log level or sampling rate.** If an operator raises verbosity while a task sits queued, you want the task to log at the new level, not the level in force when it was submitted. - **Feature kill switches / circuit state.** A task should observe the switch as it stands when it runs; acting on a stale "enabled" snapshot is precisely what a kill switch exists to prevent. - **Current configuration generation.** Late-bound config is often the point. Notice what these have in common: they are process-global settings that happen to be read through an ambient mechanism, not per-operation identity. In practice they are usually better stored in an explicit global/config holder rather than thread-local storage at all — which is a good tell that if a value truly wants re-read semantics, it probably should not be thread-local. ## The trap in between: capturing a mutable reference A design that captures a *reference to a mutable context object* gets neither semantics cleanly. The task holds the object from hand-off time, but its fields can be mutated afterwards by the originating thread — so what the task observes depends on timing, which is a data race in the ordinary sense and needs synchronization to even be well-defined. The fix is to capture an **immutable snapshot**: copy the values into a frozen object at hand-off. Then "captured" means exactly one set of values, forever, and no synchronization is required because nothing can change. If a per-operation value genuinely needs to be updated mid-flight and seen by already-dispatched work, thread-local context is the wrong mechanism — use a shared, explicitly synchronized structure with an obvious owner, so the sharing is visible in the code. ## Deadlines: the instructive middle case A request deadline shows why the distinction needs care. Capturing the *remaining duration* ("you have 300 ms") is wrong: if the task sits queued for 250 ms, it starts believing it has 300 ms left and blows the caller's budget. Capturing the *absolute instant* ("finish by T") is right: the task computes its remaining budget when it runs and correctly discovers there are only 50 ms left — or that the deadline has already passed and the work should be abandoned. So the rule is capture the invariant fact, then derive time-relative quantities at execution. The same reasoning applies to anything expressed as a duration rather than a point. ## Per-key policy, not a blanket rule A mature propagation layer treats this as configuration per key rather than one global choice: - Correlation id, tenant, user → **capture**, immutable. - Trace context → **capture the parent**, and *derive* a new child span at execution, so the trace shows a proper parent-child relationship rather than one span with impossible overlapping activity. - Deadline → **capture the absolute instant**, compute the remainder at execution. - Log level, kill switches, config generation → **re-read** (and consider whether they should be in thread-local storage at all). The operational rule that follows: whichever semantics a key uses, the value installed on the executing thread must still be removed or restored in a finally, or you have a leak on top of a semantics question. ## How to answer in an interview State the two semantics in one sentence each, say capture is the default because context identifies the operation rather than the environment, give the stale-context failure of re-read on a pool, then show sophistication with the two edge cases: capture immutable snapshots rather than mutable references, and capture absolute deadlines rather than remaining durations.

  • Why is capturing a reference to a mutable context object unsatisfactory?
    Because it gives neither semantics. The task holds the object from hand-off time, but the originating thread can still mutate its fields afterwards, so what the task observes depends on timing and is an ordinary data race unless you synchronize. Capturing a frozen, immutable snapshot makes 'captured' mean exactly one set of values forever and removes the need for synchronization entirely.
  • How should trace context be handled across an async hand-off?
    Capture the parent span's identity, then derive a new child span when the task executes and link it to that parent. Copying the span verbatim makes the trace show one span with overlapping concurrent activity, so durations and the parent-child structure become meaningless. This is the clearest case where propagation is per-key policy rather than a blanket copy of everything in the context.
  • Give a value where re-read semantics is genuinely the right choice.
    A dynamic log level or a feature kill switch. If an operator raises verbosity or flips a switch while a task is still queued, you want the task to observe the new setting when it actually runs, not the value that was in force at submission — acting on a stale 'enabled' snapshot is exactly what a kill switch exists to prevent. Such values describe the environment at execution time rather than the operation, and they usually belong in an explicit config holder rather than in thread-local storage.

Capture is stapling a photocopy of today's instructions to the job ticket; re-read is telling the next shift to check whatever is pinned to the noticeboard when they get to it. For 'who ordered this and why', the staple is right. For 'is the machine currently allowed to run', the noticeboard is right.

saying these in an interview costs you the question

  • Assuming re-read will 'just work' on a pool, when the executing thread's slot is empty or holds another task's leftovers
  • Capturing a mutable context object by reference and calling that a snapshot
  • Capturing a remaining duration for a deadline instead of the absolute instant
  • Applying one blanket semantics to every key rather than deciding per key
  • Focusing on capture-versus-re-read while forgetting the installed value must still be removed or restored in a finally

context