Why does a tenant id stashed in per-worker ambient storage come back empty after a pipeline stage hops workers?
answer
- attached to the wrong thing
- worker, not request
- a hop changes the slot
- pooled workers reuse their slots
- bind it to the subscription
basics
~20 sAmbient storage is attached to the worker, not to the request: a read resolves against whatever worker is executing right now. Once a stage runs on a different worker, nothing ever wrote that worker's slot.
solid answer
~40 sA per-worker ambient slot is filled by a write executed on that worker and read by a read executed on that same worker. It only behaves like request state when one worker is dedicated to one request end to end. A pipeline breaks that: the chain is assembled by one piece of code and executed stage by stage, a hop hands a stage to a different worker, and a source may emit on a worker of its own — so the read runs where no write ever ran. The nastier variant is not emptiness: workers are pooled and reused, so the slot may still hold a value an earlier subscription left behind, and the reader gets another request's tenant with no error at all. The values have to ride the subscription instead.
code
pseudocode · 10 linesfunction handleRequest(request)
ambient.put("tenant", request.tenant) // slot on the arriving worker
return rows(request)
.map(function(row) return enrich(row))
.continueOnAnotherWorker() // later stages run elsewhere
.map(function(row)
tenant = ambient.get("tenant") // empty here, or stale from an earlier run
return writeAudit(row, tenant)
)go deeper
Recall that ambient storage is keyed by the worker running the code, so a value written on one worker is simply not there when another worker reads.
Explain the mechanics end to end: where the write landed, which worker runs each stage after a hop, and why a reused worker can still hold an earlier request's value.
Show how you would find it in production — audit rows with blank or mismatched tenants, traces losing identity mid-chain — and argue that the fix is to bind the values to the subscription rather than to pin the schedule.
Weigh what an implicit ambient channel costs when many services share it, and decide where request identity is written and who is allowed to read it.
## Where an ambient value actually lives Ambient storage attaches a slot to the **worker** that runs code, not to the **request** the code is serving. A write fills the slot belonging to whichever worker executes the write; a read returns the contents of the slot belonging to whichever worker executes the read. Nothing in that mechanism mentions a request. It behaves like request state under exactly one arrangement: a single worker is dedicated to one request from the moment it arrives until the response is finished, so *the current worker* and *the current request* happen to name the same thing for the whole call. ## The ways a pipeline breaks that identity A pipeline whose stages can move between workers violates the assumption from both sides: - The chain is **assembled** before it runs, and the code that assembles it is usually not the code that later executes each stage. - A stage may be handed to a different worker than the stage before it, so a write and a read straddle that boundary. - A source may begin emitting on a worker of its own, so even the first stage need not run where the subscription was opened. - Workers are **pooled and reused**: one worker serves many subscriptions over its life, and frequently several at once, interleaved. - Nothing copies a slot. A slot is filled only by an explicit write, and no write ever ran on the second worker. ## Two failure shapes, and which one hurts | what the reader finds | why | how it surfaces | |---|---|---| | an empty slot | no write ever ran on this worker | a blank field, a default value, or an absence branch taken quietly | | a value from another request | an earlier subscription wrote this reused worker's slot and nothing cleared it | a plausible wrong value: an audit row attributed to the wrong tenant | The empty read is the one people expect. The **stale** read is the one that costs money. The audit writer receives a tenant identifier that looks exactly like a real one, the row is written, nothing is raised, and the only evidence is a comparison of totals across tenants after the fact. Pooled workers are not scrubbed between units of work unless something scrubs them, and a pipeline offers no natural place to do it, because the stage that finishes is not necessarily the stage that started. ## Binding the value to the subscription instead The repair changes what the value is attached to: 1. The boundary that opens the subscription collects the values describing this unit of work — trace identity, tenant, caller identity — into an immutable key/value **context**. 2. That context rides with the **subscription**, so every stage of that subscription resolves a key against the same entries, whichever worker executes it. 3. Because the context is immutable, a derived context is a new value rather than a mutation, so two subscriptions sharing a worker cannot overwrite each other's entries. The hop stops mattering, because the lookup path no longer mentions a worker at all. ## Why pinning the stages is not the fix The tempting repair is to keep every stage on the worker that received the request. That removes the symptom and leaves the defect: the pipeline still depends on a property of the **schedule** rather than a property of the **subscription**. Any later change — an operator that introduces a hop, a source that emits elsewhere, a stage moved off a small pool because it began to block — reinstates the failure, in a change whose author has no reason to suspect it. Pinning does not even cover interleaving: one worker serving two subscriptions in turn still has exactly one slot between them. ## What still needs care after the move - Code you do not own may read only worker-bound storage. Bridging is then explicit and local: copy from the subscription context into the slot immediately around that call, on the worker that makes it, and clear the slot afterwards. - Absence has to be **loud**. A read that returns nothing while the stage carries on is the same defect one level down; a mandatory key should fail the subscription rather than fall back to a default. - The context describes the subscription and is fixed for its life. Anything computed per element belongs with the elements, not in the context. The short version an interviewer wants: the slot is keyed by the executing worker, a pipeline changes which worker that is, and reused workers make the wrong answer look right.
- Why is a stale value read from a pooled worker's slot more dangerous than an empty one?An empty read at least tends to produce a blank field somebody eventually questions. A stale read produces a well-formed identifier belonging to a request that already finished, so the audit row is written, every shape check passes, and nothing fails. It is discovered only by comparing records across tenants.
- Does keeping every stage on one worker fix it?It hides it. Pinning removes the hop but not the dependency on where the code runs, and any later scheduling change brings the failure back. It also does not help when one worker interleaves several subscriptions, since they share the single slot. The value belongs to the subscription, not to a worker.
- What has to happen when a stage must call code that only reads worker-bound storage?Bridge explicitly and narrowly: on the worker about to make the call, copy the values out of the subscription context into the slot, make the call, then clear the slot in a finally-style path. Leaving the copy behind is precisely what creates the stale reads for the next subscription on that worker.
A per-worker slot is a note pinned to a desk rather than tucked into the file. Send the file to another desk and it arrives blank; leave the note behind and whoever sits there next reads it as if it were theirs.
saying these in an interview costs you the question
- Says the value is lost because the pipeline is asynchronous, not because it hops workers.
- Believes an empty read surfaces as an error rather than a quiet default.
- Assumes a pooled worker's slot is cleared automatically between units of work.
- Claims the ambient value travels along with the values flowing through the operators.
- Thinks copying the slot to every worker at startup would solve it.