An indexing failure's stack trace names only worker-pool frames and none of your code — what happened to the caller?
answer
- the trace belongs to the worker
- a hop is a handoff, not a call
- caller frames already returned
- carry context as data
- assembly traces cost a capture each
basics
~20 sThe trace was captured on the pool worker that raised the failure, and that worker's stack begins at its task loop. The code that built the chain and subscribed ran elsewhere and returned long ago.
solid answer
~50 sA stack trace describes one thread at one instant. Once a stage has hopped onto a worker pool, the failure is raised on a worker whose stack starts at the loop that pulled the task off a queue — a queue handoff is not a call, so the caller's frames were never pushed onto that stack and are already gone. Nothing is being hidden; the link simply does not exist in the trace. You restore it by carrying the link as data instead: put the work item's identity and a correlation identifier into the failure when it is created, place named markers at the boundaries of the chain so the failure records the last one it passed, and enable the assembly-time diagnostic mode only in test or under sampling, because it captures a stack per stage per subscription.
code
pseudocode · 8 linessource
.transform(fetch)
.marker("fetch")
.hop_to(parsing_pool) // frames below this point are the pool's
.transform(doc -> parse(doc)) // failure carries doc.id in its message
.marker("parse")
.on_failure(f ->
log(f, last_marker = f.last_marker, run = context.correlation_id))go deeper
Remember that a stack trace only describes the thread that raised the failure. Once work moves to a pool worker, the trace shows that worker and not whoever started the job.
Explain why: a hop is a queue handoff, so no frames from the caller were ever pushed onto the worker's stack, and the assembling code returned long before the failure existed.
Show the production repair: carry the work item's identity in the failure, mark chain boundaries so the last one is recorded, add a correlation identifier, and reserve the assembly-trace mode for sampling or test.
Treat it as a budget: decide how much per-signal diagnostic overhead the platform buys by default, and which paths are worth targeted deep tracing during an incident rather than always.
## A trace is a photograph of one thread A stack trace is captured at the moment a failure is created, and it contains exactly the frames that were active **on the capturing thread** at that moment. That is all it can ever contain. In synchronous code that is enough, because the thread that started the work is still the thread doing it: every caller is still on the stack below you. In an asynchronous pipeline that assumption breaks the first time a stage moves work onto another worker. ## What a hop actually is A hop is **not a call**. The upstream stage puts a task on a queue and returns; a worker picks it up later and runs it. Two consequences follow, and they explain everything about the useless trace: - The worker's stack starts at its own task loop — pull from queue, run, repeat — and contains nothing above that, because nothing above it called into the worker. - The frames of the code that assembled the chain and subscribed are gone. They returned when subscription completed, which may have been milliseconds or hours earlier. So a trace captured after one hop reads as: worker loop, scheduling glue, the pipeline's internal stages, the failure. It names your code only where your code is literally inside a stage body, and even then it does not name the caller that started the work. ## Restore the link as data, not as frames Since the frames cannot be recovered, everything useful must be **carried with the signal**: 1. **Work-item identity in the failure itself.** When the parsing stage fails, create the failure with the document identifier already in its message. That survives every hop, because it is data on the signal rather than a property of a thread. 2. **Named markers at chain boundaries.** Place a small named marker after each meaningful stage — fetch, parse, write. The failure records the last marker it passed, so even a trace full of worker frames comes with 'died after parse'. This is the cheapest thing on the list and the one most worth doing by default. 3. **A correlation identifier in subscription-scoped state.** Attach the identifier of the request or batch that started the work to the subscription, so every log line the run writes can be joined back to its origin. 4. **The assembly-time diagnostic mode.** Many implementations offer a mode that captures, for each stage at the time the chain was built, the stack of the code that built it, and attaches that to any failure travelling through. It genuinely answers 'which line assembled this pipeline'. | Technique | What it names | What it costs | |---|---|---| | Identity inside the failure | The item that failed | One field, set where the failure is created | | Named boundary markers | The stage the run died after | A small constant per marked boundary | | Correlation identifier | The request or batch that started it | A value carried with the subscription | | Assembly-time diagnostic mode | The code that built the chain | A stack capture per stage per subscription | ## Why the diagnostic mode is not simply left on Capturing a stack is not free: it walks frames and allocates, and this mode does it **per stage, per subscription**. A chain of ten stages taking a thousand subscriptions a second is ten thousand captures a second, all on the hot path, most of them for runs that never fail. That is why the honest production answer is: leave it on in test and in local debugging, sample it in production, or enable it for a targeted window on a targeted path — and rely on markers and identities the rest of the time. ## The indexing job, in practice The job fetches, hops to a pool to parse, and writes. Parsing fails, and the recorded trace names the pool loop and the pipeline's internals only. The on-call engineer cannot tell which document, which batch, or which stage. After the repair, the same failure records: the document identifier, the marker 'parse', and the batch correlation identifier. No stack trace improved at all — and the incident is now a two-minute diagnosis, because the missing information was never in the stack to begin with. ## What a strong answer sounds like Name the mechanism first — the trace belongs to the worker, and a queue handoff pushes no frames — so that the answer is clearly not a complaint about tooling. Then show that you know the repair is to carry context as data, and close with the cost of the one technique that does reconstruct assembly frames, because knowing when *not* to enable it is the senior half of the answer.
- Why is the assembly-time diagnostic mode usually unaffordable as an always-on setting?It captures a stack for every stage at every subscription, so cost scales with chain length times subscription rate and is paid on the hot path for runs that never fail. Sample it, scope it to one path, or keep it to test environments, and use markers and carried identity as the always-on substitute.
- Which parts of the failure's context survive every hop, and why those?Whatever you put into the signal itself: the work item's identity, marker names, and any subscription-scoped values you attached. They survive because they are data travelling with the signal, not properties of a thread. Everything derived from a thread's stack is recreated at each hop and describes only that worker.
saying these in an interview costs you the question
- Expects the worker's stack to contain the frames that subscribed
- Concludes the failure is inside the framework because framework frames dominate
- Blames the logging configuration rather than the thread handoff
- Claims the full diagnostic trace mode is cheap enough to leave on
- Assumes each hop rethrows and therefore chains the original stack