In runtime/trace, how would you find which stage of a hung ingestion workflow never finished?
answer
- look for what never ended
- the absence of an event is the clue
- the last region start names the stage
- the log messages name the item
- no end event, no entry in the latency
basics
~20 sCapture a trace while the process is still stuck and look for the task with no end event. Its last recorded events are the start of a region that never ended and the log messages before it, which name the stage and the item.
solid answer
~50 sAnnotation turns a hang into an absence you can point at. Collect a trace while the workflow is still wedged, then open the user-task view: the stuck unit of work is the task that has a start and no end. Inside it, the last event is a `trace.StartRegion` with no matching end — that region's type is the stage name — preceded by whatever `trace.Log` messages you emitted, which carry the document id. The runtime records the goroutine parking at the moment it blocks, and after that it emits nothing at all, so the diagnosis is literally "the events stop here". What the trace cannot tell you is *why*: it names the stage and the item, and a goroutine stack dump names the line. In a workflow whose stages wait on a context with no deadline, the pair is usually enough — the stage that never ends is the one whose wait was never bounded.
code
go · 6 linesreg := trace.StartRegion(ctx, "enrich")
trace.Logf(ctx, "doc", "id=%s waiting for enrichment", doc.ID)
doc.Enriched = <-enriched // ctx carries no deadline: this can wait forever
reg.End() // never reached, so the region and the task have no end eventgo deeper
Understand that a stage which never finishes leaves its annotation unclosed: the region and the task have a start event and no end. The missing end is the evidence, not an error message.
Explain what the sequence of events under one task looks like when a stage hangs, and why a task with no end event cannot appear in the latency figures for its type.
Show the whole diagnosis: capture while the process is still stuck, find the task with no end, read the last region type and the log messages for the item, and pair that with a goroutine stack to get the line.
Own the operational posture that makes this possible: whether trace collection is reachable in production, what the standing annotation vocabulary is, and which signal pages you when work starts and never finishes.
## The shape of the problem A multi-stage ingestion workflow stops making progress. Its own application logs show `stage 3 started` and then nothing — which tells you where it got to and nothing about why it stayed there. Every stage waits on something: a channel, a network call, another stage. One of them is waiting on a context that will never be cancelled, and the workflow will sit there until the process is restarted. This is the case that user annotations are unusually good at, for a slightly counter-intuitive reason: **they turn the hang into a missing event rather than a missing log line.** ## Reading the incomplete task With `trace.NewTask` per document and `trace.StartRegion` per stage, a healthy document produces a tidy sequence: task start, region start, region end, region start, region end, task end. A hung document produces a truncated one: task start, some regions that completed, then a region start with no end, and no task end. That truncation is the answer to "which stage": - The **last region started with no matching end** is the stage that never returned, named by its region type. - The **`trace.Log` events on that task** give you the identity — the document id, the size, the retry count — so you know which item, not just which stage. - The **task has no end event**, which is why the operation is invisible in the aggregate latency for its type. Latency needs two events; this one has one. That last point is worth sitting with, because it explains a familiar and dangerous symptom: while a workflow is silently wedging, the latency distributions can look perfectly healthy. Everything that finishes, finishes fast. The stuck work never enters the statistics at all. Counting tasks that start and never end is a different signal from measuring the ones that complete, and only the first one sees this failure. ## What the runtime adds and where it stops When a goroutine blocks — on a channel receive, on a mutex, in a network wait — the runtime records that it parked. After that the goroutine emits nothing, because it is doing nothing. So the trace does not contain a helpful "still waiting" heartbeat; it contains a final event and then silence for the rest of the recording. Reading a stuck trace is reading a gap. This also sets the limit of what the trace can tell you. It names **which logical work item** and **which stage**, in your own vocabulary. It does not name the line of code or the reason the wait was unbounded. A goroutine stack dump does that, and the two are complementary: the trace tells you *which* of the forty goroutines named `worker` matters and what it was doing on behalf of whom; the stack tells you the exact receive it is parked on. ## Capturing it in the first place The practical constraint is that annotations exist only inside a trace that was actually being collected. A trace taken after the restart contains nothing about the hang. So the collection has to happen while the process is still wedged — which is a strong argument for making trace collection reachable in production rather than something you add to a build afterwards. A second and duller failure is worth ruling out before you conclude anything: if the trace shows **no user tasks at all**, the likely cause is not the workflow but the plumbing. Either tracing was not running when that code executed, or the stage never received the context returned by `NewTask` — someone in the middle of the chain took no context, or substituted a fresh one — so its events were attributed to no task. ## Turning the finding into a fix The finding is usually not subtle once you have it: stage 3 waits on a result that a cancelled or failed upstream will never deliver, and the context it was handed has no deadline, so the wait is unbounded. The fixes are ordinary — give the stage a bounded wait, make the upstream close its channel on every exit path, ensure the cancellation actually propagates to the goroutine doing the waiting — and the annotation you already added becomes the verification: after the change, the same task completes, the stage's region has an end event, and the stage appears in the region view with a real latency distribution instead of vanishing. ## The habit worth keeping Instrument the stages before you need to. The task and region vocabulary costs almost nothing when tracing is off, and its value is highest in exactly the situation where you cannot reproduce the problem: the process is stuck right now, and you have one chance to record what it is stuck on. An unannotated trace of a hung process shows you parked goroutines with no idea which document, which stage, or which retry they belong to.
- Why does a hung workflow of this kind leave the latency distributions looking healthy?Because a latency needs a start and an end event, and the stuck task only ever recorded a start. Everything that completes still completes quickly, so the aggregates over completed work look fine while the wedged work is invisible in them. The signal that catches it is counting operations that begin and never finish, not measuring the ones that do.
- The trace shows no user tasks at all. What are the likely causes?Either tracing was not running while that code executed — annotations are recorded only during collection — or the stage never received the context returned by NewTask, so its events belonged to no task. The usual culprit for the second is a function in the middle of the chain that takes no context or creates a fresh one, cutting the task off from everything below it.
- What can the annotated trace not tell you about this hang?Why the wait was unbounded. It names the work item and the stage in your own vocabulary and shows that the goroutine parked and never resumed, but it does not name the line or the missing cancellation. A goroutine stack dump supplies that, and the trace is what tells you which of many identical-looking goroutines to read.
It reads like a relay handoff sheet where the last runner signed out of the exchange and never signed in at the next one. The missing signature, not any entry on the sheet, tells you which leg failed.
saying these in an interview costs you the question
- Expects a blocked stage to keep emitting trace events
- Trusts the latency distribution while tasks never complete
- Collects the trace only after restarting the process
- Believes the trace shows why a context never cancelled
- Cannot say which events an unfinished task is missing