Stepping over a channel receive in a Go program, why does the debugger resume in a different goroutine?
answer
- goroutines are not threads
- the runtime reused the thread you paused
- a breakpoint is an address, not an owner
- a hundred workers, one instruction
- wall-clock time keeps running while you read
basics
~20 sBecause the goroutine you were stepping blocked, and Go's runtime scheduler immediately put a different goroutine on that OS thread. A breakpoint address is process-wide, so you must use goroutine-aware stepping and switch goroutines explicitly rather than trusting the current thread.
solid answer
~50 sGoroutines are multiplexed onto OS threads by Go's runtime scheduler, and a debugger's primitives are thread-level. When the goroutine you are stepping blocks on a channel receive, a mutex, or a syscall, the runtime parks it and schedules another runnable goroutine onto the same thread — so a naive step resumes in code that has nothing to do with your goroutine. That is why a Go-aware debugger keeps a goroutine list and offers goroutine-scoped stepping and switching. Breakpoints have the same problem in the other direction: a breakpoint is an address, so every goroutine running that function stops the whole process, and with many concurrent workers you must narrow it by goroutine or by a condition on request data. Also remember that stopping distorts the program: `context` deadlines, server timeouts and heartbeats expire while you are paused, so stepping is a poor tool for timing bugs.
go deeper
Know that goroutines are scheduled onto OS threads by Go's runtime rather than being threads themselves, and that a breakpoint stops the whole process for whichever goroutine hits it first.
Explain the mechanism: when the goroutine you are stepping blocks on a channel or a mutex, the scheduler parks it and runs another on that thread, so goroutine-scoped stepping and an explicit goroutine list are required rather than optional.
Show the judgment: reduce concurrency or condition the breakpoint before stepping, and recognise which questions stepping cannot answer because pausing expires deadlines and serialises interleavings. Reach for -race or a goroutine snapshot instead.
Set the expectation for the team about which diagnostic fits which failure, so that incidents are not chased with a debugger on a paused process, and so that services carry the observability needed when stepping is not an option at all.
## Two different notions of "where I am" A debugger, at the level of the operating system, knows about *processes* and *threads*. It stops threads, single-steps a thread's instructions, and reads a thread's registers. Go's execution model does not line up with that. Goroutines are scheduled in user space onto a bounded set of OS threads by Go's runtime scheduler; a goroutine is not a thread, has no OS-visible identity, and does not stay on one thread. So "step" in the raw sense means "advance this thread", which is not the same as "advance the goroutine I was reading". ## Why a step over a blocking operation lands somewhere strange Suppose you step over `v := <-results`. If a value is already buffered the receive completes and nothing surprising happens. If it is not, the runtime does something entirely reasonable and entirely inconvenient for you: it parks the goroutine and hands the thread to another runnable goroutine so the machine keeps doing work. The thread you were stepping is now executing unrelated code, and the debugger, faithfully advancing that thread, drops you into a function you never asked about. The same happens on a mutex your goroutine cannot take, on a `time.Sleep`, and — differently but with the same visible effect — around a blocking syscall, where the runtime hands the P off so another thread can keep running Go code. This is why a debugger for Go has to be *Go-aware* rather than merely correct: it reads the runtime's own data structures to enumerate goroutines, shows each one's state and stack, lets you select one, and implements stepping scoped to that goroutine rather than to whatever the thread happens to be running. Without that, concurrent Go is close to un-steppable. ## Breakpoints are process-wide, and that is the other half of the problem A breakpoint is a modified instruction at an address. It has no idea which goroutine executes it. Put one inside a decode function that a hundred goroutines call, and the process stops for whichever one arrives first; continue, and it stops again for the next. If you are chasing the one input that produces a wrong value, you will step through ninety-nine irrelevant ones. The practical answers, in the order you should try them: 1. **Reduce the concurrency for the session.** Run the tool against one input, or with a worker count of one. A bug that still reproduces serially is far cheaper to inspect. 2. **Condition the breakpoint.** Make it fire only when a field identifies the case you want — the file name, the record id, a count that is zero. 3. **Scope it to a goroutine.** Once you have stopped in the right one, keep stepping inside it rather than trusting "continue". ## Stopping changes the program This is the judgment part, and it is the thing that separates someone who has actually debugged a live Go process from someone who has read about it. A stopped process is still subject to wall-clock time: - A `context.Context` created with a deadline will be cancelled by the time you resume, so the code after your breakpoint takes the cancellation path, not the path that failed in production. - Server read and write deadlines, client timeouts, and heartbeats all expire. - If you are debugging through `go test`, its default ten-minute timeout kills the binary while you are reading a struct. So interactive stepping is excellent for one class of question — *what is this value, and why is it the zero value?* — and actively misleading for another: anything whose answer depends on timing or on interleaving. A race is the clearest example. Setting a breakpoint serialises the very interleaving you are trying to observe, and "it works when I step through it" is not evidence of anything. Build with `-race` and let the detector report the conflicting accesses instead. Deadlocks and leaked goroutines are similar: you learn far more from a snapshot of every goroutine's stack than from advancing one of them a line at a time. ## The habit to carry Before reaching for a breakpoint in concurrent code, ask which question you are answering. If it is about a **value**, reduce the concurrency to one and step; that is exactly what the debugger is for. If it is about **ordering, timing or contention**, put the debugger down and reach for a tool that observes the program running at full speed.
- You set a breakpoint inside a function that two hundred goroutines run concurrently. What actually happens?The process stops for whichever goroutine reaches that address first, then again for the next, and so on — the breakpoint belongs to the code, not to a goroutine. Either cut the concurrency to one for the session, or condition the breakpoint on data that identifies the case you want, such as a file name or a record id.
- Why is stepping a bad way to investigate a suspected data race?Stopping serialises the interleaving you are trying to observe, so the bug usually stops reproducing, and "it works when I step" proves nothing. Build with `-race` and exercise the path instead: the detector reports the conflicting accesses with both stacks. Remember it only finds races on paths that actually execute.
- A handler behaves differently under a debugger than in production. What is the most common reason?Time kept passing while the process was stopped. Any `context.Context` deadline, server read or write deadline, client timeout or heartbeat will have expired, so on resume the code takes a cancellation or timeout branch it would never have taken at full speed. If the behaviour under investigation depends on a deadline, stepping is the wrong instrument.
saying these in an interview costs you the question
- Assumes a breakpoint belongs to one goroutine
- Treats the current OS thread as the current goroutine
- Concludes there is no bug because it works while stepping
- Uses a breakpoint to investigate a race instead of -race
- Forgets that context deadlines expire while the process is paused
- Steps through a hundred workers rather than reducing concurrency