Why send on a semaphore `chan struct{}` before the `go` statement rather than inside the goroutine?
answer
- the loop can be the throttle
- goroutines are cheap, not free
- who parks: the spawner or the child
- N running versus N alive
basics
~10 sAcquiring before the go statement makes the spawning loop block, so only about N goroutines ever exist. Acquiring inside the goroutine starts one per item immediately: it bounds work in flight, not memory.
solid answer
~50 sBoth placements bound how many goroutines are past the acquire, so both cap the concurrent work. They differ in who waits. With `sem <- struct{}{}` before `go`, the blocking send happens on the loop's own goroutine, so the loop stops spawning and the live goroutine count stays around N. With the acquire as the first line inside the goroutine, the loop runs to completion instantly and you get one goroutine per input item, each parked on a channel send — a million-item input means a million parked goroutines, each holding a stack and everything it captured. For an in-memory slice of a few hundred items that is harmless; for a large or streaming input it is an unbounded memory footprint that looks fine in tests. Acquire outside unless you specifically need the spawn loop to keep running.
code
go · 8 linesfor _, part := range parts {
go func() {
sem <- struct{}{} // acquired inside: the loop never waits
defer func() { <-sem }()
uploadPart(part)
}()
}
// len(parts) goroutines now exist, each holding its stack and capturesgo deeper
Know that both placements limit how much work runs at once, and that only the earlier one also limits how many goroutines exist. Recall the shape: send, then go.
Explain precisely which goroutine parks in each version and what a parked goroutine still holds — a stack and every variable its closure captured — even though it uses no CPU and no thread.
Tie the choice to the input: a fixed slice makes them equivalent, a stream or queue makes the inner acquire an unbounded retention. Show the select on ctx.Done() that keeps the outer acquire cancellable.
Own the guidance your codebase gives: a single reviewed helper that acquires before spawning removes a whole class of quiet memory growth that only appears under production-sized inputs.
## Two placements that look equivalent Both of these bound concurrent uploads to four: ```go // A: acquire before go for _, part := range parts { sem <- struct{}{} go func() { defer func() { <-sem }(); uploadPart(part) }() } // B: acquire inside the goroutine for _, part := range parts { go func() { sem <- struct{}{} defer func() { <-sem }() uploadPart(part) }() } ``` In both, at most four goroutines are executing `uploadPart` at any moment. The difference is **which goroutine does the waiting**, and that decides what stays bounded. ## Who blocks In A, the send runs on the goroutine executing the loop. When four permits are out, the loop itself parks on the send. Nothing further is created. The number of goroutines alive because of this loop is roughly N+1: the N holding permits, plus the loop. In B, the loop never blocks. It creates a goroutine per item as fast as it can iterate and finishes. Every created goroutine immediately parks on the same full channel. The four with permits work; the rest sit in the channel's wait queue. ## Why parked goroutines are not free A goroutine parked on a channel send occupies no OS thread and burns no CPU — that part of the intuition is right, and it is why Go can hold very large numbers of them. But it is not weightless: - Each has a stack (small at creation, but it grew to whatever the goroutine needed before parking) that cannot be reclaimed while the goroutine is alive. - Each keeps alive everything its closure captured. If the closure captures a per-item buffer, a decoded record, or a `[]byte` read from disk, the whole set is retained simultaneously — the very memory the concurrency limit was supposed to bound. - Each is runtime bookkeeping the scheduler and the garbage collector must account for, and each appears in every goroutine dump you take while debugging something else. The practical failure is not "Go cannot handle many goroutines". It is that placement B converts a bounded-concurrency limit into a limit on *active* work only, while the *retained* work scales with the input. ## When the input size decides it Over a fixed slice of two hundred items, both placements are fine and B is arguably tidier. The moment the input is a stream, a directory walk, a database cursor, or a queue that never ends, B is a memory leak in slow motion: the producer races ahead of the consumers with nothing to stop it. Placement A gives you the stall for free — the loop simply cannot get ahead of the permits. This is why "acquire before `go`" is the default advice. It is not that B is wrong; it is that A is bounded in one more dimension for the same number of characters. ## What A costs you The cost of A is that the spawning goroutine is stuck. If that goroutine also has to do something else — poll for cancellation, read from another source, respond to a shutdown signal — it cannot while it is parked on the send. The fix is not to move the acquire inside; it is to make the acquire itself abandonable: ```go select { case sem <- struct{}{}: // acquired case <-ctx.Done(): return ctx.Err() } ``` Now the loop stops spawning when the permits are exhausted *and* can still give up promptly when the caller cancels. ## What neither placement gives you Neither bounds anything about goroutines started elsewhere in the process, and neither bounds memory held by the work itself. A permit says "one more unit may run", not "one more unit may allocate a bounded amount". If units differ wildly in footprint, a count-based limit lets N large ones coincide; that is a property of counting units rather than of where the count is taken. ## Saying it in an interview The crisp answer is one sentence and one consequence: acquiring before `go` puts the blocking on the spawner, so live goroutines are bounded too; acquiring inside puts the blocking on the child, so you bound the work but retain a goroutine and its captures per input item.
- If the spawning loop blocks on the acquire, what keeps it responsive to cancellation?Replace the bare send with a `select` whose cases are `sem <- struct{}{}` and `<-ctx.Done()`. The loop still stalls when permits are exhausted, but a cancelled context unparks it immediately and it can return rather than waiting for a permit that may never come back.
- Does acquiring before `go` bound the program's peak memory?No, only the goroutines this loop starts, to roughly N+1. Each running unit can still allocate as much as it likes, and code elsewhere starts its own goroutines. It bounds the concurrency you control, not the process, and it counts units rather than their footprint.
- Is acquiring inside the goroutine ever the right choice?Yes, when the spawning goroutine must keep making progress — draining a source, servicing a control channel — and the number of items is small and known. Over a bounded slice the retained goroutines are a fixed, acceptable cost; over a stream they are unbounded.
saying these in an interview costs you the question
- Assumes acquiring inside the goroutine bounds memory as well as work
- Says goroutines are free, so a million parked ones cost nothing
- Thinks the placement changes how many units run concurrently
- Believes each parked goroutine holds an OS thread
- Insists the spawning loop must never be allowed to block