How do you pin a specific goroutine interleaving in a Go test without time.Sleep?
answer
- a rendezvous, not a wait
- the fake tells you it is inside
- zero capacity makes both sides meet
- entered, then act, then release
- one schedule proven, not all of them
basics
~20 sGive a test-only fake collaborator two unbuffered channels: it sends on one when it is entered and blocks receiving on the other. The test receives the first signal, so it knows the goroutine is stopped inside that call, does its second operation, then releases the fake.
solid answer
~50 sUse an unbuffered channel as a rendezvous inside a test-only fake of a collaborator. The fake's method does `entered <- struct{}{}` and then `<-release`; the test receives from `entered`, which completes only when the fake has actually reached that line, so the goroutine is provably parked mid-call. The test now performs the second operation it wants to overlap, then closes `release` to let the first call finish. Two properties make this work: an unbuffered send completes only when a receiver takes it, so the two sides meet at a known point, and the fake blocks on `release`, so the window stays open for as long as the test needs rather than for however long a sleep guessed. Guard both receives with a `select` and a `time.After` case so a wrong assumption fails the test instead of hanging it. Keep the hook in the fake, never in production code.
code
go · 10 linestype gateStore struct {
entered chan struct{}
release chan struct{}
}
func (g *gateStore) Save(v string) error {
g.entered <- struct{}{} // completes only when the test receives
<-g.release // the test decides when Save may return
return nil
}go deeper
Know that an unbuffered channel send waits for a receiver, and that this is what lets a test observe another goroutine at an exact line. Being able to read such a test is enough at this level.
Explain the two-channel rendezvous — enter, act, release — and why buffering the first channel destroys the guarantee. Show where the hook lives so production code is untouched.
Demonstrate the judgment call: pin an ordering once you understand the bug, and keep it as the regression test, while separate repeated runs keep hunting for orderings nobody named. Bound every receive so a wrong assumption fails rather than hangs.
Own where testing seams are allowed to exist in shipped code and what the team pays for them. Be ready to argue when a deterministic reproduction is worth the coupling to internals and when the test belongs at a coarser boundary instead.
## The problem this solves Some bugs only appear when one operation is halfway through another: a second writer arriving while the first is between its read and its write, a shutdown running while a request is in flight, a cache entry being replaced while a reader holds the old one. To write a test for such a bug you must place one goroutine at an exact point and hold it there while another one runs. Widening the window with a sleep in the production path is not a test, it is a modification of the thing under test — and it still only makes the interleaving *likely*. ## The rendezvous An unbuffered channel is a rendezvous: a send completes only when a receiver takes the value, and the receive completes only when a sender offers one. Both goroutines are therefore at a known line at the same moment. That is the primitive to build on. Put it in a fake that satisfies the same interface as the real collaborator: ```go type gateStore struct { entered chan struct{} release chan struct{} } func (g *gateStore) Save(v string) error { g.entered <- struct{}{} // completes only when the test receives <-g.release // the test decides when Save may finish return nil } ``` The test then drives the ordering explicitly: ```go g := &gateStore{entered: make(chan struct{}), release: make(chan struct{})} svc := NewService(g) go svc.Update("first") <-g.entered // the first Update is provably inside Save second := svc.Update("second") // meets the first mid-flight close(g.release) // let the first one finish ``` After `<-g.entered` returns there is no guessing left: `Save` has begun and cannot proceed. The overlap you wanted to test is a fact of the program's control flow, not a probability. ## Why unbuffered specifically With `make(chan struct{}, 1)` the fake's send would complete into the buffer and `Save` would sail straight on to the `<-release`. You would still get *an* ordering, but the moment the test learns the call started is decoupled from the moment it started, and the two sides no longer meet. Zero capacity is what turns a notification into a synchronisation point. ## Bound every wait Every one of these receives is a potential hang, and a hung test is much worse than a failing one: it produces no message and takes the whole package's timeout with it. Wrap them: ```go select { case <-g.entered: case <-time.After(2 * time.Second): t.Fatal("Save was never called") } ``` That converts a wrong assumption — the code path you expected does not actually call `Save`, or calls it twice — into a named failure at the right line. Note the `t.Fatal` is on the test goroutine, which is where the fail-fast methods work. ## Where the hook lives The blocking belongs in test code. Three shapes, in order of preference: 1. **A fake implementing the collaborator's interface**, as above. Production code never sees it. 2. **An unexported function field** on the type under test that production leaves nil and the test sets — `if t.afterRead != nil { t.afterRead() }`. Cheap, but it is a testing seam in shipped code, so it should be rare and unexported. 3. **A sleep or an exported knob in production** — never. This is the shape a reviewer should reject: it changes the timing of the real system to make a test convenient. ## What a pinned interleaving does and does not prove It proves one thing extremely well: the specific ordering you named produces the specific outcome you assert, deterministically, on every machine, at full speed, with no sleeps. That makes it an excellent *regression* test — once you understand a concurrency bug, this is how you nail it down so it cannot come back silently. It proves nothing about the orderings you did not name. There is no combinatorial exploration here; you chose one schedule out of many. So the two techniques are complements, not alternatives: repetition under `go test -race` searches broadly for interleavings you did not think of, and a pinned rendezvous test locks down the one you did. A suite that has only the first kind cannot reproduce its own bugs; a suite that has only the second kind stops finding new ones. ## Cleanup Close or drain everything before the test returns. A fake left parked on `<-g.release` is a goroutine that outlives the test, and if it later touches `testing.T` it panics with a log-after-completion message. Registering the release in `t.Cleanup` is a tidy way to guarantee it even when an assertion fails early.
- Why must the entered channel be unbuffered rather than have capacity one?With a buffer the fake's send completes immediately into the buffer and the method runs on to its next line before the test has observed anything. Zero capacity forces the send to wait for the test's receive, so both goroutines are at a known point at the same instant — that is what pins the ordering.
- What does a pinned-interleaving test fail to prove?Anything about the schedules you did not choose. It verifies one named ordering deterministically, which makes it a strong regression test, but it explores nothing. Finding orderings nobody thought of needs repetition under the race detector; the two techniques cover different halves of the problem.
- Where should the blocking hook live so production code stays clean?In a test-only fake that implements the collaborator's interface, so shipped code never knows about it. If there is no seam at all, an unexported nil-by-default function field is an acceptable second choice. A sleep or an exported knob in the production path is not: it changes the real system's timing to suit a test.
- How do you stop the fake's goroutine outliving the test when an assertion fails early?Release it from t.Cleanup, which runs even after a fatal failure, or close the release channel in a defer at the top of the test. Otherwise the goroutine stays parked, and if it later calls a testing.T method it panics with a log-after-completion message that hides the real failure.
saying these in an interview costs you the question
- Widens the window with a sleep in production code
- Uses a buffered channel, so the fake never waits
- Never releases the fake, leaving it parked after the test
- Exports a test-only hook on the production type
- Claims one pinned ordering proves the code race free