skip to content

Why does a hung `go test` run report a test timeout instead of Go's deadlock error?

level: middleimportance: nice to knowfreq 32%

answer

  1. the harness has a clock of its own
  2. an armed timer blocks the runtime's proof
  3. ten minutes unless you say otherwise
  4. the panic dumps every goroutine
  5. -timeout 0 removes the harness timer

basics

~10 s

The test harness arms its own timeout timer, and any pending timer stops the runtime from claiming a deadlock. When it fires, the harness panics and dumps every goroutine's stack — the same evidence.

solid answer

~40 s

`go test` starts a timer for its `-timeout` budget, ten minutes by default, and covers the whole test binary run with it. A pending timer means future work exists, so the runtime's whole-program deadlock proof can never conclude inside a test binary. When the budget expires the harness raises the traceback level, prints which tests were still running, panics with `test timed out after 10m0s` and dumps every goroutine's stack — which is exactly the evidence you wanted. In practice you shorten `-timeout` in CI so a wedged test produces that dump quickly instead of being killed by the runner with no output at all. If you want the runtime's own verdict on a minimal reproduction, run it with `-timeout 0` and strip out every `time.Sleep`, ticker and background server so nothing stays armed.

code

text · 7 lines
text
$ go test -timeout 30s ./controller
panic: test timed out after 30s
	running tests:
		TestReconcileLoop (30s)

goroutine 62 [sync.Mutex.Lock, 30s]:
...

go deeper

for a junior

Recognise the test timed out after panic as your friend: it prints every goroutine's stack and names the tests still running, which is usually enough to find where the run got stuck.

for a middle

Explain the causal chain — the harness arms a timer, a pending timer disqualifies the runtime's whole-program deadlock proof, so the timeout panic is what you get instead — and know that the budget covers the whole run.

for a senior

Show you tune the budget so the harness wins the race against the CI runner's kill, and that you can strip a reproduction of timers and I/O to get the runtime's own verdict when you need proof rather than stacks.

for a principal

Set the policy: per-package timeout values that guarantee evidence on every hang, and the expectation that a hung test is investigated as a real deadlock rather than retried as flakiness.

## Why the runtime stays quiet in tests The runtime reports `all goroutines are asleep - deadlock!` only when nothing in the process is runnable, nothing sits in a system call, and no timer is armed anywhere. The `testing` harness violates the last condition on purpose: at the start of a run it arms a timer for the `-timeout` budget. From that moment there is always pending work in the scheduler, so the runtime's proof is unavailable for the entire run. A test that deadlocks therefore hangs silently rather than crashing with the fatal error people expect. ## What you get instead When the budget expires, the harness does three useful things before dying: 1. It raises the traceback level so the dump covers *every* goroutine, not just the one that panics. 2. It prints which tests were still running and for how long, which immediately narrows the search. 3. It panics with `test timed out after <duration>`, so the process exits non-zero and the whole goroutine dump lands in the test output. That dump is the same artefact you would collect from a hung production process, and it is where you do the real work: find the goroutines parked in a lock acquisition or a channel operation, group them, and identify the one holding things up. ## The budget covers the run, not each test A common misreading is that `-timeout` is a per-test allowance. It is not: it bounds the whole test binary's execution. A package of many slow tests can exhaust the budget with no single test being slow, and the panic message names whichever tests happened to still be running when it fired. If you need a per-test bound, that is your own code's job — a `context.Context` with a deadline, or `t.Deadline` to see how much of the shared budget is left. ## Why you should shorten it in CI The default ten minutes is longer than most CI job step limits are generous about, and there is a failure mode worth avoiding: if the CI runner kills the process first, you get no goroutine dump at all — just a job that was terminated. Setting something like `-timeout 60s` for a fast package means the harness always wins the race, and a wedged test hands you complete evidence on the first red build instead of on the fourth attempt with extra logging bolted on. The cost of a short timeout is flakiness on genuinely slow packages, so pick per-package values rather than one global number, and treat a package that legitimately needs minutes as a signal in its own right. ## Turning the harness off to use the runtime's proof Sometimes you want the runtime to *confirm* a suspected cycle rather than hand you stacks to interpret. Passing `-timeout 0` disables the harness timer entirely. Then, in a reduced reproduction: - delete every `time.Sleep` — a sleeping goroutine arms a timer; - stop every `time.Ticker` and avoid `time.AfterFunc`; - start no HTTP server, open no listener and touch no files, since a goroutine parked in the network poller or a thread in a system call also disqualifies the check. With all of that gone, a genuine total deadlock produces the runtime's fatal error and its dump directly, which is a much stronger statement than "it did not finish in time": the runtime has *proved* no progress is possible, rather than merely observed that none happened. ## What this does not give you Neither mechanism detects a partial deadlock. If a test's goroutines are wedged but the test itself keeps making progress, nothing fires and the test simply passes with leaked goroutines. That is a different problem with different tools, and it is worth being explicit that a green test suite is not evidence that a cycle is impossible — only that the paths you exercised did not hit it in the order that triggers it. ## The habit to build Treat a hung test as a first-class diagnostic opportunity rather than an annoyance: it is a deadlock reproducing itself, in an environment you control completely, with a mechanism that will hand you every goroutine's stack for free. That is a far better starting position than the same cycle appearing once, at three in the morning, in production.

  • What do you set so a hung test still yields evidence in CI?
    A short `-timeout` for that package — often tens of seconds. The harness then panics and dumps every goroutine's stack before the CI runner's own kill, which produces nothing. Pick per-package values rather than one global number, so genuinely slow packages do not turn flaky.
  • How do you make a minimal reproduction produce the runtime's own deadlock error?
    Run it with `-timeout 0`, or as a plain program, and remove everything that keeps the scheduler expecting work: `time.Sleep`, tickers, `time.AfterFunc`, listeners and file I/O. With nothing armed and no thread in a system call, a true total deadlock makes the runtime print its fatal error.
  • Does the -timeout budget apply per test function?
    No — it bounds the entire test binary run. Many moderately slow tests can exhaust it with no single test being slow, and the panic simply names whichever tests were still running. For a per-test bound, use a `context.Context` with a deadline in your own code.

saying these in an interview costs you the question

  • Expects the runtime deadlock error to appear during tests
  • Lets the CI runner kill the job instead of setting -timeout
  • Leaves time.Sleep in a minimal deadlock reproduction
  • Thinks -timeout is a per-test-function budget
  • Treats a hung test as flakiness to retry