A Go test fails at teardown with 'TempDir RemoveAll cleanup: access is denied' on Windows — what causes it?
answer
- the assertions were never the problem
- one platform hides it, the other reports it
- somebody never closed what they opened
- registration order decides teardown order
- log inside each cleanup and run with -v
basics
~20 sSomething still holds an open file inside the temp directory. Windows refuses to delete a file with a live handle, so the automatic removal registered by t.TempDir fails and reports an error, turning an otherwise passing test red.
solid answer
~50 sThe assertions really did pass; the failure comes from teardown. `t.TempDir()` registers its `RemoveAll` through `t.Cleanup` at the moment you call it, and if that removal fails it is reported as a test error. On Unix a file can be unlinked while open, so a leaked `*os.File` is invisible; on Windows an open handle blocks deletion, and the same test goes red. The leak is usually in the code under test — a CLI that rewrites files in place and returns before closing its output — or in a test helper that opened a file and never closed it. The fix is to close every handle, and to make the ordering deliberate: because cleanups run last-registered-first, register the close **after** calling `t.TempDir()` so the close runs before the removal. Check `Close`'s error while you are there. Running with `-v` and logging from each cleanup shows the real order in one run.
code
go · 15 linesfunc TestRewriteInPlace(t *testing.T) {
dir := t.TempDir() // its removal is registered here, so it runs LAST
path := filepath.Join(dir, "out.txt")
f, err := os.Create(path)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { // registered later, so it runs FIRST
t.Logf("cleanup: closing %s", path)
if err := f.Close(); err != nil && !errors.Is(err, os.ErrClosed) {
t.Errorf("close: %v", err)
}
})
// ... exercise the rewriter against dir ...
}go deeper
Recall that the temp directory is removed for you and that the removal can itself fail the test. Closing every file you open is the habit being checked.
Explain the mechanics: the removal is a cleanup registered when t.TempDir was called, cleanups unwind LIFO, and registering the close afterwards makes it run first.
Diagnose it end to end: read the failure as a teardown error, use -v log lines to prove the cleanup order, and trace the leaked handle back to the code under test rather than patching the test.
Treat the platform split as a signal about your test matrix: if only one platform reports leaked handles, decide whether the pipeline runs there at all, and set the expectation that leaks are fixed at the source rather than retried around.
## Read the failure literally A line of the form `TempDir RemoveAll cleanup: ...` is not an assertion failure. It comes from the cleanup function that `t.TempDir()` registered on your behalf: it called `RemoveAll` on the directory, the removal failed, and it reported the error against the test. Everything the test was actually checking succeeded. That is why the message appears *after* all of the test's other output, including every subtest's, and why the same test looks fine on a developer's machine and red on a Windows agent. ## Why the platform split exists On Unix-like systems, unlinking a file that some process still has open is legal — the directory entry disappears and the data lives until the last descriptor closes. `RemoveAll` therefore succeeds even when the test leaked a handle, and the leak stays hidden. Windows, by default, refuses to delete a file that has an open handle, so the same leak surfaces as a permission-shaped error. Go's testing package retries the removal briefly on Windows before giving up, which is why the failure can also be intermittent: if the handle happens to be released by a finalizer or a straggling goroutine within the retry window, the removal succeeds and nobody sees anything. So the Windows failure is not a Windows bug. It is the honest report of a resource leak your other platform was hiding. ## Where the handle usually is Three places, in order of likelihood: 1. **The code under test.** A tool that unpacks an archive into a working directory and rewrites files in place opens each destination file, writes, and returns without closing — often because the close was written as `defer f.Close()` inside a loop body's enclosing *function*, so it does not run until the whole walk finishes, or because an early error path returns before the close. 2. **A test helper.** A helper opens a fixture file to seed data and returns the path, leaving the `*os.File` open. 3. **A goroutine the test started** that is still writing when the test body returns. The test then finishes, cleanup runs, and the write is still in flight. The first is the valuable find: the leak exists in production too, where a long-running process eventually exhausts file descriptors. The test did you a favour. ## Making the ordering deliberate Cleanup functions run last-registered, first-called. `t.TempDir()` registers its removal *when you call it*, so anything you register afterwards runs before the removal. That gives you the correct pattern almost for free: ```go dir := t.TempDir() // registered first -> runs last f, err := os.Create(filepath.Join(dir, "out.txt")) if err != nil { t.Fatal(err) } t.Cleanup(func() { // registered second -> runs first if err := f.Close(); err != nil && !errors.Is(err, os.ErrClosed) { t.Errorf("close: %v", err) } }) ``` Invert those two statements — obtain the file from a helper that registers its close, *then* call `t.TempDir()` — and the removal is now the later registration, so it runs first and fails. When a helper owns both, have the helper take the directory as a parameter so the registration order follows the dependency order. ## Confirming it rather than guessing Add a `t.Logf` at the top of each cleanup and run the package with `-v`. The log lines appear in the order the cleanups actually ran, and you can see at a glance whether the close ran before the removal or after it. That takes the argument out of the realm of opinion, and it is far quicker than reasoning about registration sites scattered across helpers. If the handle is not in test code at all, the same `-v` run tells you that — the close ran first and the removal still failed, which points at the code under test. ## Fixes that are not fixes - **Retrying `RemoveAll` in a loop.** This converts a real leak into a slow test that usually passes, and leaves the descriptor leak in production code. - **Swapping `t.TempDir()` for `os.MkdirTemp` and ignoring the removal error.** Now the failure is silent and the agent's disk fills up over a few thousand runs. - **Skipping the test on Windows.** The leak is still there; you have only removed the one platform that reports it. - **Calling `runtime.GC()` before cleanup** in the hope a finalizer closes the file. Unreliable by construction and a sign the ownership of the handle is unclear. The fix is to give every opened file a definite owner that closes it, check the error from `Close` on files you wrote to (a failed close can mean lost data, not just a leaked handle), and let the LIFO ordering of cleanups line up teardown with acquisition.
- Why does the same test pass on Linux?Unix lets you unlink a file that is still open: the entry vanishes and the data survives until the last descriptor closes, so RemoveAll succeeds and the leak stays invisible. Windows blocks deletion while a handle is live, which turns the same leak into a visible teardown failure.
- The close in test code demonstrably runs before the removal, and removal still fails. Where do you look next?At the code under test and any goroutine it started. A tool that rewrites files in place commonly returns before closing its output, or an error path skips the close. That is a real descriptor leak that would also bite a long-running process, not just this test.
- Would retrying RemoveAll in a loop be an acceptable fix?No. It converts a genuine resource leak into a slower test that usually passes, hides the same bug in production code, and leaves the run flaky whenever the handle outlives the retry window. Close the handle and give it a clear owner instead.
- Why check the error from Close on a file you wrote to?Because buffered data can be flushed at close time, so a failed Close can mean the write never reached disk. Ignoring it turns a lost-data bug into a silently passing test; on a file you only read, the error is far less interesting.
saying these in an interview costs you the question
- Blames a flaky CI agent instead of an unclosed handle
- Assumes a passing assertion means the test cannot fail afterwards
- Adds a retry loop around RemoveAll and calls it fixed
- Skips the test on Windows to make the pipeline green
- Thinks cleanup functions run in registration order
- Ignores the error returned by File.Close