skip to content

If a rewrite fails after os.CreateTemp, what is left in the directory and how should the code clean it up?

level: middleimportance: should knowfreq 32%

answer

  1. the create already touched the filesystem
  2. every early return leaves it behind
  3. it sits where the reader is looking
  4. register the cleanup on the next line
  5. removing a name that is gone is harmless

basics

~20 s

A partially written temporary file is left beside the target, where it had to be created. Register defer os.Remove(f.Name()) right after os.CreateTemp: it clears the leftover on every failure path and is a harmless no-op once the rename has succeeded.

solid answer

~50 s

`os.CreateTemp` creates a real file in the target's directory, and every `return err` after that line leaves it there. Over time a controller that fails often enough litters the directory with half-written files, and if the consuming process reads the whole directory rather than the exact path, it eventually parses one of them. The fix is one line placed immediately after the create: `defer os.Remove(f.Name())`. On failure it removes the scratch file; on success the rename has already consumed that name, so the remove fails with a not-exist error you deliberately ignore, and the published file is untouched because you removed the temp *name*, not the target. Use `f.Name()` rather than a name you composed, since `os.CreateTemp` chose the random part. Give the pattern a distinctive shape - a leading dot and a `.tmp` suffix - so a directory-scanning reader skips it even in the window before cleanup.

code

go · 9 lines
go
f, err := os.CreateTemp(filepath.Dir(path), ".config-*.tmp")
if err != nil {
	return err
}
tmp := f.Name() // os.CreateTemp chose the random part
defer func() {
	f.Close()      // backstop for the failure paths
	os.Remove(tmp) // already gone if the rename succeeded
}()

go deeper

for a junior

Remember that os.CreateTemp creates a real file straight away, so anything that returns early afterwards leaves it on disk unless you removed it.

for a middle

Explain why the cleanup is a single deferred os.Remove registered right after the create, and why that same line does nothing harmful once the rename has succeeded.

for a senior

Connect the litter to a real incident: scratch files in the directory the consumer scans, or inode exhaustion from a retry loop, plus what you would sweep at startup after shipping the bug.

for a principal

Set the contract for a shared publish directory - who may write there, what filename shapes are reserved for scratch, and whether consumers are allowed to scan it at all rather than open one known path.

## What the failure path leaves behind The atomic-publish idiom starts by creating a scratch file: `os.CreateTemp(filepath.Dir(target), ".config-*.tmp")`. That call has a side effect on the filesystem - a new, empty, uniquely named file exists from that moment. Everything after it can fail: the write, the sync, the close, the rename. Each of those failures is normally handled by returning an error up to a caller that will retry later. If nothing removes the scratch file, each failed attempt leaves one behind. And it is not left in some tidy temp area, because the whole point of the idiom is that the scratch file must live on the same filesystem as the target - in practice, in the target's own directory. So the litter accumulates exactly where the consumer is looking. ## Why the litter is more than untidy Think of a controller that rewrites a small config file that a sidecar process picks up. Two things go wrong: 1. **A reader that scans the directory.** Plenty of consumers load every file in a config directory rather than one exact path. A leftover `.config-9182.tmp` from a failed attempt is a half-written document, and it gets parsed. Now a failure that should have been invisible becomes an outage, and the postmortem is about a file nobody meant to publish. 2. **Unbounded growth.** Each attempt is a fresh unique name, so retries do not overwrite each other. A tight reconcile loop failing on a full disk produces thousands of tiny files and burns inodes - and the condition that caused the failures is now harder to clear. ## The cleanup idiom ```go f, err := os.CreateTemp(filepath.Dir(path), ".config-*.tmp") if err != nil { return err } defer os.Remove(f.Name()) ``` Three properties make this correct: - **It is registered immediately.** The defer goes on the line after the create, before any code that can return, so there is no path out of the function that skips it. - **It is safe on success.** After `os.Rename(f.Name(), path)` succeeds, the temp name no longer exists. `os.Remove` then returns a `*fs.PathError` wrapping `fs.ErrNotExist`, which you ignore. Nothing else is at risk, because the argument is the temp name and the published target has a different name. - **It uses `f.Name()`.** `os.CreateTemp` substitutes a random string for the `*` in your pattern, so the actual path is only knowable from the returned file. Reconstructing the name yourself is how people end up removing the wrong thing, or nothing. The file handle deserves the same treatment. On the success path you close the temp file explicitly and check that error before renaming, because a buffered write problem can surface at close. On the failure paths you still need it closed, so a deferred close as a backstop is reasonable - a second close on an already-closed `*os.File` simply returns an error, which the backstop ignores. What matters is that the explicit close on the success path is the one whose error you act on. ## Naming the pattern defensively Cleanup covers the interval after a failure, but there is still a window while the temp file is being written in which it exists and is incomplete. Choose the pattern so that window is harmless: a leading dot hides it from shell globs and from readers that skip dotfiles, and a `.tmp` suffix makes a directory-scanning consumer's filter easy to write and easy to justify in review. The best readers open the exact path they want and never scan; the pattern is defence for the ones that do not. ## Cleaning up what earlier versions left If a service has been running the buggy version, the directory already holds junk that no running process will ever remove. A sweep at startup - list the directory, delete entries matching the temp pattern that are older than some age - is a reasonable belt-and-braces addition, and the age check keeps it from deleting a scratch file that another instance is writing right now. Treat it as recovery from a past bug, not as the primary mechanism: the deferred remove is what keeps the steady state clean. ## What an interviewer is listening for That you recognise `os.CreateTemp` as having a filesystem side effect the moment it returns; that the cleanup is registered right there rather than repeated on each error branch; that you can explain why the same line is harmless after a successful rename; and that you noticed the temp file is sitting in the directory your consumer reads.

  • Why is defer os.Remove(f.Name()) harmless after the rename has already succeeded?
    The rename consumed the temp name, so os.Remove returns a *fs.PathError wrapping fs.ErrNotExist and changes nothing. The published file is safe because the argument is the temp name, which os.CreateTemp made unique, and never the target path.
  • Why call f.Name() instead of building the temp path yourself?
    os.CreateTemp replaces the * in your pattern with a random string it chooses, so the real path exists only on the returned file. Composing your own guess means the cleanup removes nothing, or worse, removes some other file that happens to match the name you invented.
  • How would you deal with temp files a previous buggy release already left in the directory?
    Sweep them at startup: list the directory, and delete entries matching the temp pattern that are older than a generous age. The age check prevents deleting a scratch file another instance is writing right now. Treat the sweep as recovery from a past bug; the deferred remove keeps the steady state clean.

saying these in an interview costs you the question

  • Repeats the remove on every error branch and misses one
  • Removes the target path instead of the temp name
  • Thinks a failed create still needs cleaning up
  • Reconstructs the temp path instead of using f.Name()
  • Leaves scratch files where the consumer scans the directory