skip to content

Atomic Writes and Renames

Write to a temporary file beside the target, force it to disk, then rename over the original so no reader ever sees half a file. Interviewers ask after a crash has left a truncated config behind.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How does the write-to-temp-then-os.replace pattern make a file update atomic?

level: middleimportance: must knowfreq 45%

answer

  1. Never modify the destination in place
  2. Build the new version beside the old one
  3. One call does the visible swap
  4. Same directory keeps it on one filesystem
  5. os.replace rebinds the directory entry

basics

~20 s

The new contents go into a temporary file in the target's own directory; os.replace(tmp, target) then swaps the directory entry in one step, so a reader sees either the whole old file or the whole new one.

solid answer

~50 s

Create the temporary file with `tempfile.mkstemp(dir=...)` pointed at the target's own directory so it lands on the same filesystem, wrap the returned descriptor with `os.fdopen`, write it, close it, then call `os.replace(tmp, target)`. `os.replace` performs a rename that overwrites the destination, and a rename inside one filesystem is atomic with respect to other processes: the name points at the old file or the new one, never at something half-written. A process that already had the old file open keeps reading it until it closes - nothing is pulled out from under it. Two details bite in practice: `mkstemp` hands you ownership of the temp file, so you must unlink it yourself if the write fails, and it creates the file mode 0600, so `os.chmod` it before the swap if the target must be readable by others.

code

python · 21 lines
python
import json
import os
import tempfile


def write_atomic(path, payload):
    directory = os.path.dirname(os.path.abspath(path))
    fd, tmp = tempfile.mkstemp(dir=directory, prefix=".routes-", suffix=".json")
    try:
        with os.fdopen(fd, "w", encoding="utf-8") as f:
            json.dump(payload, f)
    except BaseException:
        os.unlink(tmp)
        raise
    os.replace(tmp, path)


target = os.path.join(tempfile.mkdtemp(), "routes.json")
write_atomic(target, {"services": 17})
with open(target, encoding="utf-8") as f:
    print(f.read())

go deeper

for a junior

Learn the shape before the details: write the new contents somewhere else, then rename that file over the target. Knowing that a rename replaces a name in one step is enough at this level.

for a middle

Be able to write the helper from memory - temp file in the target's directory, write, close, os.replace - and to say precisely which step is the atomic one and why the others do not need to be.

for a senior

Show the operational details: cleanup of strays on failure, the 0600 mode a temp file starts with, readers keeping the old file open, and the fact that atomicity is not durability.

for a principal

An interviewer expects you to argue for one shared implementation rather than each service rolling its own, and to say which files deserve the treatment - state and config yes, high-rate telemetry probably not.

## The recipe 1. Determine the target's parent directory, typically `os.path.dirname(os.path.abspath(path))`. 2. Create the temporary file **inside that directory** with `tempfile.mkstemp(dir=directory)`, which returns an open descriptor and the path it created. 3. Wrap the descriptor with `os.fdopen` and write the complete new contents. 4. Fix the permission bits with `os.chmod` if the target needs to be readable by anyone but the owner. 5. Close the file. 6. Call `os.replace(tmp, path)`. 7. On any failure between steps 2 and 6, unlink the temporary file and re-raise. ## Why the rename is the atomic part A directory entry is a name bound to a file. Renaming over an existing name rebinds that one entry, and the filesystem performs the rebinding as a single indivisible update: any process resolving the path either gets the old binding or the new one. Nothing else in the sequence needs to be atomic, because until step 6 the new contents live under a name nobody is reading. All the slow, failure-prone work - serializing, encoding, writing megabytes - happens off to the side, and only the instantaneous rebinding is exposed. `os.replace` is the call to use rather than `os.rename`, because it overwrites an existing destination on every supported platform. `os.rename` does the same on Unix but refuses on Windows when the destination exists, so code that uses it and is only tested on Linux fails the first time it runs on Windows. ## Why the temporary file goes in the target's directory The new entry has to be created in the target's parent directory, so a temporary file created in that same directory is guaranteed to be on the same filesystem - no device probing, no comparing mount tables. A temporary file in the system temp directory very often is not, and the rename then fails outright rather than degrading gracefully. ## What happens to processes holding the old file A reader that opened the path before the swap keeps its descriptor, and that descriptor still refers to the old file. It reads the old contents to the end, undisturbed, and the storage is reclaimed once it closes. This is why the pattern is safe for long-running readers and why such readers must reopen the path to notice a new version. It is also why there is no window in which a reader gets an error: the old file is never removed while someone is using it. ## The details that bite **Cleanup.** `tempfile.mkstemp` deliberately does not delete anything for you - it returns a descriptor and a path and hands you ownership. If serialization raises after the file exists, an aborted run leaves a stray file behind, and a job that retries every few minutes accumulates them in the target's directory. Wrap the write in `try`/`except`, unlink on failure, and re-raise. A leading dot or a distinctive prefix in the temp name makes strays easy to spot and keeps them out of a glob that scans the directory. **Permissions.** `mkstemp` creates the file readable and writable by the owner only. That is the right default for a temp file and the wrong result for, say, a configuration file another account has to read - after the swap the file carries the temp file's mode, not the previous target's. Set the mode explicitly before the rename if it matters. **Metadata.** The replacement is a different file: inode number, creation time and hard links to the old name do not carry over. Anything watching the old file by identity rather than by name will be surprised. ## What the pattern does not give you It does not give you **durability**: after `os.replace` returns, the swap can still be sitting in the kernel's cache and a power loss can undo it, which is what `os.fsync` is for. It does not give you **mutual exclusion**: two processes replacing the same path both succeed and the last rename silently wins, so if losing an update matters you need a marker file created with `os.O_CREAT | os.O_EXCL`, or a single writer. And it is per-file - replacing three files that must change together is three separate atomic swaps, and a reader can observe the intermediate combination. A route-optimisation job publishing a 17-service dependency graph gets exactly what it needs from this: the graph file always parses, whichever version a consumer happens to load.

  • What happens to a process that already had the old file open when os.replace runs?
    Nothing visible to it. Its descriptor still refers to the old file, so it reads the old contents to the end and the storage is freed only when it closes. The rename changes which file the *name* resolves to, not what an existing descriptor points at. That is why a long-lived reader has to reopen the path to pick up a new version, and why it can never observe a mixture of the two.
  • Why must you unlink the temporary file yourself when the write fails?
    `tempfile.mkstemp` gives you a descriptor and a path and leaves the file's lifetime entirely to you - it is not removed at close or at interpreter exit. If serialization raises after the file exists, the stray stays in the target's directory, and a job retrying every few minutes accumulates them. Wrap the write in try/except, unlink, and re-raise.
  • Does the swap protect you from two writers replacing the same file at once?
    No. The rename protects readers, not writers: both processes succeed and the last rename wins, silently discarding the other's contents. If losing an update is unacceptable you need explicit exclusion - a marker file created with `os.O_CREAT | os.O_EXCL`, or funnelling all writes through one process. Atomicity here means no partial file, not mutual exclusion.

saying these in an interview costs you the question

  • Removes the target first, then renames the temp file in
  • Creates the temporary file in the system temp directory
  • Uses os.rename over an existing path and calls it portable
  • Thinks the rename also guarantees survival of a power cut
  • Leaves the temporary file behind when the write raises
  • Expects readers holding the old file to get an error

context

open as a page

Why can open(path, 'w') expose an empty or half-written file to readers?

level: juniorimportance: should knowfreq 30%

basics

~20 s

Mode 'w' truncates the file the instant it is opened, before a single byte is written, and later writes reach it piecemeal as buffers flush. A reader that opens the path meanwhile sees an empty or half-written file.

open as a page

Why does os.replace raise OSError with errno EXDEV when the temp file is in /tmp?

level: middleimportance: should knowfreq 28%

basics

~20 s

A rename only rewrites a directory entry inside one filesystem; it cannot move data between filesystems. When /tmp is a separate mount from the target, os.replace fails with EXDEV. Create the temporary file in the target's own directory instead.

open as a page

After os.replace returns, why is os.fsync on the containing directory still needed?

level: seniorimportance: should knowfreq 30%

basics

~20 s

os.replace makes the swap visible to other processes at once, but the new directory entry may still live only in the kernel's cache. A power loss can revert the name, so durability needs os.fsync on the file and on its directory.

open as a page