skip to content

After os.replace returns, why is os.fsync on the containing directory still needed?

level: seniorimportance: should knowfreq 30%

answer

  1. Visible to processes is not the same as safe
  2. Two different guarantees, two different mechanisms
  3. The new name lives in the parent directory
  4. Sync the data before the swap
  5. os.open the directory read-only, then os.fsync

basics

~20 s

os.replace makes the swap visible to other processes at once, but the new directory entry may still live only in the kernel's cache. A power loss can revert the name, so durability needs os.fsync on the file and on its directory.

solid answer

~50 s

Atomicity and durability are different guarantees. `os.replace` gives atomicity - no process ever resolves the name to a half-written file - but the rename is recorded in the kernel's cache first and reaches the device whenever the filesystem gets around to it. To make the new version survive a power cut you need two syncs: `os.fsync(f.fileno())` on the temporary file **before** the rename, so its bytes are on the device, and an `os.fsync` on a descriptor for the parent directory **after** the rename, so the new entry itself is recorded. The order matters: syncing the data only after the rename risks a crash leaving the name attached to a file whose contents never landed. Directories cannot be opened with builtin `open()`, so use `os.open(directory, os.O_RDONLY)`, sync that descriptor, and close it. Both syncs cost a device round-trip, so spend them where losing the last update is a correctness problem.

code

python · 19 lines
python
import os
import tempfile

directory = tempfile.mkdtemp()
target = os.path.join(directory, "plan.bin")

fd, tmp = tempfile.mkstemp(dir=directory)
with os.fdopen(fd, "wb") as f:
    f.write(b"17 services")
    f.flush()
    os.fsync(f.fileno())      # the new file's bytes are on the device
os.replace(tmp, target)       # the swap is visible, but only in the cache

dir_fd = os.open(directory, os.O_RDONLY)
try:
    os.fsync(dir_fd)          # now the directory entry is durable too
finally:
    os.close(dir_fd)
print(os.path.exists(target))

go deeper

for a junior

Take away the distinction: a rename decides what other programs can read right now, while fsync decides what survives a power cut. You are not expected to write the sync sequence yet.

for a middle

Explain that flush() only reaches the operating system while os.fsync reaches the device, and be able to place both syncs correctly around the rename rather than reciting them as a pair.

for a senior

Demonstrate the failure modes: data synced after the rename gives an empty file, no directory sync gives a reverted name. Be ready to say which files in a service deserve the cost.

for a principal

An interviewer expects a position on durability as policy - what the service promises about the last acknowledged update, which storage that promise depends on, and where paying per-write syncs would be the wrong tradeoff.

## Two guarantees people run together **Atomicity** is about what other processes can observe: with the temp-file-then-rename pattern, a concurrent reader resolves the name either to the complete old file or to the complete new one. **Durability** is about what survives the machine losing power: whether the change is on the storage device or still only in memory. The rename buys the first outright. It buys none of the second, because like almost every filesystem operation it is applied to in-memory state and written to the device later, in whatever order the filesystem prefers. So after `os.replace` returns successfully there are three things that may still be only in memory: the new file's contents, the new directory entry that names it, and the removal of the old entry. A power loss at that moment can leave the directory looking exactly as it did before. ## The two syncs and their order ```python fd, tmp = tempfile.mkstemp(dir=directory) with os.fdopen(fd, "wb") as f: f.write(payload) f.flush() # Python buffer to the operating system os.fsync(f.fileno()) # operating system to the device os.replace(tmp, target) dir_fd = os.open(directory, os.O_RDONLY) try: os.fsync(dir_fd) finally: os.close(dir_fd) ``` **`flush()` is not `fsync`.** Flushing pushes bytes out of the file object's userspace buffer into the operating system. `os.fsync` asks the operating system to push its own cached copy down to the storage device and, on most platforms, to wait for it. Skipping the flush before the fsync means syncing a file the kernel has not been told about yet. **Sync the data before the rename.** If the rename reaches the device before the data does, a crash in between leaves a perfectly valid directory entry pointing at a file whose blocks were never written - the classic zero-length-file-after-a-crash report. Syncing first makes the ordering explicit rather than relying on a filesystem's own heuristics, which vary between filesystems and mount options. **Sync the directory after the rename.** The entry lives in the parent directory, so making the entry durable means syncing the directory itself. Builtin `open()` refuses a directory, so obtain a descriptor with the low-level `os.open(directory, os.O_RDONLY)`, pass it to `os.fsync`, and close it in a `finally`. You never read from it. This is a Unix technique; on Windows you cannot get a directory handle this way, and durability there is handled differently by the filesystem. ## What each sync buys you concretely * Data sync only: after a crash the name may still resolve to the old file, but you never see corruption. * Data sync plus directory sync: after a crash the name resolves to the new file, and its contents are complete. * Directory sync without a data sync: the worst of the options - the name may point at a file with no contents. Notice that even the full sequence does not promise *which* version you get after a crash; it promises that whichever one you get is whole and that once the directory sync returns, the new one is committed. ## The cost, and when to skip it `os.fsync` is expensive. It is a round-trip to the device, it can be milliseconds even on fast storage, and on a shared volume it can stall other writers. A route-optimisation job rewriting a 17-service dependency graph every few minutes should pay it without hesitation. The same job writing per-request scratch files thousands of times a minute should not - the syncs would dominate its runtime for data a restart would recompute anyway. The useful split is by what the file is for. Files whose loss changes correctness - state that cannot be recomputed, a checkpoint, an acknowledgement someone else relies on - get both syncs. Files that are caches or derived artefacts get the rename, which is nearly free and prevents corruption, and skip the syncs. The rename is about your readers; the syncs are about your recovery story, and the two decisions are independent. ## Diagnosing the absence The symptom of missing syncs is a file that is *old* after a hard reboot rather than corrupt: the service restarts with the previous version of its state and nobody can explain where the last update went, because the application logged a successful write. Reproducing it requires a real power cut or an equivalent - a clean process kill will not do it, because a kill does not discard the kernel's cache. That gap between how the bug is reported and how it is reproduced is why it survives so long in production systems.

  • Where exactly in the sequence do the two fsync calls belong?
    Flush the file object's buffer and call `os.fsync(f.fileno())` while the temporary file is still open, then close it, then `os.replace`, then open the parent directory read-only and `os.fsync` that descriptor before closing it. Syncing the data before the rename is what prevents a crash from leaving the name attached to a file whose blocks never reached the device.
  • How do you fsync a directory when open() will not open one?
    Use the low-level `os.open(directory, os.O_RDONLY)` to get a raw descriptor, pass it to `os.fsync`, and close it in a `finally`. You never read from it, so there is no need to wrap it with `os.fdopen`. This is a Unix technique - on Windows you cannot obtain a directory handle this way, and directory syncing is not part of the recipe there.
  • When is skipping fsync the right engineering call?
    When the file is a cache or a derived artefact that a restart can rebuild. An fsync is a device round-trip and on a shared volume can stall unrelated writers, so paying it on every update in a hot path is often worse than losing the last write. Keep the rename regardless - it costs almost nothing and prevents corruption - and add the syncs only where losing the newest committed version breaks correctness.

saying these in an interview costs you the question

  • Thinks os.replace alone guarantees survival of a power cut
  • Syncs the file after the rename rather than before
  • Assumes closing a file forces it to the physical device
  • Never syncs the directory, so the new name can vanish
  • Confuses flushing the Python buffer with an fsync
  • Calls fsync on every small write without weighing the cost

context