You deleted a multi-gigabyte log file on a Linux server, but df still reports the space as used and the filesystem stays full. What is the kernel doing, and how do you get the space back?
answer
- the directory entry is not the file
- two references must both drop
- df and du disagree for a reason
- the holder is visible under /proc
- truncate through the descriptor
basics
~20 sA process still holds the file open. Deleting removed the directory entry and dropped the link count to zero, but the kernel frees an inode and its blocks only when no open file descriptor references it either, so the space stays allocated until that process closes it.
solid answer
~50 s`rm` does not delete data; it calls `unlink()`, which removes a directory entry and decrements the inode's link count. The kernel releases the inode and its data blocks only when **both** counters reach zero: the on-disk link count and the number of open file descriptions referencing it. A logger that still has the file open keeps the blocks alive, so `df` — which reports allocated blocks — still shows them, while `du`, which walks names, no longer sees the file at all. That `df`/`du` divergence is the diagnostic tell. The kernel exposes the holders under `/proc/<pid>/fd`, where the symlink target reads with a `(deleted)` suffix. To recover the space, make the holder let go: restart it, or signal a daemon that reopens its logs. If you cannot restart, truncate through the descriptor with `: > /proc/<pid>/fd/<n>`, which frees the blocks immediately.
go deeper
Know that removing a file only removes its name, and that a process which still has it open keeps the data — and therefore the disk space — alive until it closes or exits.
Explain the two reference counts the kernel checks before reclaiming an inode, and why df sees the blocks while du, which walks names, does not.
Drive the diagnosis end to end: spot the df/du gap, identify the holder through /proc/<pid>/fd, choose between restarting, signalling a reopen, or truncating through the descriptor, and say what each costs.
Own the prevention: rotation policy that never depends on an operator deleting a live log, size-based caps and volume budgets per service, and a standard for whether services must handle a reopen signal or be rotated by truncation.
## Deletion is reference counting, not erasure A file on a Unix filesystem is an inode. Names are directory entries pointing at it, and the inode carries a **link count** of how many such entries exist. `rm` calls `unlink()`: remove one entry, decrement the count. If the count is still above zero, the file plainly still exists under its other names. When the count reaches zero, the file has no name — but that is not sufficient to free it. Every open file description held by any process is also a reference. The rule the kernel enforces is: reclaim the inode and its data blocks when the link count is zero **and** no open file description refers to it. Until then the file is nameless yet fully alive. Its readers can keep reading, its writer can keep appending, and its blocks remain allocated. This is deliberate and load-bearing. It is why you can replace a running program's files during an upgrade without killing the process, why a temporary file created and immediately unlinked cannot leak on the filesystem even if the process crashes, and why no reader of an open file ever gets the rug pulled out from under it. ## Why df and du disagree `df` asks the filesystem how many blocks are allocated. The deleted-but-open file's blocks are allocated, so they are counted. `du` walks the directory tree and sums the files it finds by name — and the file has no name, so `du` cannot see it. A large, unexplained gap between the two on a full filesystem is the signature of this situation, and recognising that gap is most of the answer an interviewer is listening for. ## Finding what still holds it The kernel publishes each process's descriptors as symlinks in `/proc/<pid>/fd/`. For a file whose last name is gone, reading that symlink shows the path it had, with ` (deleted)` appended: ``` /proc/2417/fd/3 -> /var/log/app.log (deleted) ``` That gives you the PID, the descriptor number and the original path in one place. `/proc/<pid>/fdinfo/<n>` additionally shows the current file offset and the open flags, which matters for what comes next. ## Getting the space back The honest fix is to make the holder release the descriptor: - **Restart the process.** All its descriptors close, the last reference disappears, and the blocks are freed instantly. - **Signal a daemon that reopens its logs.** Many long-running services close and reopen their log files on `SIGHUP`; that is the mechanism log rotation depends on, and it drops the old inode as a side effect. If restarting is not acceptable, you can empty the file through the descriptor without closing it: ```sh : > /proc/2417/fd/3 ``` Opening that `/proc` entry with truncation truncates the underlying inode, so the blocks are released while the descriptor stays valid and the process keeps running. There is a nuance: if the writer opened the file with `O_APPEND` it will simply continue at the new end and everything is fine, but a writer that tracks its own offset resumes at the old, now enormous offset. The result is a **sparse** file whose apparent size is huge again while consuming almost no blocks — confusing to look at, but not itself a space problem. Nothing recovers the log lines already written; you are trading the content for the space. What does **not** help is worth knowing too: `sync` only flushes dirty data, and dropping caches releases cached metadata rather than allocated blocks. Neither touches a live reference. ## The design lesson The same semantics you are debugging here are what makes safe log rotation possible in the first place. Rotation by rename plus reopen works because the rename never disturbed the writer's descriptor and the reopen is what finally releases the old inode. Rotation with copy-then-truncate keeps the same inode deliberately, so no reopen is needed at all — at the cost of losing anything written between the copy and the truncate. Choosing between them is a real production decision, and both choices follow directly from the fact that a descriptor references an inode while a name does not. The long-term fix for a host that keeps hitting this is usually not sharper tooling but removing the human step: rotate on size as well as time, cap total log volume, and make sure every service either handles the reopen signal or is configured for the truncate style, so nobody ever deletes a live log by hand.
- Why can you replace a running program's files on Linux without crashing the process?Because the running process holds references to the old inodes, which survive losing their names. A package upgrade that unlinks the old file and creates a new one leaves the process mapped to the old, now-nameless inode until it exits. Overwriting the file *in place* is the unsafe variant — the kernel returns `ETXTBSY` when you try to open a running executable for writing, precisely to stop that.
- How does logrotate's copytruncate option differ from rotating by rename, and what does it risk?Rename-and-signal moves the name and relies on the daemon reopening; the old inode is released when it does. `copytruncate` copies the content aside and then truncates the original in place, so the inode and the daemon's descriptor never change and no signal is needed. The risk is a race: anything written between the copy and the truncate is lost, and a writer without `O_APPEND` resumes at its old offset, leaving a sparse hole.
- Why is creating a temporary file and unlinking it immediately a common pattern?The file stays fully usable through the descriptor while having no name, so no other process can open it by path and no leftover file survives a crash — the kernel frees it when the last descriptor closes. It is the classic self-cleaning scratch file; the tradeoff is that nothing can find it by name, which is why it suits private scratch space rather than anything another process must read.
saying these in an interview costs you the question
- Says rm frees the blocks immediately in all cases
- Blames the discrepancy on df caching stale numbers
- Suggests sync or dropping caches to reclaim the space
- Thinks deleting a file invalidates open descriptors
- Reaches for fsck on a mounted, healthy filesystem