In Linux, what does a file's inode number identify, and why does copying a file with cp produce a new inode number while renaming it with mv inside the same filesystem does not?
answer
- identity lives below the name
- the number alone is not unique
- rename only edits directory entries
- create allocates a fresh inode
- descriptors follow the inode, not the path
basics
~20 sAn inode number identifies a file within one filesystem, not across the machine. Renaming rewrites only a directory entry, so the inode is untouched; copying allocates a fresh inode and new data blocks, producing a different file that merely holds the same bytes.
solid answer
~50 sThe inode number is the file's identity inside its filesystem — the number a directory entry stores next to the name. It is unique only per filesystem, so the true identifier of a file is the pair (device, inode), which is what `stat` reports as `Device` and `Inode` and what tools compare when they ask "is this the same file?". `mv` within one filesystem is a `rename()` syscall: it unlinks a name from one directory and adds it to another, leaving the inode, its data blocks, its permissions, its open file descriptors and any other hard links completely untouched. `cp` is the opposite — it creates a new file, so a new inode is allocated, blocks are written, and the result is a distinct file that happens to hold identical bytes. `mv` **across** filesystems cannot rename, so it degrades to copy-then-unlink and the inode number changes there too.
go deeper
Know that the inode number is the file's real identity and that renaming a file does not copy or rewrite it, while copying makes a genuinely separate file.
Explain rename() as a directory-entry edit that leaves the inode, its blocks, its other names and its open descriptors untouched, and contrast that with creation allocating a new inode.
Bring the operational consequences: rotation that renames needs a reopen signal, editors that write-and-rename break hard links and descriptor-following tails, and same-file checks must compare device plus inode.
Own the pattern as a design tool — atomic publish by rename, temp-file-plus-rename for crash safety, and the rules your tooling must follow so identity checks and backup deduplication stay correct across mount boundaries.
## What the number actually is Each filesystem maintains its own numbered table of inodes. A directory entry is a (name, inode number) pair, and `ls -i` or `stat -c '%i'` prints that number. Because the numbering is per filesystem, the same inode number will regularly appear on `/` and on `/home` for two entirely unrelated files. That is why identity comparisons use the pair of device ID and inode number: ```sh $ stat -c '%d %i %n' a.txt b.txt 2049 131074 a.txt 2049 131074 b.txt # same device, same inode -> hard links to one file ``` A program does this with `stat()` and comparing `st_dev` and `st_ino`; on the command line `find -samefile` does it for you. Inode numbers are also reused: once an inode is freed, the filesystem may hand the same number to the next file created. So an inode number is a *current* identity, not a durable name you can store and come back to weeks later. ## Rename moves the name, not the file `mv old new` inside a single filesystem is the `rename()` syscall. It edits directories: the entry disappears from the source directory and appears in the destination directory, still carrying the same inode number. Nothing about the file object changes. Consequences worth being able to state out loud: - Data is not read or written, so renaming a 40 GB file is instant regardless of size. - Open file descriptors keep working and keep writing to the same inode. A daemon logging to a descriptor it opened as `/var/log/app.log` continues writing after that path is renamed to `app.log.1` — it is writing to the inode, and the descriptor never consulted the name again. - Other hard links to the inode are unaffected; the link count does not change, because one entry was removed and one added. - Permissions, owner, size and `mtime` are preserved. The inode's `ctime` (status-change time) is updated, and both affected directories get new `mtime` and `ctime`, because their contents changed. - Rename onto an existing name is atomic: `rename()` replaces the destination entry in one step, and nobody ever observes the path missing. That is the basis of the write-temp-then-rename pattern used for safe config updates. ## Copy creates a different file `cp src dst` opens the source, creates the destination, and writes bytes into it. Creation allocates a new inode, so the destination has a different inode number, a fresh link count of 1, timestamps of now, and ownership by the copying user with permissions filtered through the umask. `cp -p` (or `-a`) asks for mode, ownership and timestamps to be preserved, but even then the inode number is new — preservation copies *attributes*, it does not share the inode. `cp -l` makes hard links instead of copies, and `cp --reflink` on a copy-on-write filesystem shares the underlying blocks while still creating a separate inode. Because a copy is a new file, anything holding the old inode open keeps seeing the old content, and any other hard link to the source is unaffected. ## Where this bites people The classic surprise is editing. Many editors do not write into the existing inode; they write a temporary file and rename it over the target. That is deliberately safe against a crash mid-write, but it means the file gets a **new inode**: hard links to the old inode are silently orphaned from your edit, and any process tailing the file by descriptor keeps reading the old, now-nameless inode. This is exactly why `tail -f` (follow the descriptor) and `tail -F` (follow the name, reopen when it is replaced) exist as separate options. The second is log rotation. Renaming a log file leaves the writing daemon happily appending to the same inode, now called `app.log.1`, and the newly created `app.log` stays empty forever. That is why rotation must either signal the daemon to reopen its logs after the rename, or use copy-then-truncate to keep the same inode in place. Both approaches are consequences of the same fact: the descriptor points at the inode, and the name is not part of it. The third is scripting identity checks. Comparing paths does not tell you whether two names are the same file, and comparing inode numbers alone across mount points produces false positives. Compare the device and inode together, or ask `find -samefile`.
- After a log file is renamed on the same filesystem, why must the daemon be told to reopen it?Because the daemon's open descriptor refers to the inode, not the path. The rename moved the name; the descriptor is unaffected, so the daemon keeps appending to the same inode now visible as the rotated file, while the freshly created log path stays empty. Signalling the daemon (typically SIGHUP for daemons that implement it) makes it close and reopen by name, picking up the new inode.
- Which timestamps change when you rename a file, and which do not?The file's `mtime` does not change, because its data was not modified. Its `ctime` does, because inode status changed. Both the source and destination directories get updated `mtime` and `ctime`, since their contents were edited. This is a useful tell: an unchanged `mtime` with a recent `ctime` usually means a rename, a `chmod`, a `chown`, or a link-count change rather than a write.
- How can a script reliably decide whether two paths refer to the same file?Compare the device ID and the inode number together, not the paths and not the inode alone — inode numbers are unique only within one filesystem. From a shell, `find -samefile` or comparing `stat -c '%d:%i'` on both paths does it; from code, `stat()` both and compare `st_dev` with `st_ino`. Resolving both paths and string-comparing them fails as soon as hard links are involved.
saying these in an interview costs you the question
- Treats an inode number as unique across the whole machine
- Says mv rewrites the file's data on the same filesystem
- Believes cp -p keeps the original inode number
- Thinks an open descriptor follows the pathname after a rename
- Assumes editing a file always keeps its inode