How does os.SameFile let a backup scanner notice it already archived a file under another name?
answer
- a path is a name, not an identity
- two names can reach one file
- size and modtime prove nothing about sameness
- the filesystem already assigns an id
- only FileInfo values from the os package qualify
basics
~20 sos.SameFile compares two FileInfo values by the file's underlying identity — device and inode on Unix — rather than by path, size or timestamp. A scanner keeps the FileInfos it has archived and skips any new one that SameFile matches.
solid answer
~50 s`os.SameFile(fi1, fi2 fs.FileInfo) bool` answers "are these two descriptions the same file on disk?" It looks past the path at the identity the filesystem assigns — device plus inode on Unix, volume serial plus file index on Windows — so it returns true for two hard links, or for a real path and a symlink that was resolved to it, and false for two byte-identical copies. That is exactly the check a backup scanner needs after it chooses to follow a link: keep the `FileInfo` of everything already archived, and before writing an entry, `SameFile` the candidate against that set. One constraint decides whether the technique works at all: `SameFile` is only meaningful for `FileInfo` values produced by the `os` package — `os.Stat`, `os.Lstat` or `File.Stat`. Hand it a `FileInfo` from an embedded filesystem or any other `fs.FS` implementation and it simply returns false, which fails silently in the direction of duplicating everything.
code
go · 13 lines// seen holds FileInfo values from os.Stat for everything archived so far.
func alreadyArchived(seen []fs.FileInfo, name string) (bool, error) {
fi, err := os.Stat(name) // resolved identity of the target
if err != nil {
return false, err
}
for _, s := range seen {
if os.SameFile(s, fi) {
return true, nil
}
}
return false, nil
}go deeper
Know that os.SameFile takes two FileInfo values and reports whether they describe the same underlying file, regardless of the paths they came from.
Explain what identity means here — device and inode on Unix — and why hard links and followed symbolic links make two names for one file while a copy is a genuinely different file.
Demonstrate the operational constraints: the values must come from the os package, identity is valid only within a run because inodes are reused, and the linear check must be bounded to the entries that need it.
Own what your tool promises about duplicates and links: whether a followed link is archived, deduplicated or refused, what a restore is expected to rebuild, and where that contract is written down for the team.
## The problem it solves A path is a name, not an identity. The same file can be reached through several names — a hard link creates a second directory entry for one file, a symbolic link that a tool chose to follow lands on a file that has its own name too, and a bind mount or a case-insensitive filesystem can make two different strings reach one object. A tool that decides "have I seen this already?" by comparing cleaned paths will answer wrongly in all of those cases. Comparing metadata is no better. Two files with equal `Size()` and equal `ModTime()` are routinely distinct files: a copy has both. Equality of name, size and time is a heuristic for *change*, never a proof of *identity*. `os.SameFile(fi1, fi2 fs.FileInfo) bool` is the real answer. It compares the identity the filesystem itself assigns: on Unix the device number and inode number, on Windows the volume serial number and file index. Two names, one identity, `true`. ## The constraint people miss `SameFile` only understands `FileInfo` values that the `os` package produced — from `os.Stat`, `os.Lstat` or `(*os.File).Stat`. That is because it reaches through `Sys()` for the platform record. Give it a `fs.FileInfo` from an `embed.FS`, a `fstest.MapFS`, a tar reader or any other `fs.FS` implementation and it returns **false** rather than an error. In a scanner that means every entry looks new, and the bug shows up as a silently doubled archive rather than a crash. If your code walks an abstract `fs.FS`, identity is simply not available and you need a different dedup key. Identity is also only stable *within a run*. Inode numbers are reused after a file is deleted, device numbers change across remounts, and some network filesystems do not provide stable file ids at all. Persisting a device/inode pair between runs and treating it as a permanent identifier is a real defect; use it as a within-scan set and store content hashes if you need identity across time. ## Wiring it into a scanner The scanner's rule is a pair: `Lstat` every entry so links announce themselves, and if the tool has decided to follow a link, `Stat` it and check the resolved file against everything already archived. A linear `SameFile` scan over a slice of seen `FileInfo` values is O(n) per candidate, which is fine for the handful of followed links a normal tree contains and quadratic if you apply it to every file in a million-file tree. When you do need it at scale, key a map on the platform identity directly — assert `fi.Sys().(*syscall.Stat_t)` on Unix and use `{Dev, Ino}` as a comparable struct key — and accept that this is a build-tagged, Unix-only path. `os.SameFile` is the portable, low-volume version of the same idea. What the scanner does on a hit is a product decision worth stating explicitly: archive the content once and record the second name as a link to the first, so a restore rebuilds the tree's real shape rather than materialising two full copies. ## The failure it prevents, seen from the diagnostic side The symptom that sends someone to this function is an archive that is measurably larger than the tree it came from, with a file appearing twice under two names. The diagnosis is a three-line dump for the suspicious name: the `os.Stat` result, the `os.Lstat` result, and `os.SameFile` against the entry already recorded. If `Lstat` shows a mode beginning with `L` while `Stat` shows a regular file, and `SameFile` says true against something already archived, the whole story is on one screen — a link was followed, and the target was archived a second time under the link's name. ## What a reviewer should ask When this appears in a pull request, three questions separate a correct implementation from a plausible one. Are the `FileInfo` values being compared genuinely from the `os` package, or has someone abstracted the walk behind `fs.FS` and quietly disabled the check? Is the identity set scoped to a single run, or is it being written to a state file where inode reuse will eventually make it lie? And is the O(n) comparison bounded to followed links, or is it running against every file in the tree? Answering those is what makes the difference between a check that works and one that looks like it does.
- Two names report SameFile true and neither is a symbolic link. What happened, and what should the backup tool do?They are hard links: two directory entries pointing at one inode, with no primary and no secondary. Archive the content once and record the second name as a link to the first, so a restore rebuilds one file with two names rather than two independent copies that then drift apart.
- Why not just compare filepath.Clean-ed absolute paths instead?Cleaning normalises `.`, `..` and separators, nothing more. It cannot see that two distinct strings reach one file through a hard link, a followed symlink, a bind mount, or a case-insensitive filesystem where two spellings collide. Path equality is a sufficient condition for sameness, never a necessary one.
- Can you persist the device and inode pair between runs as a file identifier?No. Inode numbers are reused once a file is deleted, so a stored pair can later name a completely different file, and device numbers change across remounts. Treat filesystem identity as valid only within one scan; for identity across time use a content hash, or a size-plus-modtime heuristic you are explicit about.
- Your walk is written against fs.FS rather than os. What happens to os.SameFile?It returns false for every pair, because the FileInfo values did not come from the os package and carry no platform identity record. There is no error and no panic, so the dedup silently stops working. Either stat through `os` for the identity check, or choose a different key such as a content hash.
saying these in an interview costs you the question
- Compares cleaned absolute paths and calls it identity
- Uses equal size and modification time as proof of sameness
- Passes FileInfo from an fs.FS implementation to os.SameFile
- Stores a device and inode pair as a permanent file id
- Assumes only symbolic links can produce two names for one file
- Runs the O(n) comparison against every file in the tree