skip to content

On a Linux filesystem, what is the difference between a hard link and a symbolic link, and what can each one do that the other cannot?

level: juniorimportance: must knowfreq 76%

answer

  1. one file, many names
  2. the name is not in the inode
  3. a link count versus a stored path
  4. crossing filesystems is the tell
  5. only one of them can dangle

basics

~20 s

A hard link is an extra name for the same inode, so every name is equal and the data survives until the last one is removed. A symbolic link is a separate small file holding a path string, resolved at every access.

solid answer

~50 s

On Unix filesystems the file *is* the inode; a filename is just a directory entry mapping a name to an inode number. `ln target name` creates a second directory entry for the **same** inode and bumps its link count, so the two names are indistinguishable — same permissions, same owner, same data — and removing either one just decrements the count. The data goes away only when the count reaches zero. `ln -s target name` instead creates a **new** inode of type symlink whose contents are the path string you gave it; the kernel resolves that string on every lookup. That difference drives everything else: a hard link cannot cross a filesystem boundary (`link()` returns `EXDEV`) and cannot point at a directory, while a symlink can do both and can also dangle when its target is renamed or deleted. `ls -l` shows a symlink with `l` and an arrow; `ls -i` reveals that hard links share one inode number.

go deeper

for a junior

Be able to say plainly that a hard link is a second name for the same file while a symlink is a small file containing a path, and that only the symlink can break when its target disappears.

for a middle

Explain the mechanics: the directory entry maps a name to an inode number, ln increments the link count, and unlink decrements it. Know why hard links cannot cross filesystems or point at directories.

for a senior

Show judgment about when a hard link is the right tool — deduplicated backups, atomic publish, rotate-while-readers-hold-the-old-name — and name its hazard: nothing in the directory reveals that the data is shared, so an in-place edit hits every name.

for a principal

Own the consequences at scale: link-based deduplication changes what a backup or sync tool must preserve, symlink farms become the deployment contract other teams depend on, and a policy on which one your tooling emits prevents whole classes of surprise.

## The idea that makes both links obvious On a Unix-style filesystem, the object that holds a file's metadata and its pointers to data blocks is the **inode**. The inode records the file type, mode bits, owner UID and GID, size, timestamps, the link count, and where the data lives. One thing it does *not* record is the file's name. Names live in directories. A directory is itself a file whose contents are a table of entries, each pairing a name with an inode number. So "the file `/etc/hosts`" really means: look up `etc` in the root directory to get an inode, read that directory, look up `hosts` in it, get an inode number, and open that inode. Once you accept that a name is only a pointer, hard links and symlinks stop being two kinds of magic and become two obviously different things. ## A hard link is another directory entry ```sh ln /var/data/report.csv /home/kim/report.csv ``` This adds a second directory entry that names the *same inode number*, and increments the inode's link count from 1 to 2. There is no "original" and no "copy": the two paths are peers. `stat` on either one reports identical size, mode, owner and timestamps, because there is only one inode to report on. `chmod` through one path changes what the other path sees. Writing through one is visible through the other immediately — it is one file with two names. `rm` does not delete files; it calls `unlink()`, which removes a directory entry and decrements the link count. Remove one name and the other still works perfectly. The inode and its blocks are released only when the count hits zero (and nothing holds the file open). Two restrictions follow from the fact that a hard link is just an inode number: an inode number is only meaningful *within one filesystem*, so `link()` across a mount point fails with `EXDEV` ("Invalid cross-device link"). And Linux refuses hard links to directories outright, because arbitrary directory links would create cycles that make tree traversal, `rm -r` and consistency checking unsafe. The only directory hard links are the ones the kernel maintains itself — `.` inside a directory and `..` inside each of its children — which is why a freshly created empty directory has a link count of 2. ## A symbolic link is a file that contains a path ```sh ln -s /var/data/report.csv /home/kim/report.csv ``` This creates a brand-new inode, of type symlink, whose data is the literal text `/var/data/report.csv`. Nothing links the two inodes; the association is a string the kernel re-resolves on every path lookup. Because it is only text, it can name anything: a path on another filesystem, a directory, or a path that does not exist at all. That last case is a **dangling symlink** — reads fail with `ENOENT`, and `ls -l` still happily prints the arrow, so the breakage is silent until something tries to open it. A symlink can be relative (`ln -s ../data/report.csv link`), in which case it is resolved relative to the directory containing the link, not your current directory. Relative symlinks survive moving a whole tree; absolute ones survive moving the link itself. Most syscalls **follow** symlinks; a few deliberately do not. `stat` follows, `lstat` does not, which is why `ls -l` can show you the link rather than its target. `chmod` and `chown` follow the link and change the target: a symlink's own mode bits read as `lrwxrwxrwx` on Linux and are not used for access control — access is decided by the target's permissions plus the search permission on every directory along the way. ## Telling them apart, and choosing ```sh $ ls -li 131074 -rw-r--r-- 2 kim kim 14 Aug 20 10:01 a.txt 131074 -rw-r--r-- 2 kim kim 14 Aug 20 10:01 b.txt # same inode, count 2 131099 lrwxrwxrwx 1 kim kim 5 Aug 20 10:02 c.txt -> a.txt ``` The first column is the inode number: identical for hard links, different for the symlink. The number after the mode is the link count. In practice symlinks are the default choice — they are visible, they cross filesystems, they can point at directories, and their intent is readable. Hard links are for cases where you need a second name that is genuinely the same file and must survive the first name disappearing: de-duplicating backup trees, atomic publish patterns, and rotating a file while readers keep the old name. The cost of a hard link is exactly its strength: nothing in the second directory tells you the data is shared, so someone editing "their copy" in place edits everyone's.

  • Why does Linux refuse to let you create a hard link to a directory?
    Because arbitrary directory hard links would let the tree contain cycles, and a cyclic tree breaks recursive traversal, `rm -r`, and filesystem checking — there is no reliable way to know when you have visited everything. The kernel makes the only directory hard links itself: `.` and `..`. That is why a brand-new empty directory has a link count of 2, and gains one more for every subdirectory created inside it.
  • A file's link count is 3. What does that tell you, and how would you find the other names?
    Three directory entries somewhere on that same filesystem point at the inode, and the data survives until all three are unlinked. To find them, search the filesystem by inode: `find /mountpoint -xdev -inum 131074`, or `find /mountpoint -xdev -samefile /path/to/known/name`. `-xdev` matters because inode numbers are only unique within one filesystem, so crossing a mount point would produce false matches.
  • If you run chmod on a symlink, what actually changes?
    The target. `chmod` follows the link, so the mode bits of the file it points at change; the symlink's own permissions always display as `lrwxrwxrwx` on Linux and play no part in access decisions. Access to the target is governed by the target's mode bits plus execute (search) permission on every directory in the resolved path. With hard links there is no ambiguity — one inode, one set of permissions, visible through every name.

saying these in an interview costs you the question

  • Calls a hard link a shortcut or pointer to the original file
  • Says deleting the original filename breaks its hard links
  • Believes a hard link can span two mounted filesystems
  • Thinks a symlink stores an inode number rather than a path
  • Assumes a symlink's own permission bits protect the target

context