skip to content

Filesystem Hierarchy

Where files live and how they are really stored — the FHS layout, inodes and links, mount points and fstab, and the trade-offs between ext4, xfs, and tmpfs. Interviewers use this to check you can find your way around and repair an unfamiliar system.

part ofLinux & distributionsoverview, primer and where to startread it →
on this pageshow

questions

22

You are handed a shell on an unfamiliar Linux server. Where do you expect a service's configuration, its persistent state and its log files to live, and what is the rule that puts those three things in different top-level directories?

level: juniorimportance: must knowfreq 70%

answer

  1. three directories, three lifetimes
  2. who writes it, and how long it lives
  3. static package code versus local edits
  4. back up two of them, reinstall the third
  5. read-only root depends on this split

basics

~20 s

Configuration lives under /etc, persistent state under /var/lib and logs under /var/log. The split is by mutability and ownership: /usr holds static package-owned code, /etc holds locally editable configuration, /var holds data the machine writes while it runs.

solid answer

~40 s

The Filesystem Hierarchy Standard splits the tree by **who writes a file and how long it lives**, not by what kind of file it is. `/usr` is static and owned by the package manager — binaries, libraries and shipped data, nothing the running system modifies, which is why it can be mounted read-only. `/etc` is host-local configuration: text, editable by the administrator, small, and the thing that makes this machine different from an identical one. `/var` is everything the machine itself writes: `/var/log` for logs, `/var/lib/<service>` for persistent state such as a database's data files, `/var/spool` for queues. So on an unknown host I look in `/etc/<name>` first, then `/var/lib/<name>` and `/var/log/<name>`. The payoff is practical: back up `/etc` and `/var`, reinstall `/usr`, and the machine comes back.

go deeper

for a junior

Be able to name the three places without hesitating: /etc for configuration, /var/log for logs, /var/lib for a service's data. Say plainly that /usr is program files installed by packages, not user home directories.

for a middle

Explain the organising rule rather than reciting directories: static and package-owned versus local and editable versus written by the running system. Be ready to place an unfamiliar file correctly by applying that rule out loud.

for a senior

Show the operational payoff: which directories a backup must cover, why /var on its own filesystem stops a runaway log from wedging a host, and how the split enables a read-only /usr. Distinguish /var/lib from /var/cache when planning recovery.

for a principal

Own the standard across a fleet: where local convention diverges from the FHS, what that costs when images become immutable, and how to make configuration ship as vendor defaults in /usr with overrides in /etc so upgrades never prompt.

## The rule, not the list The Filesystem Hierarchy Standard (FHS 3.0, maintained by the Linux Foundation; also summarised in the `hier(7)` man page) is often taught as a list of directories to memorise. That is the wrong way in. The tree is organised along two axes, and once you know them you can place a file you have never seen before. The first axis is **shareable vs unshareable**: can this data be identical across many machines, or is it specific to this host? The second is **static vs variable**: does it change only when an administrator installs or upgrades software, or does the running system write it by itself? | | static | variable | |---|---|---| | shareable | `/usr` | `/var/mail`, `/srv` | | unshareable | `/etc`, `/boot` | `/var/log`, `/var/lib`, `/run` | ## /usr — static, package-owned code `/usr` holds the operating system's programs and data: `/usr/bin` for executables, `/usr/lib` for shared libraries and private helper files, `/usr/share` for architecture-independent data (documentation, icons, locale data, man pages), `/usr/include` for headers. Every file here is owned by a package. Nothing the system does at runtime modifies it. That property is what makes read-only-root images, atomic-update distributions and shared or verified `/usr` trees possible. The name is a historical accident — it did once hold user home directories — but it is now read as "Unix system resources". Home directories live in `/home`, and root's home is `/root`. ## /etc — the machine's identity `/etc` is host-specific configuration. The FHS is explicit that it must contain no binaries: it is configuration, and it should be small enough that a human can read it. `/etc/passwd`, `/etc/fstab`, `/etc/hosts`, `/etc/ssh/sshd_config` are what distinguish this host from a freshly installed one. Many projects now ship their *defaults* under `/usr/lib/<project>` or `/usr/share/<project>` and let `/etc` hold only the administrator's overrides, often as drop-in files in a `*.d` directory. That is a refinement of the same rule: vendor content is static and package-owned, so it belongs in `/usr`; your edit is local, so it belongs in `/etc`. It also means an upgrade never has to ask "you modified this config file, what should I do?" for files you never touched. ## /var — what the machine writes `/var` is variable data whose *existence* the administrator does not manage file by file: - `/var/log` — log files. - `/var/lib/<service>` — persistent application state: a PostgreSQL cluster's data directory, a package manager's database, a service's cached credentials. This is the directory people most often misplace. - `/var/cache` — regenerable data; deleting it must cost only performance, never correctness. - `/var/spool` — queues (mail, print, cron jobs). - `/var/tmp` — temporary files that must survive a reboot. The key discipline is separating `/var/lib` (lose it and you lose data) from `/var/cache` (lose it and nothing breaks). That distinction drives backups, disk-pressure cleanup and container volume design. ## The remaining ephemera `/run` holds runtime state valid only since boot — PID files, Unix sockets, lock files — and is cleared at every boot. `/tmp` is scratch space with no guarantee of surviving a reboot. `/srv` is data served by this system to the outside world (web roots, exported trees), and `/opt` is self-contained third-party software. `/boot` holds the kernel and initramfs; `/dev` holds device nodes managed by the kernel. ## Why the split earns its keep Three everyday consequences follow directly: 1. **Backups.** `/etc` and `/var/lib` are the machine. `/usr` is reinstallable from packages, `/var/cache` and `/tmp` are disposable. A backup policy falls out of the hierarchy for free. 2. **Read-only and immutable systems.** Because `/usr` is static, an image-based or container-based system can mount it read-only and keep only `/etc`, `/var` and `/run` writable. Container images lean on exactly this: the layers are `/usr`-shaped, the volumes are `/var`-shaped. 3. **Partitioning and disk pressure.** Putting `/var` on its own filesystem means a runaway log cannot fill the root filesystem and wedge the whole host. ```sh # the three places to look for an unfamiliar service called "foo" ls /etc/foo /var/lib/foo /var/log/foo ``` When you are unsure where something you are writing belongs, ask the two questions in order: does the package manager own it (`/usr`), does an administrator edit it (`/etc`), or does the program write it itself (`/var`)?

  • Why does the FHS say /etc must contain no binaries?
    `/etc` is meant to be small, host-specific, human-readable configuration that an administrator can read end to end and copy between machines. Executables are package-owned, architecture-dependent and static, so they belong in `/usr/bin` or `/usr/lib`. Keeping binaries out also means `/etc` can be captured in configuration management or version control without dragging along compiled artefacts.
  • If a system mounts /usr read-only, what still has to be writable for it to work?
    `/etc` for configuration changes, `/var` for logs and service state, `/run` for runtime sockets and PID files, and `/tmp` for scratch. Image-based systems that want `/etc` read-only too usually provide it as an overlay or a tmpfs seeded from `/usr`, so local edits live somewhere writable while the shipped defaults stay in the read-only tree.
  • What is the difference between /var/lib and /var/cache, and why does it matter operationally?
    `/var/lib` is authoritative state — deleting it loses data. `/var/cache` is regenerable: deleting it costs only time, never correctness. The distinction drives backups (back up `lib`, skip `cache`) and disk-pressure response (a cleanup script may safely empty `/var/cache`, never `/var/lib`). Services that blur the two make both jobs unsafe.

saying these in an interview costs you the question

  • Says /usr is where user home directories live
  • Treats /etc as the place for logs or runtime state
  • Thinks /var is only for log files
  • Puts a service's database files under /etc because they are configuration
  • Backs up /usr and skips /var/lib

context

open as a page

An /etc/fstab entry names a disk as /dev/sdb1. Why is that fragile on a Linux server, and how do UUID=, LABEL= and PARTUUID= differ as replacements?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Kernel names like /dev/sdb1 depend on device discovery order, so a new disk or a slow controller can renumber them and mount the wrong filesystem. Prefer UUID= from the filesystem superblock, or PARTUUID= from the partition table.

open as a page

You build a tool from source on a Linux server and run `make install`. Under the FHS, why does it belong in /usr/local rather than /usr, and when is /opt the right home instead?

level: middleimportance: must knowfreq 55%

basics

~20 s

/usr is owned by the distribution's package manager, so an upgrade can overwrite or conflict with anything you drop there. /usr/local is reserved for software the administrator installs and packages never touch. /opt holds self-contained third-party bundles in their own subtree.

open as a page

On an ext4 filesystem, what does the journal actually protect after an unclean shutdown, and how do the data=ordered, data=writeback and data=journal mount options differ?

level: middleimportance: must knowfreq 62%

basics

~20 s

ext4's journal protects filesystem metadata consistency, not file contents. The default data=ordered flushes data blocks before committing the metadata that points at them; data=writeback drops that ordering and can expose stale bytes; data=journal journals data too, more slowly.

open as a page

`umount /data` returns "target is busy". What kinds of references make a Linux filesystem busy, and what does `umount -l` actually do that a plain unmount does not?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A filesystem is busy while anything still references it: open file descriptors, a process whose working directory or root is inside it, memory-mapped files, an active swap file, or a submount. umount -l (MNT_DETACH) only detaches the subtree from the path namespace — the filesystem stays alive until the last reference goes away.

open as a page

You are provisioning two volumes on a Linux server: one for a busy relational database's data directory that may have to grow later, and one for a CI build cache holding millions of small files that is wiped weekly. How would you choose between ext4, XFS and Btrfs for each, and which property of each filesystem drives the choice?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Pick XFS or ext4 for the database — both journal metadata and overwrite in place, and XFS grows online but can never shrink. Btrfs suits the disposable build cache, where cheap snapshots, compression and checksums outweigh its copy-on-write fragmentation.

open as a page

A Linux host has a tmpfs filesystem mounted at /dev/shm whose size is reported as half the machine's RAM. What kind of storage is tmpfs, what does that size figure actually mean, and what happens to files written into it?

level: juniorimportance: should knowfreq 52%

basics

~20 s

tmpfs is a memory-backed filesystem: its files live in the kernel's page cache, may be pushed out to swap under pressure, and disappear on unmount or reboot. The reported size is a ceiling, not a reservation.

open as a page

A normal user gets "command not found" for an administrative tool, but `ls /usr/sbin` shows the binary is right there and root runs it fine. On Linux, what does the /usr/bin versus /usr/sbin split mean, and why does the shell not find the program?

level: middleimportance: should knowfreq 45%

basics

~20 s

The sbin directories hold system-administration programs, a convention about intended audience rather than a permission boundary. A normal user's default PATH usually omits them, so the shell never looks there — the file is executable, just not on the search path. Invoking it by absolute path works.

open as a page

On Linux, what does `mount --bind /srv/data /var/www/html` actually create, how does it differ from putting a symlink at that path, and what does `--rbind` add?

level: middleimportance: should knowfreq 50%

basics

~20 s

A bind mount makes an existing directory (or file) appear at a second path as a real mount entry, with both paths referring to the same underlying data. Unlike a symlink it is a genuine mount, not a link the caller can detect or refuse; --rbind also replicates any submounts.

open as a page

In an /etc/fstab entry, what do the mount options nosuid, nodev and noexec each enforce on Linux, and why is noexec on /tmp weaker protection than people assume?

level: middleimportance: should knowfreq 45%

basics

~20 s

nosuid makes the kernel ignore setuid and setgid bits (and file capabilities) on that mount; nodev makes device special files there non-functional; noexec makes execve fail for files under it. noexec is weak because an interpreter can still read and run a script as data.

open as a page

After an unclean reboot an ext4 filesystem mounts in seconds, and a colleague insists that XFS 'has no fsck'. What actually happens at mount time on each of these filesystems, and how do you repair an XFS filesystem that refuses to mount?

level: middleimportance: should knowfreq 44%

basics

~20 s

Both replay their journal at mount, which is why recovery is fast — that is not a full check. XFS really does ship a do-nothing fsck.xfs; genuine repair uses xfs_repair on an unmounted filesystem, with -n to inspect and -L as a data-losing last resort.

open as a page

On a Linux system, what is /run for, and why does a daemon whose installation script does `mkdir /run/myservice` stop working after the first reboot?

level: seniorimportance: should knowfreq 40%

basics

~20 s

/run holds runtime state valid only since the current boot — PID files, Unix sockets, lock files. It is a memory-backed filesystem mounted early and empty at every boot, so a directory created once at install time is gone after a restart and must be recreated during startup instead.

open as a page

A long-running batch job writes intermediate files under /tmp and occasionally fails days later with "No such file or directory" for a file it created. On a Linux host, how do /tmp and /var/tmp differ, and where should that scratch data live?

level: seniorimportance: should knowfreq 50%

basics

~20 s

/tmp is volatile: it may be cleared at boot, is aged out by a periodic cleanup, and on many distributions is a memory-backed filesystem with its own size limit. /var/tmp is temporary storage that must survive reboots. Multi-day scratch data belongs in /var/tmp or the service's own directory.

open as a page

You add an fstab entry for a new data disk, reboot, and the machine drops to an emergency shell instead of coming up. Why can one fstab line stop the whole boot, and which mount options let a non-critical filesystem fail without taking the system down?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Every fstab entry becomes a boot-time dependency: on a systemd machine the fstab generator turns each line into a mount unit that local-fs.target requires, so a missing device or failed mount fails the target and drops you to emergency mode. Mark optional filesystems nofail.

open as a page

On a current Linux distribution, `ls -ld /bin` shows that /bin is a symbolic link to usr/bin. What is the "/usr merge", why was it done, and what does it change in practice?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

The /usr merge turns /bin, /sbin, /lib and /lib64 into symbolic links into /usr, so every program and library has exactly one real location. It removes a split that only existed so a tiny root filesystem could mount a separate /usr, a job the initramfs now does.

open as a page

Linux mounts carry a propagation type — shared, private or slave. What do those mean, and why does a filesystem mounted inside a separate mount namespace sometimes appear on the host and sometimes not?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Propagation controls whether mount and unmount events on one mount are replicated to its peers. Shared mounts propagate both ways, private mounts propagate nothing, and slave mounts receive events from their master but send none back — which is why a mount made in an isolated tree may or may not become visible elsewhere.

open as a page

On a copy-on-write filesystem such as Btrfs or ZFS, what is a snapshot physically, and why do free space and file layout behave in ways that surprise people once snapshots exist?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

A snapshot is a second reference to the same on-disk blocks at a moment in time, created instantly with no data copied. Space is consumed later, as the live copy diverges, which is why deleting files frees nothing while a snapshot still references them.

open as a page