skip to content

Why does a typical Linux distribution boot through an initramfs instead of letting the kernel mount the root filesystem directly, and what does that early user space do before it hands over?

level: middleimportance: should knowfreq 56%

answer

  1. a chicken-and-egg driver problem
  2. generic kernel, per-host hardware
  3. temporary root that lives in RAM
  4. assemble root, then switch_root onto it
  5. host-only images are not portable

basics

~20 s

An initramfs is a compressed cpio archive the bootloader loads into RAM as a temporary root filesystem. It lets a single generic kernel load the drivers and assemble the devices needed to reach the real root, mount it, and then switch onto it.

solid answer

~50 s

A distribution ships one kernel for every machine, so the drivers for your disk controller and root filesystem are usually modules on the root filesystem — a chicken-and-egg problem. The initramfs breaks it. The bootloader loads it into memory next to the kernel; the kernel unpacks that cpio archive into a RAM-backed root and runs `/init` from it. That early user space loads the storage and filesystem drivers, waits for the device named by `root=` to appear, assembles anything that needs assembling — software RAID, an encrypted volume, a network-attached disk — mounts the real root read-only at `/sysroot`, and then `switch_root`s: it moves `/proc`, `/sys` and `/dev` over, frees the RAM filesystem and execs the real init. It is built per host by `dracut` on RHEL-family and SUSE systems, `update-initramfs` on Debian/Ubuntu, and `mkinitcpio` on Arch.

code

bash · 9 lines
bash
# what is inside the image that will be used at the next boot
lsinitrd /boot/initramfs-"$(uname -r)".img 2>/dev/null ||
  lsinitramfs /boot/initrd.img-"$(uname -r)"

# regenerate it after a config change (RHEL-family, then Debian/Ubuntu)
sudo dracut -f 2>/dev/null || sudo update-initramfs -u -k all

# check there is room in /boot before any kernel update
df -h /boot

go deeper

for a junior

Know that the initramfs is a small temporary root loaded into memory, and that it exists so the kernel can find and mount the real root filesystem.

for a middle

Explain the driver chicken-and-egg problem, name what early user space does before the pivot, and say that switch_root execs the real init rather than starting a new process.

for a senior

Connect a failed boot to a bad initramfs: a full /boot truncating the rebuild, a host-only image moved to different hardware, or a stale root UUID — and know how to regenerate it from a rescue shell.

for a principal

Decide the fleet policy: generic versus host-only images for portability, /boot sizing so regeneration cannot silently truncate, and whether kernel updates are validated by an actual reboot before the change is considered done.

## The problem it solves A distribution kernel has to boot laptops, NVMe servers, virtual machines with paravirtual disks, and hardware RAID controllers, all from one image. Compiling every driver in would make a huge kernel, so most drivers are modules — and modules live under `/lib/modules/` on the root filesystem, which the kernel cannot mount until it has the driver. That circularity is what the initramfs exists to break. There is a second reason, just as important on real servers: the root device may not exist as a device yet. If root sits on software RAID, on an encrypted volume that needs a passphrase, or on a network-attached disk, something in user space has to assemble or unlock it first. That "something" needs a filesystem to live on, and the only one available before root is mounted is one that lives in memory. ## What it physically is An **initramfs** is a cpio archive, compressed with gzip or zstd, sitting in `/boot` next to the kernel — `/boot/initramfs-<version>.img` on RHEL-family systems, `/boot/initrd.img-<version>` on Debian/Ubuntu. The bootloader loads it into memory as an opaque blob; the kernel unpacks it into a RAM-backed root filesystem and executes `/init` from it. (The older `initrd` mechanism was a *block device* image that the kernel mounted and later `pivot_root`ed away from; the modern initramfs is unpacked, not mounted, which is why it can simply be freed.) You can look inside one: ```bash lsinitrd /boot/initramfs-$(uname -r).img # dracut-based systems lsinitramfs /boot/initrd.img-$(uname -r) # Debian/Ubuntu ``` ## What its /init actually does Roughly, in order: start device management so hardware events are processed; load the modules for the storage controller and the root filesystem; wait for the device named by `root=` to appear, which is why `root=UUID=…` is used rather than a kernel device name that can change between boots; assemble or unlock the root device if needed; optionally check it; mount it **read-only** at `/sysroot`; and finally `switch_root` — move the `/proc`, `/sys` and `/dev` mounts into the new root, release the RAM filesystem, chroot, and `exec` the real init. Because that last step is an exec, the process keeps PID 1: there is exactly one PID 1 for the whole life of the machine. ## Host-only versus generic images This is where interviews get interesting. `dracut` defaults to a **host-only** image on Fedora and RHEL: it includes only the modules this machine currently needs, which keeps the image small and boot fast. The consequence is that the image is not portable — move that disk into a machine with a different disk controller, or change the storage backing of a VM, and the initramfs no longer contains the driver for the new root device. You get a kernel that boots fine and then panics with `VFS: Unable to mount root fs`. Rebuilding with a generic image (`dracut --no-hostonly`) is the fix, and it is the reason golden images for varied hardware are built non-host-only. ## Why it breaks The initramfs is **regenerated** whenever a kernel is installed or a relevant config changes, so it is a file that can be silently written wrong: - `/boot` is full, so the new image is truncated — the classic "it booted fine until the next reboot, three weeks after the update" failure. - The rebuild ran with a config that omits the module for this host's disk. - `root=` points to a UUID that no longer exists after a disk was replaced or cloned. All three present the same way: the kernel starts, prints messages, and then cannot find or mount root. The diagnostic move is to interrupt the initramfs before the pivot — dracut offers `rd.break` on the kernel command line, Debian's initramfs-tools offers `break=` — and look at what devices are actually present from that shell. ## The rule of thumb If the failure happens *before* any user-space service messages appear but *after* the kernel is clearly running, suspect the initramfs and the `root=` parameter. Regenerating it (`dracut -f` or `update-initramfs -u -k all`) from a rescue environment fixes a large share of "the server won't boot after the update" incidents.

  • What is the practical difference between the old initrd mechanism and today's initramfs?
    An initrd is a block-device image the kernel mounts as a real filesystem, using `pivot_root` to move away from it and needing an explicit release afterwards. An initramfs is a cpio archive unpacked into a RAM-backed root; there is no block device and no filesystem driver involved, and `switch_root` simply frees the memory. Distributions kept the old `initrd.img` filename long after switching to initramfs content.
  • Why do distributions use root=UUID=… on the kernel command line instead of naming the device directly?
    Kernel device names depend on probe order, which can change when disks are added, a controller is swapped, or a VM's storage layout changes — yesterday's first disk may not be today's. A UUID is stored in the filesystem itself, so early user space can wait for and identify the correct device regardless of enumeration order. It is the difference between a machine that survives a hardware change and one that panics.
  • When is it safe to run a system with no initramfs at all?
    When the kernel has the disk controller and root filesystem drivers built in, and the root device needs no assembly or unlocking — typical of custom-built and embedded kernels tuned for one known machine. You gain a slightly faster, simpler boot and lose portability: any hardware change that needs a module now leaves the kernel unable to mount root.

saying these in an interview costs you the question

  • Thinks the initramfs is the root filesystem itself
  • Says the kernel loads it from disk after booting
  • Believes it stays mounted while the system runs
  • Assumes one initramfs works on any hardware
  • Cannot explain why root= uses a UUID

context