skip to content

A Linux server no longer completes its boot after a kernel package update, and you have console access only. How do you get a root shell on it and work out what broke?

level: seniorimportance: must knowfreq 55%

answer

  1. the last working stage is your tool
  2. the menu still works, use it
  3. previous kernel entry first
  4. edit the command line for one boot
  5. root comes up read-only

basics

~20 s

Interrupt the bootloader menu, boot the previous kernel entry to get a working system, and if that fails edit the entry's kernel command line for a one-off recovery boot that stops in the initramfs or execs a shell instead of init.

solid answer

~50 s

Start at the GRUB2 menu, because it is the last stage that still works. Distributions keep several kernels installed, so the first move is booting the previous entry — if that succeeds you have a working machine and a fault isolated to the new kernel or its initramfs. If it does not, press `e` to edit the entry's kernel command line for this boot only. `rd.break` on a dracut system stops in the initramfs before the root pivot, so you can see whether the root device exists at all and remount `/sysroot` read-write to fix things. `init=/bin/bash` boots further and execs a shell instead of init, which isolates whether the fault is the kernel or user space; root comes up read-only, so remount it before editing anything, and never just exit that shell because the kernel panics when PID 1 dies. Then check `/proc/cmdline`, whether `/boot` is full, and whether the initramfs for the new kernel is a plausible size.

code

bash · 11 lines
bash
# after reaching a recovery shell via init=/bin/bash: root is read-only
mount -o remount,rw /

# the three checks that explain most post-update boot failures
cat /proc/cmdline
df -h /boot
ls -l /boot/initramfs-* /boot/initrd.img-* 2>/dev/null

# finish safely: this shell IS pid 1, exiting it panics the kernel
sync
reboot -f

go deeper

for a junior

Know that the bootloader menu lets you pick an older kernel, and that a boot failure right after an update is usually the new kernel or its initramfs rather than the hardware.

for a middle

Explain what editing the kernel command line at the menu does, that the change applies to this boot only, and why root is mounted read-only when you reach a recovery shell.

for a senior

Drive the escalation in order — previous kernel, one-off command-line edit, rescue media — and read the evidence: /proc/cmdline, /boot free space, whether the root device is visible from the initramfs shell.

for a principal

Address the systemic gap: kernel updates validated by an actual reboot, retention of known-good entries, /boot sized for regeneration, and out-of-band console access so recovery never depends on the network.

## Work backwards along the chain Boot recovery is the boot chain read in reverse. Whatever stage you can still reach is the tool you use to reach the one you cannot. After a kernel update the bootloader itself is almost always fine, so the GRUB2 menu is your entry point. ## Move one: boot the previous kernel Distributions deliberately keep the previous kernel or two installed. Interrupt the menu countdown and pick the older entry — often under an "Advanced options" submenu. This is the highest-value action available: it takes seconds, it either restores service or eliminates a whole class of cause, and it leaves the broken kernel installed for analysis. If the old kernel boots, you have a running system with logs, package history and a network, and the incident becomes an investigation rather than an outage. ## Move two: edit the command line for one boot Pressing `e` at the menu opens the entry for editing; the change applies to this boot only and is not written to disk, which is exactly the property you want in a recovery. Two edits matter. **`rd.break`** (dracut-based systems — RHEL family, Fedora, SUSE; Debian's initramfs-tools uses `break=` instead) stops early user space *before* it pivots to the real root. You land in a shell inside the initramfs. From there the question is simply whether the root device exists: ```bash # in the initramfs shell after rd.break ls /dev/disk/by-uuid/ # is the UUID from root= actually present? mount -o remount,rw /sysroot # the real root, mounted read-only at /sysroot chroot /sysroot ``` If the UUID is absent, the initramfs lacks the driver for this host's storage, or `root=` is wrong — the two dominant causes of a post-update boot failure. **`init=/bin/bash`** goes further: the kernel mounts root normally and then execs a shell instead of the system's init. Reaching this shell proves the kernel, the initramfs and the root filesystem are all fine and moves suspicion into user space. Two rules apply. Root is mounted **read-only**, so `mount -o remount,rw /` before changing anything. And that shell **is** PID 1 — exiting it kills init and panics the kernel, so finish with `sync` and a forced reboot rather than `exit`. There are gentler variants for a system that boots but fails partway into starting services: `single` and, on systemd systems, passing a rescue unit on the command line so the machine comes up with a root shell and minimal services. Which you reach for depends on how far the boot gets before it stops. ## Move three: the rescue environment If no menu entry works at all — a corrupt bootloader config, a wiped `/boot` — boot installation or netboot media in rescue mode, mount the root filesystem and the EFI System Partition under it, `chroot` in, and repair from there: regenerate the initramfs (`dracut -f --regenerate-all` or `update-initramfs -u -k all`), regenerate the bootloader config (`grub2-mkconfig -o …` or `update-grub`), and reinstall the bootloader if needed. ## What to actually look for Once you have any shell, a small checklist covers most post-update failures: - **Is `/boot` full?** A full `/boot` truncates the initramfs written during the update. The kernel starts and then cannot mount root. Compare the new image's size against the previous kernel's. - **Does `/proc/cmdline` match what you expect?** It is the authoritative record of what the bootloader passed, and it will show a stale `root=` UUID after a disk replacement or clone. - **Does the initramfs contain the storage driver?** Host-only images built on different hardware, or built while a module was unavailable, produce exactly this symptom. - **Did the update finish?** An interrupted package transaction can leave a kernel installed with no matching initramfs at all. ## The part people get wrong The instinct is to reinstall or to reach for the rescue media immediately. The disciplined sequence is cheaper: previous kernel first, then a one-off command-line edit, then rescue media — each step only if the one before failed. And afterwards, the finding that matters to the organisation is usually not the specific driver: it is that a kernel update was considered complete without a reboot to validate it, so the failure surfaced weeks later on an unrelated restart.

  • Why must you not simply type exit in a shell you reached with init=/bin/bash?
    Because that shell is PID 1. The kernel panics when init exits — it has no other process to own and reap the process tree — so you get an immediate halt, potentially with unsynced writes if you had remounted root read-write. Finish the session with `sync` and then a forced reboot or the kernel's own reset path instead.
  • Why is /boot filling up such a common cause of a machine that boots fine until its next reboot?
    Kernel updates write a new kernel and a freshly generated initramfs into /boot, and old kernels accumulate. If the partition is full, the initramfs is written truncated and the update may still report success. Nothing fails until the machine actually reboots onto that entry — which can be weeks later, making the correlation with the update invisible unless you look.
  • How would you make one-off recovery command-line changes permanent, and why is editing the generated GRUB config directly wrong?
    Set the parameter in `/etc/default/grub` and regenerate with `grub2-mkconfig -o …` or `update-grub`, or on RHEL 9 use `grubby` to update the Boot Loader Specification entries. The generated `grub.cfg` is rewritten on every kernel update, so an edit made directly there survives until the next update and then silently vanishes — the worst possible failure mode for a fix you are relying on.

saying these in an interview costs you the question

  • Reaches for reinstall before trying the previous kernel
  • Edits changes into the generated grub.cfg permanently
  • Forgets root is mounted read-only in recovery
  • Exits the init=/bin/bash shell and panics the box
  • Never checks whether /boot ran out of space

context