skip to content

System & Service Management

Everything you do to a Linux box once it is running: starting and supervising services, kernel-level storage and tracing, installing packages, watching processes, wiring up networking and the firewall, and getting in from somewhere else. Interviews lean on this layer because it separates people who have operated servers from people who have only deployed to them.

on this pageshow

explore

questions

164 · 6 sections

On a systemd-managed Linux host, a service unit has just failed. Which journalctl filters do you combine to show only that unit's messages, only errors, and only the current boot — and what does each flag do?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Combine journalctl -u nginx.service to select the unit, -b to scope to the current boot, and -p err to keep priority error and above. Add -f to follow live, -n 100 for the last lines, or --since for a time window.

open as a page

On a systemd host, what is the difference between `systemctl reload nginx` and `systemctl restart nginx`, and why does `systemctl reload` fail outright on some units?

level: juniorimportance: must knowfreq 68%
basics
~20 s

systemctl restart stops the service and starts a new process, so the PID changes and in-flight work is dropped. systemctl reload runs only the unit's ExecReload= command, leaving the same process running; a unit that defines no ExecReload= rejects reload with an error.

open as a page

On a systemd host you add a backup.timer unit with an OnCalendar= schedule, but the backup never runs. How does a systemd timer unit actually get its work done, and what has to be in place for the schedule to be active?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A .timer unit only schedules; the work lives in a separate .service unit that it activates — by default the same-named one. You must enable and start the timer itself (systemctl enable --now backup.timer), not the service.

open as a page

Walk through a minimal systemd service unit file: which three sections does it have, and which of Description=, After=, ExecStart=, User= and WantedBy= belongs in each? What does systemd do with a key placed in the wrong section?

level: juniorimportance: must knowfreq 64%
basics
~20 s

[Unit] holds Description= and relationship keys such as After=; [Service] holds how to run the process, including ExecStart= and User=; [Install] holds WantedBy= and is read only when the unit is enabled. A key in the wrong section is ignored with a warning.

open as a page

A systemd service unit declares Requires=postgresql.service but has no After= line. Why can it still start before PostgreSQL is up, and what does After= actually change?

level: middleimportance: must knowfreq 72%
basics
~20 s

Requirement and ordering are separate axes in systemd. Requires= only says the other unit must be pulled in and must not fail; it never says "wait for it". Without After=, both units are started in parallel, so either can win the race.

open as a page

In eBPF, what is a map, and how does a user-space program get at the data that a kernel-side eBPF program writes into one?

level: juniorimportance: must knowfreq 78%
basics
~20 s

An eBPF map is a kernel-resident key/value store shared by BPF programs and user space. BPF code uses helpers like bpf_map_lookup_elem(); user space copies values in and out via the bpf() syscall on the map's file descriptor.

open as a page

An eBPF program is loaded into the Linux kernel with a fixed program type such as BPF_PROG_TYPE_KPROBE or BPF_PROG_TYPE_XDP. What does that program type decide about the program, and why can't you attach one program anywhere you like?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An eBPF program's type fixes which kernel hooks it may attach to, the layout of the single context argument it receives, which BPF helper functions it may call, and how the kernel interprets its return value.

open as a page

In Linux eBPF, at what point does the kernel's verifier examine your program, and what happens to a program that fails verification?

level: juniorimportance: must knowfreq 70%
basics
~20 s

The eBPF verifier runs inside the kernel at load time, when the bpf() syscall submits the program — before it is attached to anything and before a single instruction executes. A rejected program never runs: the load call fails and the kernel returns a verifier log saying why.

open as a page

On a Linux server with two blank disks, explain LVM's three layers — physical volume, volume group and logical volume — and give the commands, in order, that turn those disks into a mounted filesystem.

level: juniorimportance: must knowfreq 78%
basics
~20 s

LVM stacks three layers: pvcreate initialises each disk as a physical volume, vgcreate pools those PVs into a volume group, and lvcreate carves logical volumes out of the pool. Then you mkfs and mount the LV like any block device.

open as a page

On a Linux host using LVM, you run `lvextend -L +50G /dev/vg0/data` and it reports success, yet `df -h` still shows the old size for the mounted filesystem. Why, and what do you run to make the filesystem use the new space?

level: juniorimportance: must knowfreq 76%
basics
~20 s

An LVM logical volume and the filesystem on it are separate layers. lvextend grew only the block device, so the filesystem still records its old size. Run resize2fs (ext4) or xfs_growfs (XFS) afterwards, or pass lvextend -r to do both.

open as a page

On a Debian or Ubuntu server, `apt install curl` fails with "Unable to locate package curl", and on a different, rarely-touched box an install fails with a 404 while downloading the .deb. What does `apt update` actually do, and why do both failures go away once you run it?

level: juniorimportance: must knowfreq 78%
basics
~20 s

apt update refreshes the package index from every configured repository and installs nothing. With no index apt cannot resolve the package name at all; with a stale index it asks the mirror for a .deb the archive has already replaced, so the download 404s.

open as a page

On a RHEL 8 or 9 host, `yum install httpd` still works even though the system's package manager is dnf. What is the `yum` command on such a host, and what happened to /etc/yum.conf and /etc/yum.repos.d?

level: juniorimportance: must knowfreq 62%
basics
~20 s

On RHEL 8 and 9 the yum command is a symlink to dnf-3, so running yum runs dnf. /etc/yum.conf is a symlink to /etc/dnf/dnf.conf, and /etc/yum.repos.d is still the directory dnf reads repository files from.

open as a page

On a Debian or Ubuntu host, what is the difference between `apt upgrade`, `apt full-upgrade` and `apt-get dist-upgrade`, and which of them is allowed to remove an installed package?

level: middleimportance: must knowfreq 62%
basics
~20 s

apt upgrade never removes an installed package; anything whose upgrade would require a removal is kept back. apt full-upgrade and apt-get dist-upgrade are the same command and may remove packages to resolve conflicts. apt-get upgrade is stricter still: it also refuses to add new packages.

open as a page

On a RHEL or Fedora host, a `dnf upgrade` run an hour ago has left a service broken. How do you find out exactly what that transaction changed, how do you reverse it, and what makes the reversal fail?

level: middleimportance: must knowfreq 64%
basics
~20 s

dnf records every transaction locally. dnf history lists them with IDs, dnf history info <id> shows every package the transaction added, removed or upgraded, and dnf history undo <id> reverses it — provided the older RPMs are still obtainable.

open as a page

A snap-installed application refuses to open a file under /mnt/data even though the file is world-readable and you can read it with `cat` as the same user. Why does snap's strict confinement block it, and how do you grant the access?

level: middleimportance: must knowfreq 50%
basics
~20 s

Strict snap confinement sandboxes the process with AppArmor and seccomp, so access depends on the interfaces the snap has connected rather than on Unix permissions. Paths under /mnt need the removable-media interface, which does not auto-connect; attach it with snap connect.

open as a page

A Linux service vanished overnight with nothing useful in its own log. How do you confirm that the kernel's out-of-memory killer took it, and what does the kernel's log line tell you about the process it killed?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Check the kernel log with dmesg -T or journalctl -k, searching for "Out of memory" or "oom-kill". The kill line names the PID and command and reports total-vm, anon-rss, file-rss, shmem-rss and oom_score_adj — anon-rss being what the process actually held in RAM.

open as a page

You have a shell on an unfamiliar Linux server and need a full process listing. What is the difference between `ps aux` and `ps -ef`, and when would you reach for `ps -eo` instead of either?

level: juniorimportance: must knowfreq 74%
basics
~20 s

ps aux and ps -ef are BSD-style and UNIX-style invocations of the same tool, differing in default columns: aux prints %CPU, %MEM, VSZ and RSS; -ef prints PPID and start time. ps -eo lets you choose columns and sort order.

open as a page

A Linux application server feels slow and `vmstat 1` shows the `wa` column sitting around 40%. What does iowait actually measure, and why is a high value on its own not proof that the disk is the problem?

level: middleimportance: must knowfreq 62%
basics
~20 s

Iowait is idle CPU time that happened while at least one I/O request was outstanding. It is a subset of idle, so a high value means the CPUs had nothing else to run — it locates spare capacity, not a slow disk. Confirm with iostat -x latency before blaming storage.

open as a page

Overnight, a database host's disk latency doubled. In `iostat -x` output, which fields tell you whether the device itself got slower or the queue in front of it got deeper, and why can `%util` at 100% be meaningless on an SSD or a RAID array?

level: seniorimportance: must knowfreq 58%
basics
~20 s

Read the await columns for per-request latency, aqu-sz for queue depth, and r/s/w/s for offered load. Latency up with load and queue flat means the device slowed; latency up with a deeper queue means you sent more work. %util only counts time with at least one request in flight, so on parallel devices 100% is not saturation.

open as a page

You run `vmstat 1` on a Linux server and the first line shows an almost idle machine, while every line after it shows 90% system time. Why does the first line disagree with the rest, and which lines should you actually read?

level: juniorimportance: should knowfreq 45%
basics
~10 s

The first report from vmstat (and from iostat, mpstat and pidstat) is an average since system boot, not a sample of the last second. Ignore it and read the interval lines that follow.

open as a page

A monitoring check reports a Linux server as unreachable because `ping` gets no reply, yet the HTTPS service on that same host is serving traffic normally. What does a failed ping actually prove, and how would you test reachability instead?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A failed ping proves only that ICMP echo replies are not coming back, which is usually firewall or security-group policy rather than a dead host. Test the port the service actually listens on, using nc -zv or curl -v.

open as a page

On a Linux host you need to confirm whether anything is listening on TCP port 8080 and, if so, which process owns it. How do you do that with `ss`, what does each letter in `ss -tulpn` select, and why does the process column sometimes come back empty?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Run ss -tulpn as root: t selects TCP, u UDP, l listening sockets only, p the owning process, n numeric ports. The process column is empty for sockets owned by other users unless you are root.

open as a page

In nftables, what does the `inet` family mean when you write `table inet filter`, and what does it change compared with maintaining separate iptables and ip6tables rulesets?

level: juniorimportance: must knowfreq 60%
basics
~20 s

The inet family is nftables' dual-stack address family: one table whose base chains are evaluated for both IPv4 and IPv6 packets. It replaces keeping two parallel rulesets, so each rule is written, reviewed and audited once.

open as a page

On a Linux server you run `iptables -A INPUT -p tcp --dport 443 -j ACCEPT`, but the INPUT chain already ends with a catch-all `-j DROP` rule and port 443 is still unreachable. Explain why the new rule has no effect, and what you would run instead.

level: middleimportance: must knowfreq 72%
basics
~20 s

iptables walks a chain from top to bottom and stops at the first matching rule with a terminating target, so an ACCEPT appended below an existing catch-all DROP is never reached. Insert it above that rule with -I and a position instead of appending with -A.

open as a page

A database on a Linux server is running and `ss -tlnp` shows it in LISTEN state on port 5432, yet remote clients get no connection while a client on the same host connects fine. What in the ss Local Address:Port column explains this, and how do you fix it?

level: middleimportance: must knowfreq 68%
basics
~10 s

The Local Address is almost certainly 127.0.0.1, so the socket is bound to loopback only and is unreachable from any other host. Fix it in the application's own bind configuration, not in the firewall.

open as a page

Explain what a trailing slash on an rsync SOURCE path changes, using `rsync -a /var/www /backup/` versus `rsync -a /var/www/ /backup/`, and why getting it wrong is dangerous once `--delete` is added.

level: juniorimportance: must knowfreq 78%
basics
~20 s

A trailing slash on an rsync source means "copy this directory's contents"; without it rsync copies the directory itself. So /var/www/ fills /backup, while /var/www creates /backup/www. A trailing slash on the destination changes nothing.

open as a page

You generated an SSH key pair on your laptop with `ssh-keygen`. What does `ssh-copy-id deploy@server` then do, which of the two files it produced ends up on the server, and where exactly does it go?

level: juniorimportance: must knowfreq 68%
basics
~20 s

ssh-copy-id logs in with the credentials you already have, then appends your public key file (for example id_ed25519.pub) to the remote account's ~/.ssh/authorized_keys, creating the directory and file with safe permissions. The private key never leaves your machine.

open as a page

A 2 GB file on a remote host changes by a few kilobytes a day, yet `rsync -a` over a slow link finishes in seconds. Describe the delta-transfer algorithm that makes that possible, and name a common situation where rsync does not use it at all.

level: middleimportance: must knowfreq 62%
basics
~20 s

The receiver splits its existing copy into blocks and sends the sender a weak rolling checksum plus a strong checksum for each. The sender rolls a window byte by byte through its version, and transmits only unmatched literal data plus references to blocks the receiver already has. Local-to-local copies skip this and send whole files.

open as a page

Key-based SSH login to a Linux server fails with `Permission denied (publickey)` even though the right public key is present in `/home/deploy/.ssh/authorized_keys`, and the server's auth log shows `Authentication refused: bad ownership or modes for directory /home/deploy/.ssh`. What is sshd checking, and which ownership and permissions does it require?

level: middleimportance: must knowfreq 62%
basics
~20 s

OpenSSH's sshd runs with StrictModes yes by default: it ignores a key file that anyone but the owner could modify. The account's home directory must not be group- or world-writable, ~/.ssh should be 700 and authorized_keys 600, all owned by that user.

open as a page

In rsync, what does the `-a` (archive) option expand to, what does it deliberately not preserve, and which part of it silently does nothing unless the transfer runs with superuser rights?

level: middleimportance: should knowfreq 52%
basics
~20 s

rsync's -a is shorthand for -rlptgoD: recurse, copy symlinks as symlinks, preserve permissions, times, group, owner, and device and special files. It does not cover hard links (-H), ACLs (-A), extended attributes (-X) or sparse files (-S), and preserving owner requires superuser rights on the receiving side.

open as a page