skip to content

Linux internals & storage

The kernel-side subsystems you reach for when ordinary userspace tools stop explaining things: flexible storage with LVM, and in-kernel observability through eBPF and the tracing toolchain. Interviewers go here for senior systems and SRE roles, where the expectation is that you can see inside the kernel rather than only around it.

on this pageshow

explore

questions

page 2 of 2

How does a thin snapshot of an LVM thin logical volume differ from a classic `lvcreate -s -L` snapshot in the way it stores data, and why can you keep dozens of thin snapshots but not dozens of classic ones?

level: middleimportance: should knowfreq 45%

basics

~20 s

A classic LVM snapshot owns a fixed copy-on-write area and preserves old chunks by copying them into it. A thin snapshot just shares block references in the thin pool: a write allocates a new pool block, nothing is copied, and the cost does not grow with the number of snapshots.

open as a page

As an ordinary non-root user on a Linux host, `perf record` fails with a permission error about performance monitoring operations. Which sysctl governs that, what do its values mean, and what additionally blocks perf inside a container?

level: middleimportance: should knowfreq 38%

basics

~10 s

The sysctl is kernel.perf_event_paranoid, exposed at /proc/sys/kernel/perf_event_paranoid. Higher values restrict unprivileged use: 2 permits user-space measurement only, 1 also allows kernel profiling, -1 removes restrictions. Granting CAP_PERFMON is the privileged alternative.

open as a page

A Linux daemon is already running and burning most of its time in the kernel. What does `strace -c -p <pid>` give you that a full trace does not, and how do you read its output?

level: middleimportance: should knowfreq 45%

basics

~20 s

strace -c aggregates instead of printing every call: one row per system call with call count, error count and time. It answers which syscalls dominate and which fail, at the cost of losing the sequence in which they happened.

open as a page

An eBPF program counts events by looking a key up in a BPF_MAP_TYPE_HASH and doing (*val)++, and under load the totals come out too low. What is wrong, and what do __sync_fetch_and_add() and BPF_MAP_TYPE_PERCPU_HASH each change about it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Incrementing a value fetched from a shared map is a non-atomic read-modify-write, so CPUs running the hook concurrently overwrite each other's counts. __sync_fetch_and_add() makes the update atomic; a per-CPU map gives each CPU its own value, and user space sums across CPUs.

open as a page

In an XDP-based L4 load balancer, what is the difference between returning XDP_TX and returning XDP_REDIRECT for a forwarded packet, and why do such designs typically encapsulate the packet and use direct server return rather than plain NAT?

level: seniorimportance: should knowfreq 33%

basics

~20 s

XDP_TX resends the packet out the interface it arrived on; XDP_REDIRECT hands it to a different target through a redirect map. Encapsulation with direct server return lets backends reply straight to the client, so only the request half crosses the load balancer.

open as a page

You attach an XDP drop program to a 25 GbE Linux server's interface. The filter works, but throughput and CPU use are no better than the iptables rule it replaced, and `ip link show` reports the attachment as xdpgeneric. What are XDP's three attach modes, and what has happened here?

level: seniorimportance: should knowfreq 40%

basics

~20 s

XDP attaches in native mode inside the driver, generic mode in the shared kernel receive path after the sk_buff is allocated, or offloaded onto a capable NIC. Generic mode is a correctness fallback with none of the performance benefit — the driver did not support native XDP.

open as a page

A colleague loaded an XDP program onto a production interface from a terminal that has since been closed, and traffic is still being affected. Why do some eBPF attachments survive the process that created them while others vanish with it, and how would you find and remove this one?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Attachments made through a bpf_link end when the last file descriptor to that link closes, while netlink XDP attachments, tc filters and cgroup attachments hold their own kernel reference and outlive the loader. Those have to be detached explicitly.

open as a page

In eBPF, fentry and fexit programs (BPF_PROG_TYPE_TRACING) can target many of the same kernel functions as kprobe and kretprobe programs. How do they attach differently, and what does an fexit program see that a kretprobe cannot?

level: seniorimportance: should knowfreq 40%

basics

~20 s

fentry and fexit programs attach through a BPF trampoline at the compiler-inserted patch site and read typed arguments derived from BTF. An fexit program sees the function's arguments and its return value together; a kretprobe sees only the return.

open as a page

A libbpf CO-RE binary that works on a current Ubuntu server fails to load on an older Linux host, and `/sys/kernel/btf/vmlinux` does not exist there. What is missing, and what are your options for that host?

level: seniorimportance: should knowfreq 38%

basics

~20 s

That kernel was built without CONFIG_DEBUG_INFO_BTF, so it exposes no BTF and libbpf has nothing to resolve the binary's CO-RE relocations against. Options are a kernel built with BTF enabled, supplying an externally generated BTF for that exact kernel, or falling back to a runtime-compiled bcc tool.

open as a page

Your bpftrace one-liner runs fine as root on a Linux host but fails to load when you run it inside a container, and when you run it from the host against a containerised workload the PIDs and file paths it prints do not match what you see inside the container. Explain both problems and how you would work around each.

level: seniorimportance: should knowfreq 28%

basics

~20 s

Inside a container the loading fails because the bpf and perf_event_open system calls are usually blocked by the default seccomp profile and the required capabilities are dropped. From the host, the PIDs printed are the host's, and the file paths belong to the container's mount namespace, so neither matches what the container shows.

open as a page

A large eBPF program is rejected with the verifier reporting that it processed too many instructions, even though the program itself is far shorter than the instruction limit. What is the verifier counting, and how do you restructure the program so it loads?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The verifier counts instructions it simulates across every reachable path, not instructions in the program, so branches and loops multiply the work. Reduce path explosion: cut branches, help its bounds tracking, move iteration into bpf_loop(), and split logic into independently verified global subprograms or tail calls.

open as a page

You are adding a new data disk to a Linux server for LVM. Should you run pvcreate on the whole disk (/dev/sdb) or on a partition (/dev/sdb1), and what does each choice cost you?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Both work. A whole-disk PV is simpler and self-aligning but leaves no partition table, so other tools and humans can mistake the disk for blank. A partitioned PV advertises its use through the Linux LVM type code and lets the disk be shared, at the cost of an extra layer to manage.

open as a page

On a Linux server, /dev/sdb is one of three physical volumes in an LVM volume group serving live databases, and it has started logging SMART errors and I/O timeouts. How do you get the data off that disk and remove it from the volume group without downtime?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Confirm the remaining physical volumes hold enough free extents, then run pvmove /dev/sdb to relocate its extents onto them while the volumes stay online. Finish with vgreduce to drop the disk from the group and pvremove to clear its LVM label.

open as a page

Before a risky package upgrade you snapshot the Linux server's root-adjacent data volume with `lvcreate -s`. The upgrade goes badly and you want to roll back. What does `lvconvert --merge vg0/snap` do, and why might the rollback not take effect the moment the command returns?

level: seniorimportance: should knowfreq 40%

basics

~20 s

lvconvert --merge schedules the snapshot's preserved contents to be written back over the origin, reverting it to the snapshot's point in time and deleting the snapshot when finished. If the origin is open — mounted or in use — the merge is deferred until the volume is next deactivated and reactivated.

open as a page

An LVM thin pool created with `lvcreate -L 200G -T vg0/pool` backs thin volumes whose virtual sizes total 800G. What happens to those volumes when the pool's data space — or its metadata space — actually runs out, and how do you keep that from happening?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Once the pool has no free blocks, any write needing a new allocation fails or hangs, filesystems on the thin volumes go read-only or corrupt, and metadata exhaustion is worse — the pool goes read-only and needs offline repair. Prevent it with Data%/Meta% alerting, autoextend, and volume group headroom.

open as a page

A `perf record -g` capture of a production Linux service produces a report where nearly every call stack is only one or two frames deep. Why does frame-pointer unwinding fail, and what do `--call-graph dwarf` and `--call-graph lbr` do differently?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Plain -g means frame-pointer unwinding, and optimised x86-64 builds usually compile with -fomit-frame-pointer, freeing that register for general use — so there is no chain to walk. --call-graph dwarf copies part of the user stack per sample and unwinds it offline using DWARF debug data; --call-graph lbr uses the CPU's branch records.

open as a page

Running strace inside a Linux container fails with "ptrace: Operation not permitted" even though the shell is root. What is blocking it, and how would you trace the process anyway?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Root inside a container runs with a reduced capability set that normally omits CAP_SYS_PTRACE, so the ptrace attach is refused. Either grant that capability when the container starts, or run strace from the host against the process's host PID.

open as a page

You need packet filtering and load balancing at close to line rate on commodity Linux servers. How would you decide between an XDP/eBPF datapath and a kernel-bypass framework such as DPDK, and what does each approach cost you?

level: principalimportance: should knowfreq 42%

basics

~20 s

DPDK takes the NIC away from the kernel for maximum throughput, at the cost of the whole kernel network stack, its tooling, and dedicated busy-polling cores. XDP stays inside the kernel, slightly slower but selective, interrupt-driven and operable with normal tools. Choose by how much of the traffic is exceptional.

open as a page

You own the eBPF-based tracing tools for a fleet of Linux servers spanning several kernel versions and distributions. How would you decide between shipping bcc-based tools and libbpf CO-RE binaries, and what would you standardise in the build and distribution pipeline?

level: principalimportance: should knowfreq 26%

basics

~20 s

Standardise on libbpf CO-RE binaries for anything shipped, because they need no compiler or headers on production hosts, and keep bcc for ad-hoc work on lab machines. Then fix a minimum kernel and BTF baseline, a pinned build toolchain, and a load-test matrix covering every kernel in the fleet.

open as a page

Your XDP program on a Linux host is dropping flood packets with XDP_DROP, but `tcpdump -i eth0` on that host shows none of those packets at all. Why can tcpdump not see them, and where does its own filter actually run?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

tcpdump reads packets from an AF_PACKET socket that is fed later in the receive path, after the driver has built an sk_buff. Native XDP runs before that, so packets it drops never reach the tap. Count drops in a BPF map instead.

open as a page

In a libbpf project on Linux, what does `bpftool gen skeleton prog.bpf.o > prog.skel.h` produce, and how does the generated skeleton change the user-space loader code?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

It generates a C header that embeds the compiled BPF object and declares a struct named after it, with open, load, attach and destroy functions plus typed members for every map, program, link and global variable — replacing hand-written libbpf boilerplate and file loading at runtime.

open as a page

An LVM logical volume appears on Linux as /dev/vg0/data, /dev/mapper/vg0-data and /dev/dm-2. What is each of those three paths, and how does device-mapper name a volume whose VG or LV name contains a dash?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

/dev/dm-2 is the real device-mapper device node, whose number is assigned at activation and is not stable. /dev/mapper/vg0-data and /dev/vg0/data are udev-created symlinks to it. Dashes inside a VG or LV name are doubled in the /dev/mapper form.

open as a page

You run `perf stat` against a CPU-bound Linux process and see roughly 0.3 instructions per cycle together with a high cache-miss count. What is that telling you about the workload, and what would an IPC near 3 mean instead?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

Low instructions-per-cycle with heavy cache misses means the CPU is stalling on memory rather than doing work — the fix is data layout and access patterns. A high IPC means the core is genuinely retiring instructions, so the code is doing too much work, and you optimise the algorithm.

open as a page

On Linux, what does ltrace show that strace does not, and why does ltrace often print nothing useful for a Go binary or a statically linked program?

level: middleimportance: nice to knowfreq 25%

basics

~20 s

ltrace traces calls into shared libraries — malloc, strcmp, OpenSSL functions — which never appear in a system-call trace. It works by breakpointing dynamic-linking stubs, so a statically linked or Go binary that makes no such calls gives an empty trace.

open as a page

A long-running eBPF tool keys a BPF_MAP_TYPE_HASH by connection and never observes some of those connections close, so entries accumulate. What happens once the map reaches max_entries, and what does switching to BPF_MAP_TYPE_LRU_HASH change?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

A BPF hash map has a fixed max_entries; once full, inserting a new key fails with -E2BIG and, if the program ignores the return value, the event is silently dropped. BPF_MAP_TYPE_LRU_HASH instead evicts an approximately least-recently-used entry so inserts keep succeeding.

open as a page

In eBPF, a cgroup program and a BPF LSM program both sit outside the tracing hooks and both can make the kernel refuse an operation. What does each one attach to, and what does its return value control?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A cgroup program attaches to a cgroup v2 directory and governs every process inside it; a BPF LSM program attaches to a kernel security hook. In both cases the return value decides whether the kernel allows the operation to proceed.

open as a page

You want to time a function inside a running user-space application using bpftrace's `uprobe` and `uretprobe` probes. How does a uprobe actually attach to the process, roughly what does each hit cost compared with a kernel tracepoint, and which kinds of binary make this approach unreliable?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

A uprobe patches a breakpoint instruction into a private copy of the target's executable page, so every thread that reaches that address traps into the kernel and runs the BPF program. Each hit costs roughly a microsecond, far more than a tracepoint. Stripped, statically linked, Go and JIT-compiled binaries make it unreliable.

open as a page

You take an LVM snapshot of a mounted, actively written filesystem on a Linux host. What consistency does the resulting snapshot actually give you, and what does running `fsfreeze -f` on the mount point immediately before the snapshot add?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The snapshot is crash-consistent: it looks exactly like the volume would after a power cut, so a journalling filesystem mounts it after replaying its journal, but anything still buffered above the block layer is missing. fsfreeze -f flushes and quiesces the filesystem first, giving a cleanly-shut-down image instead.

open as a page

Most Linux distributions ship with unprivileged eBPF disabled via kernel.unprivileged_bpf_disabled. If the verifier already proves every program safe, why is unprivileged loading turned off, and how would you decide what privileges to grant eBPF tooling across a fleet?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

The verifier proves architectural safety, not speculative-execution safety, and it is itself a large piece of attack surface whose bugs have been privilege escalations. Distributions therefore disable unprivileged loading and expect eBPF tooling to run with CAP_BPF plus the capability its program type needs.

open as a page

When you standardise LVM provisioning across a fleet of Linux servers, how do you decide how much of each volume group to allocate to logical volumes up front rather than leaving extents unallocated?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Allocate lean and grow on demand. LVM growth is online and cheap, while shrinking is offline at best and impossible on XFS, so unallocated extents are the reversible position. Free extents also give operations like pvmove somewhere to relocate data.

open as a page

showing 31–60 of 60