On a Linux host, what is the relationship between the `sysctl` command and the /proc/sys directory, and why does a value set with `sysctl -w` disappear after a reboot?
answer
- dots become slashes under /proc/sys
- the write is a command, not a record
- boot replays a config file, not your shell
- numeric prefixes decide who wins
- some values are sampled only once
basics
~20 ssysctl is a thin wrapper over /proc/sys: each dotted key maps to a path, so net.ipv4.ip_forward is /proc/sys/net/ipv4/ip_forward. Writes there change a live kernel variable and store nothing, so persistence requires a config file under /etc/sysctl.d.
solid answer
~40 sEvery kernel tunable is exposed as a file under /proc/sys, and `sysctl` simply translates a dotted key into that path: `net.ipv4.ip_forward` is `/proc/sys/net/ipv4/ip_forward`. `sysctl -w key=value` writes to that file, the kernel parses the bytes and updates an in-memory variable, and nothing is recorded anywhere — procfs has no storage — so the setting is gone at the next boot. To make it survive, you drop a `key = value` line in a file under `/etc/sysctl.d/`, which `systemd-sysctl.service` applies during boot; `sysctl --system` re-applies all of those files immediately so runtime and configuration agree. Two things bite people: a key only exists if the subsystem that owns it is loaded, and some values are only consulted at a specific moment, so setting them does not retroactively change objects that already exist.
code
bash · 5 linessysctl -n net.ipv4.ip_forward
sudo sysctl -w net.ipv4.ip_forward=1
cat /proc/sys/net/ipv4/ip_forward
printf 'net.ipv4.ip_forward = 1\n' | sudo tee /etc/sysctl.d/90-forwarding.conf
sudo sysctl --systemgo deeper
Know that sysctl reads and writes files under /proc/sys, that -w changes only the running kernel, and that persistence means putting the key in a file under /etc/sysctl.d.
Explain the mapping from dotted key to path, why nothing is stored by a write, and how the boot-time service applies sysctl.d files with lexicographic ordering and /etc masking /usr/lib.
Demonstrate the failure modes you have actually hit: keys missing because their module is not loaded yet, values sampled once at object creation so a live change does nothing, and config drifting from the running kernel until a reboot exposes it.
Own the policy question — which tunables are baked into the machine image versus set per workload, how you keep them reviewed and reproducible, and why blanket copy-pasted tuning files are a liability across a heterogeneous fleet.
## One namespace, two spellings The kernel's tunable parameters live in a single tree exposed at /proc/sys. The `sysctl` command from procps-ng is a convenience layer over that tree, and the mapping is purely mechanical: replace each dot with a slash and prefix /proc/sys. These three are the same operation: ``` sysctl net.ipv4.ip_forward sysctl -n net.ipv4.ip_forward cat /proc/sys/net/ipv4/ip_forward ``` and these two are the same write: ``` sysctl -w net.ipv4.ip_forward=1 echo 1 > /proc/sys/net/ipv4/ip_forward ``` `sysctl -a` dumps every key with its current value, which is the practical way to discover what exists on the kernel you are actually running — the set of keys depends on kernel version and on which subsystems are compiled in or loaded. Writing requires root, because /proc/sys entries are owned by root and mode 0644 for read-only-to-users keys or 0600/0644 with write reserved to root. ## Why the change evaporates Procfs stores nothing. When you write `1` into ip_forward, the kernel's handler parses the string and assigns an integer to a variable in kernel memory; there is no file, no block, no inode content. A reboot starts a fresh kernel with its compiled-in defaults, and nothing replays your write. That is the whole answer to the interview question: the write was a command, not a record. ## Making it persist Persistence is a userspace convention layered on top. Configuration files contain plain `key = value` lines and are read at boot: ``` # /etc/sysctl.d/90-forwarding.conf net.ipv4.ip_forward = 1 vm.swappiness = 10 ``` On a systemd distribution, `systemd-sysctl.service` runs early in boot and applies files from `/etc/sysctl.d/`, `/run/sysctl.d/` and `/usr/lib/sysctl.d/`, plus `/etc/sysctl.conf` for compatibility. Precedence has two layers that candidates routinely mix up. Files are processed in lexicographic order of their filename across all the directories, and later assignments win — so `99-local.conf` overrides `10-vendor.conf`. Separately, a file present in `/etc/sysctl.d/` masks a file of the same name in `/usr/lib/sysctl.d/`, which is how you override a package's defaults without editing its file. This is why the numeric prefix convention exists at all. `sysctl -p [file]` loads one file (defaulting to /etc/sysctl.conf), and `sysctl --system` walks the whole set of directories in the same order the boot service uses. Applying with `--system` after editing is the habit that prevents the classic mismatch where the file says one thing and the running kernel says another until somebody reboots months later. ## The two traps **A key only exists if its owner is loaded.** Tunables are registered by the subsystem that owns them, and many subsystems are kernel modules. `net.bridge.bridge-nf-call-iptables`, for example, does not exist until the `br_netfilter` module is loaded, so a sysctl.d file setting it at boot fails with "cannot stat /proc/sys/... : No such file or directory" if it is applied before the module. The fix is to arrange for the module to load first — for instance via a file in /etc/modules-load.d/ — rather than to blame sysctl. The same class of failure explains why a key you can set by hand after the system is up fails during boot. **A value is only consulted when the kernel looks at it.** Some tunables are read continuously and take effect instantly; others are sampled once, at the moment an object is created. `net.core.somaxconn` caps the accept queue that a socket is given when it calls `listen()`, so raising it at runtime does nothing for sockets that are already listening — the server has to create them again. Interviewers like this because it separates people who have actually tuned a machine from people who have read a tuning blog post. Sizes computed at boot from available memory, such as certain default buffer values, are similarly resistant to being changed later. A third detail worth knowing: much of the `net.*` tree is per network namespace rather than global, so the value you set on the host is not automatically what a differently-namespaced process sees. ## Reading a tunable you were asked about Good answers name real keys and say what they do rather than reciting a tuning cargo cult: `net.ipv4.ip_forward` turns the host into a router for IPv4, `vm.swappiness` biases the reclaim balance between page cache and anonymous pages, `fs.file-max` caps system-wide open file descriptors, `kernel.pid_max` sets where PID allocation wraps, and `net.ipv4.ip_local_port_range` bounds the ephemeral ports a client can source from. If you are not sure a key exists on the kernel in front of you, `sysctl -a | grep` answers it in a second — and saying that is a better answer than inventing a plausible-sounding key.
- You raise net.core.somaxconn while a busy server is running and the connection drops continue. Why?somaxconn is an upper bound applied when a socket calls `listen()`, not a live setting consulted per connection. Sockets that are already listening keep the queue length they were given, so the new ceiling only applies to sockets created afterwards — the server has to be restarted, or reload in a way that re-listens. It also does nothing unless the application asks for a larger backlog, since the effective queue is the smaller of the two.
- A file in /etc/sysctl.d sets a key correctly, but at boot the unit logs that the key cannot be found. What is the usual cause?The tunable belongs to a kernel module that has not been loaded yet, so the corresponding file under /proc/sys does not exist when the sysctl service runs. `net.bridge.*` keys before `br_netfilter` loads are the classic case. Arrange for the module to load first — for example via /etc/modules-load.d/ — rather than moving the sysctl file around.
- Two files under /etc/sysctl.d set the same key to different values. Which one wins?Files are processed in lexicographic order of filename, and the last assignment applied wins, so 99-local.conf beats 10-vendor.conf. Separately, a file in /etc/sysctl.d/ masks a same-named file in /usr/lib/sysctl.d/, which is the supported way to override a package's shipped defaults without editing the package's own file.
saying these in an interview costs you the question
- Thinks editing a file under /proc/sys makes the setting persistent
- Believes sysctl -w writes into /etc/sysctl.conf
- Assumes every sysctl takes effect on objects that already exist
- Says the numeric filename prefix in sysctl.d is only cosmetic
- Invents sysctl keys instead of checking sysctl -a