skip to content

Networking Stack

Configuring and inspecting Linux networking: interfaces and addresses, the routing table, name resolution, packet filtering with iptables or nftables, and socket inspection with ss. This is the toolkit for answering why a service is up but unreachable.

part ofLinux & distributionsoverview, primer and where to startread it →
on this pageshow

questions

28

What do the `nameserver` lines in /etc/resolv.conf do on a Linux host, and are those servers used as a load-balanced pool or in a fixed order?

level: juniorimportance: must knowfreq 60%

answer

  1. one file, one library
  2. an ordered list, not a pool
  3. next server only after a timeout
  4. glibc reads at most three

basics

~20 s

Each nameserver line gives the C library's resolver one DNS server address to query. They are tried strictly in the listed order, not load-balanced: the next one is used only after the previous fails or times out, and glibc reads at most three.

solid answer

~50 s

/etc/resolv.conf configures the **stub resolver inside the C library**, not a daemon. Each `nameserver` line is one recursive server the resolver may query, and glibc uses at most the first three; extra lines are ignored. The list is failover, not a pool: the resolver sends the query to the first server, waits up to the `timeout` value (5 seconds by default), and only then moves to the next, repeating the whole list for `attempts` rounds (2 by default). Nothing balances load across them, and nothing marks a dead server as down permanently — so a broken first entry costs every lookup a timeout. `options rotate` makes glibc start at a different entry per process, which spreads load but does not remove that penalty. On most distributions the file is generated, so edit whatever produced it rather than the file itself.

go deeper

for a junior

Know that each nameserver line is one DNS server address, that they are tried top to bottom, and that at most three are used. Say plainly that this file configures the resolver inside the C library.

for a middle

Explain the walk: query the first server, wait out the timeout, move to the next, repeat for the configured number of attempts. Name the defaults (timeout 5 seconds, 2 attempts) and what options timeout: and rotate change.

for a senior

Show that you reason about the cost of failover in production: a dead first entry adds a fixed multi-second delay to every lookup, and the resolver never remembers it. Talk about bounding it with a local caching stub rather than by piling on more nameserver lines.

for a principal

Own the policy question: where retry and failover behaviour should live. Argue for a single local resolver component with one well-understood timeout budget instead of leaving thousands of processes each re-implementing failover through this file.

## What /etc/resolv.conf actually is `/etc/resolv.conf` is the configuration file for the **stub resolver**: a small piece of code that lives inside the C library (glibc on nearly every mainstream Linux distribution) and is linked into every program that resolves names. It is not a service, and the kernel never reads it. When your program calls `getaddrinfo()`, that resolver code runs *in your process*, reads this file, builds a DNS query, and sends it over UDP (falling back to TCP for large answers) to one of the addresses you listed. That also explains a common surprise: there is no "DNS service" to restart after changing it. Modern glibc notices that the file's timestamp changed and reloads it, but older versions and some runtimes read it once per process, so a very long-lived daemon may keep using the old contents until it is restarted. ## The keywords Four directives matter in practice: ``` nameserver 10.0.0.53 nameserver 10.0.1.53 search corp.example.com options timeout:1 attempts:2 ``` - `nameserver` — the IP address of a recursive resolver to query. One per line. - `search` — suffixes the resolver may append to a short, unqualified name before giving up. - `domain` — a legacy single-suffix form; `search` supersedes it. - `options` — tuning knobs for the walk described below (`timeout:`, `attempts:`, `rotate`, `ndots:`, `edns0`, and others). ## How the server list is walked The glibc resolver uses **at most three** `nameserver` entries; a fourth and beyond are parsed but never queried, which is why "we listed five for redundancy" is a false comfort. For a given lookup the resolver: 1. Sends the query to the first configured server. 2. Waits up to the per-query timeout — **5 seconds** by default, changeable with `options timeout:N`. 3. On timeout or an error such as ICMP port-unreachable, moves to the next server and repeats. 4. Having exhausted the list, starts the whole sweep again, up to `attempts` rounds — **2** by default. Only when every server in every round has failed does the resolver return a failure to the caller, which surfaces to the application as something like "Temporary failure in name resolution" (EAI_AGAIN). Note the difference between *no answer* and *a negative answer*: if the first server replies NXDOMAIN, that is a perfectly good answer and the resolver stops there. Failover happens on silence and errors, never on an answer you dislike. ## Failover, not load balancing, and not high availability Two consequences follow, and interviewers ask about both. First, **all traffic goes to the first server** while it is healthy. The second entry is cold standby. `options rotate` changes the starting index per process so different processes begin at different servers, which spreads query load, but it is opt-in and it does not help a single busy process much. Second, **failover is not free and it is not sticky**. The resolver does not remember that server one was dead; every lookup pays the timeout again. That turns a dead first nameserver into a uniform multi-second stall on every lookup rather than a clean failure. Shrinking `options timeout:1 attempts:2` bounds the damage, but the real fix is to remove or repair the bad entry, or to put a local caching stub in front so applications talk to loopback and one component owns the upstream retry policy. ## Who writes this file On a modern distribution `/etc/resolv.conf` is almost always generated rather than hand-authored. A DHCP client, NetworkManager, `systemd-networkd`, the `resolvconf` package, or `systemd-resolved` may own it; with `systemd-resolved` it is typically a symlink into `/run`, and the single `nameserver` line points at a loopback stub rather than at your real servers. Hand-editing then either gets reverted on the next network event or edits a runtime file that is regenerated at boot. Always establish who owns the file before you change it. ## What it does not control `/etc/resolv.conf` configures DNS only. It has no say over whether DNS is consulted at all — that is `/etc/nsswitch.conf` — and it does not cache anything: the glibc resolver holds no answers between calls.

  • If the first nameserver is unreachable, why do users report slowness rather than an outright failure?
    Because the resolver falls back rather than failing. Every lookup sends to the dead server first, waits out the timeout — five seconds by default — then gets a correct answer from the second server. The result is right but late, on every single lookup, since the resolver keeps no memory that the first server was down.
  • You put five nameserver lines in the file for redundancy. What actually happens?
    Only the first three are used by the glibc resolver; the rest are parsed and ignored, so the extra entries provide no redundancy at all. If you need more upstreams than that, run a local caching resolver on loopback, point resolv.conf at it, and let that component manage the larger upstream list and its retry policy.
  • What is the practical effect of `options rotate`?
    It makes the resolver start at a different entry in the nameserver list per process rather than always at the first, so query load is spread across the configured servers. It does not remove the failover cost: a process that starts on a dead server still waits out the full timeout before moving on.

saying these in an interview costs you the question

  • Thinks the listed nameservers are queried in parallel
  • Believes the resolver load-balances across all nameserver lines
  • Adds a fourth and fifth nameserver expecting extra redundancy
  • Says a dead first server is skipped after the first failure
  • Thinks changing resolv.conf requires restarting a DNS service

context

open as a page

On a Linux host, what does the `default` entry in the output of `ip route show` mean, and what happens to a packet whose destination matches no route in the table at all?

level: juniorimportance: must knowfreq 75%

basics

~20 s

The default route is the 0.0.0.0/0 catch-all entry telling the kernel where to send packets that no more specific route matches, normally via a gateway. With no match at all the send fails locally and immediately with ENETUNREACH, "Network is unreachable".

open as a page

A client opening TCP connections to a Linux host gets an instant "connection refused" on one port, while a connect to another port on the same host hangs for around two minutes before failing. What is the kernel doing differently in each case?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Connection refused means the SYN reached the host and the kernel answered with a TCP RST because no socket was listening on that port. A hang means no answer came back at all, usually a dropped packet, so the client keeps retransmitting the SYN until it gives up.

open as a page

On a Linux host, `dig app.example.com` returns the address you expect, but a program running on that same host resolves the name to something else or fails outright. What explains the difference?

level: middleimportance: must knowfreq 56%

basics

~20 s

They use different code paths. dig speaks DNS straight to a nameserver, while applications call getaddrinfo(), which follows the hosts: line of /etc/nsswitch.conf — static files, local modules, then DNS. A source ahead of DNS can answer first.

open as a page

You run `ip addr add 10.0.0.50/24 dev enp3s0` on a Linux server; the address works immediately, but it is gone after a reboot — and sometimes it disappears minutes later without one. What is happening, and where does the address actually belong?

level: middleimportance: must knowfreq 62%

basics

~20 s

ip only edits live kernel state, which is never written to disk, so a reboot discards it. Persistence belongs to whichever manager owns the interface — NetworkManager, systemd-networkd, or netplan rendering to one of them — and that manager can also overwrite your address while running.

open as a page

A packet arrives on a Linux host's network interface. Which netfilter hooks does it traverse if it is addressed to a local process, and which if the host forwards it to another machine? Where does the routing decision sit in that sequence?

level: middleimportance: must knowfreq 72%

basics

~20 s

Netfilter has five hooks. An arriving packet hits PREROUTING first; the routing decision then sends it either to INPUT and a local socket, or to FORWARD and then POSTROUTING on its way out. Locally generated packets take OUTPUT then POSTROUTING.

open as a page

In Linux netfilter, what is the difference between SNAT, MASQUERADE and DNAT? At which hook must each be applied, and when is MASQUERADE the better choice than plain SNAT?

level: middleimportance: must knowfreq 66%

basics

~20 s

DNAT rewrites a packet's destination and runs in PREROUTING, before routing. SNAT rewrites the source in POSTROUTING, after routing. MASQUERADE is SNAT that takes the new source address from the outgoing interface automatically, which suits a changing address.

open as a page

A Linux routing table contains both `10.0.0.0/8 via 192.168.1.9 dev eth0 metric 500` and `default via 192.168.1.1 dev eth0 metric 100`. Which route does the kernel use for destination 10.1.2.3, and what does the metric actually decide?

level: middleimportance: must knowfreq 65%

basics

~20 s

The 10.0.0.0/8 route wins: the kernel matches the longest (most specific) prefix first, and 10.1.2.3 falls inside it. The metric never compares different prefixes — it only breaks ties between routes to the same prefix, where the lowest value wins.

open as a page

You stop a busy TCP server on Linux and start it again immediately, and it refuses to start with "Address already in use" even though no process is holding the port. Why does bind() fail, and what makes the restart succeed?

level: middleimportance: must knowfreq 72%

basics

~20 s

Connections the old server closed are still in TIME_WAIT, so the kernel still has sockets on that local port and refuses a plain bind. Setting SO_REUSEADDR on the listening socket before bind lets the server take the port back despite them.

open as a page

Why does a modern Linux server call its network card something like `enp3s0` or `eno1` rather than `eth0`, where do those names come from, and what problem was that change introduced to solve?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Kernel names like eth0 were assigned in driver probe order, which could change between boots on multi-NIC machines. systemd-udevd instead derives a stable name from firmware index, PCI slot or bus path, giving eno1, ens3 or enp3s0.

open as a page

In Linux netfilter — the kernel framework behind both iptables and nftables — what are the filter, nat and mangle tables for, and what changes about a packet's fate depending on which of them a rule lives in?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Netfilter tables group rules by the kind of work they do: filter decides whether a packet is accepted or dropped, nat rewrites source or destination addresses and ports, and mangle alters packet fields such as TTL, DSCP or the firewall mark.

open as a page

On a Linux host running systemd-resolved, /etc/resolv.conf contains the single line `nameserver 127.0.0.53`. What is answering at that address, where are the real upstream servers configured, and why does editing that file usually not stick?

level: middleimportance: should knowfreq 47%

basics

~20 s

127.0.0.53 is systemd-resolved's local stub listener, a caching forwarder running on the host itself. The real upstream servers come from the network configuration or resolved.conf and are shown by resolvectl status. The file is a generated symlink, so edits are overwritten.

open as a page

On Linux, what is a veth device, why is it always created as a pair, and what happens when you attach one end of that pair to a Linux bridge device?

level: middleimportance: should knowfreq 42%

basics

~20 s

A veth is a virtual Ethernet device created in pairs acting like a patch cable: a frame transmitted on one end is received on the other. Attaching one end to a bridge makes it a port on an in-kernel layer-2 switch that learns MACs and forwards between ports.

open as a page

Linux connection tracking assigns every packet a state such as NEW, ESTABLISHED, RELATED or INVALID. What does each of those mean, and why does a stateful firewall need the RELATED state at all?

level: middleimportance: should knowfreq 60%

basics

~20 s

NEW is a packet starting a flow the kernel has not seen, ESTABLISHED belongs to a flow already seen in both directions, RELATED is a separate flow spawned by an existing one such as an ICMP error, and INVALID fits no tracked flow at all.

open as a page

A Linux box with two NICs is supposed to route traffic between two subnets, but packets arriving on one interface are never seen leaving the other. What does the sysctl `net.ipv4.ip_forward` control, and how do you set it so the change survives a reboot?

level: middleimportance: should knowfreq 58%

basics

~20 s

By default Linux behaves as a host, not a router: it drops IP packets that are not addressed to it. Setting net.ipv4.ip_forward=1 makes the kernel forward them between interfaces. Persist it in a file under /etc/sysctl.d/ and apply with sysctl --system.

open as a page

On a Linux host where a web server and an application server run side by side, what do you gain and what do you give up by connecting them over a Unix domain socket instead of a TCP socket on 127.0.0.1?

level: middleimportance: should knowfreq 45%

basics

~20 s

A Unix domain socket skips the TCP/IP stack entirely: no handshake, no ephemeral ports, no TIME_WAIT, lower overhead, and access is controlled by filesystem permissions on the socket file. The cost is that it only works on the same host and leaves a stale file behind after an unclean exit.

open as a page

Every name lookup on a Linux server pauses for roughly five seconds and then succeeds. The network is otherwise healthy and the returned addresses are correct. What is the host doing, and what would you change?

level: seniorimportance: should knowfreq 39%

basics

~20 s

Five seconds is the glibc resolver's default per-query timeout, so the host is waiting out a query that never gets answered before falling back. The usual causes are an unreachable first nameserver or one of the parallel A and AAAA queries being dropped.

open as a page

After a Linux host is moved behind an encapsulating tunnel, small requests and ICMP echoes work fine, but SSH sessions freeze right after login and large transfers stall completely. Explain the mechanism that produces this size-dependent failure and how you would confirm it.

level: seniorimportance: should knowfreq 48%

basics

~20 s

Encapsulation shrinks the usable payload below the interface MTU, so full-size TCP segments are too big for the path. When the ICMP messages that would report this are filtered, the sender never learns and simply retransmits forever — a path MTU black hole.

open as a page

A busy Linux NAT gateway starts refusing new connections while existing ones keep working, and the kernel log shows "nf_conntrack: table full, dropping packet". What is filling up, what sets its ceiling, and what are your realistic options?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The kernel's connection-tracking table is full, so packets that would create a new flow entry are dropped while tracked flows continue. The ceiling is net.netfilter.nf_conntrack_max. Options: raise the limit and hash size, shorten timeouts, or exempt bulk traffic from tracking.

open as a page

A Linux server has two NICs on two different networks, each with its own gateway. Clients on the first network reach it fine; clients on the second reach it only sporadically or not at all. Explain what the kernel is doing with the reply packets, and how policy routing fixes it.

level: seniorimportance: should knowfreq 45%

basics

~20 s

With one routing table there is one default route, so replies to off-subnet clients on the second network leave through the first NIC's gateway. That asymmetry is dropped by reverse-path filtering or by stateful middleboxes. The fix is a second routing table selected by a source-based rule.

open as a page

A Linux service that opens many short-lived outbound TCP connections to a single backend address starts failing with "cannot assign requested address". What resource has run out, what governs its size, and how would you fix it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It has run out of ephemeral source ports for that destination. Linux picks them from net.ipv4.ip_local_port_range, and each closed connection holds its port in TIME_WAIT for 60 seconds, so a high connect rate exhausts the range. Reusing connections through a keep-alive pool is the durable fix.

open as a page

During a burst of new connections, clients of a Linux TCP server wait several seconds before their first request gets any response, while the server's own application logs show nothing wrong. Explain the two kernel queues behind listen() and what happens when the accept queue fills.

level: seniorimportance: should knowfreq 50%

basics

~20 s

A listening socket has two kernel queues: a SYN queue for half-open handshakes and an accept queue of completed connections waiting for the application to call accept(). When the accept queue overflows, Linux drops the client's final handshake packet by default, so the server retransmits its SYN-ACK and the client stalls for seconds.

open as a page

A multi-homed Linux server holds several IP addresses, and connections it initiates arrive at the peer carrying an unexpected source address. How does the kernel choose the source address for a locally generated packet, and how can you pin it?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

Unless the application binds a specific address, the source IP is a by-product of the routing decision: the kernel uses the chosen route's preferred-source attribute if it has one, otherwise an address of the egress interface. ip route get <dst> prints the address that would be used.

open as a page

A DNS record was repointed an hour ago. A fresh lookup on a Linux host returns the new address, but a long-running service on that same host keeps connecting to the old one. Where can the stale address still be held?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Not in the C library, which caches nothing between calls. Look at a local caching resolver on the host, an NSS caching daemon, the process's own in-memory address cache, and an already-established connection that was opened to the old address and never closed.

open as a page

Two 10 Gb NICs on a Linux server are bonded together with LACP (802.3ad), yet a single large TCP transfer never goes faster than 10 Gb/s. Why, and what would actually make use of both links?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A bond chooses an outgoing member by hashing packet header fields, and a single TCP connection has constant header values, so every packet takes the same link. Aggregation adds capacity across many flows, never within one.

open as a page

On current Debian and RHEL, the iptables command is a compatibility front-end that writes nftables rules into the kernel. What does that change about how rules are stored, and what goes wrong on a host where some software still uses the legacy iptables backend?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

The iptables command translates classic syntax into nf_tables rules, so one ruleset holds everything. The legacy binaries write to the older x_tables subsystem instead; if both are in use, each tool shows only half the rules while the kernel enforces both.

open as a page

Linux's SO_REUSEPORT lets several processes hold listening sockets on the same TCP port. How would you decide whether to scale a service that way rather than having one process accept connections and hand them to workers?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

SO_REUSEPORT gives every worker its own accept queue and lets the kernel spread new connections across them by hashing the connection's four-tuple, removing the single-acceptor bottleneck. The tradeoffs are hashed rather than balanced load, and connections lost from a worker's queue when it exits.

open as a page