skip to content

questions

5

What do the `nameserver` lines in /etc/resolv.conf do on a Linux host, and are those servers used as a load-balanced pool or in a fixed order?

level: juniorimportance: must knowfreq 60%

answer

  1. one file, one library
  2. an ordered list, not a pool
  3. next server only after a timeout
  4. glibc reads at most three

basics

~20 s

Each nameserver line gives the C library's resolver one DNS server address to query. They are tried strictly in the listed order, not load-balanced: the next one is used only after the previous fails or times out, and glibc reads at most three.

solid answer

~50 s

/etc/resolv.conf configures the **stub resolver inside the C library**, not a daemon. Each `nameserver` line is one recursive server the resolver may query, and glibc uses at most the first three; extra lines are ignored. The list is failover, not a pool: the resolver sends the query to the first server, waits up to the `timeout` value (5 seconds by default), and only then moves to the next, repeating the whole list for `attempts` rounds (2 by default). Nothing balances load across them, and nothing marks a dead server as down permanently — so a broken first entry costs every lookup a timeout. `options rotate` makes glibc start at a different entry per process, which spreads load but does not remove that penalty. On most distributions the file is generated, so edit whatever produced it rather than the file itself.

go deeper

for a junior

Know that each nameserver line is one DNS server address, that they are tried top to bottom, and that at most three are used. Say plainly that this file configures the resolver inside the C library.

for a middle

Explain the walk: query the first server, wait out the timeout, move to the next, repeat for the configured number of attempts. Name the defaults (timeout 5 seconds, 2 attempts) and what options timeout: and rotate change.

for a senior

Show that you reason about the cost of failover in production: a dead first entry adds a fixed multi-second delay to every lookup, and the resolver never remembers it. Talk about bounding it with a local caching stub rather than by piling on more nameserver lines.

for a principal

Own the policy question: where retry and failover behaviour should live. Argue for a single local resolver component with one well-understood timeout budget instead of leaving thousands of processes each re-implementing failover through this file.

## What /etc/resolv.conf actually is `/etc/resolv.conf` is the configuration file for the **stub resolver**: a small piece of code that lives inside the C library (glibc on nearly every mainstream Linux distribution) and is linked into every program that resolves names. It is not a service, and the kernel never reads it. When your program calls `getaddrinfo()`, that resolver code runs *in your process*, reads this file, builds a DNS query, and sends it over UDP (falling back to TCP for large answers) to one of the addresses you listed. That also explains a common surprise: there is no "DNS service" to restart after changing it. Modern glibc notices that the file's timestamp changed and reloads it, but older versions and some runtimes read it once per process, so a very long-lived daemon may keep using the old contents until it is restarted. ## The keywords Four directives matter in practice: ``` nameserver 10.0.0.53 nameserver 10.0.1.53 search corp.example.com options timeout:1 attempts:2 ``` - `nameserver` — the IP address of a recursive resolver to query. One per line. - `search` — suffixes the resolver may append to a short, unqualified name before giving up. - `domain` — a legacy single-suffix form; `search` supersedes it. - `options` — tuning knobs for the walk described below (`timeout:`, `attempts:`, `rotate`, `ndots:`, `edns0`, and others). ## How the server list is walked The glibc resolver uses **at most three** `nameserver` entries; a fourth and beyond are parsed but never queried, which is why "we listed five for redundancy" is a false comfort. For a given lookup the resolver: 1. Sends the query to the first configured server. 2. Waits up to the per-query timeout — **5 seconds** by default, changeable with `options timeout:N`. 3. On timeout or an error such as ICMP port-unreachable, moves to the next server and repeats. 4. Having exhausted the list, starts the whole sweep again, up to `attempts` rounds — **2** by default. Only when every server in every round has failed does the resolver return a failure to the caller, which surfaces to the application as something like "Temporary failure in name resolution" (EAI_AGAIN). Note the difference between *no answer* and *a negative answer*: if the first server replies NXDOMAIN, that is a perfectly good answer and the resolver stops there. Failover happens on silence and errors, never on an answer you dislike. ## Failover, not load balancing, and not high availability Two consequences follow, and interviewers ask about both. First, **all traffic goes to the first server** while it is healthy. The second entry is cold standby. `options rotate` changes the starting index per process so different processes begin at different servers, which spreads query load, but it is opt-in and it does not help a single busy process much. Second, **failover is not free and it is not sticky**. The resolver does not remember that server one was dead; every lookup pays the timeout again. That turns a dead first nameserver into a uniform multi-second stall on every lookup rather than a clean failure. Shrinking `options timeout:1 attempts:2` bounds the damage, but the real fix is to remove or repair the bad entry, or to put a local caching stub in front so applications talk to loopback and one component owns the upstream retry policy. ## Who writes this file On a modern distribution `/etc/resolv.conf` is almost always generated rather than hand-authored. A DHCP client, NetworkManager, `systemd-networkd`, the `resolvconf` package, or `systemd-resolved` may own it; with `systemd-resolved` it is typically a symlink into `/run`, and the single `nameserver` line points at a loopback stub rather than at your real servers. Hand-editing then either gets reverted on the next network event or edits a runtime file that is regenerated at boot. Always establish who owns the file before you change it. ## What it does not control `/etc/resolv.conf` configures DNS only. It has no say over whether DNS is consulted at all — that is `/etc/nsswitch.conf` — and it does not cache anything: the glibc resolver holds no answers between calls.

  • If the first nameserver is unreachable, why do users report slowness rather than an outright failure?
    Because the resolver falls back rather than failing. Every lookup sends to the dead server first, waits out the timeout — five seconds by default — then gets a correct answer from the second server. The result is right but late, on every single lookup, since the resolver keeps no memory that the first server was down.
  • You put five nameserver lines in the file for redundancy. What actually happens?
    Only the first three are used by the glibc resolver; the rest are parsed and ignored, so the extra entries provide no redundancy at all. If you need more upstreams than that, run a local caching resolver on loopback, point resolv.conf at it, and let that component manage the larger upstream list and its retry policy.
  • What is the practical effect of `options rotate`?
    It makes the resolver start at a different entry in the nameserver list per process rather than always at the first, so query load is spread across the configured servers. It does not remove the failover cost: a process that starts on a dead server still waits out the full timeout before moving on.

saying these in an interview costs you the question

  • Thinks the listed nameservers are queried in parallel
  • Believes the resolver load-balances across all nameserver lines
  • Adds a fourth and fifth nameserver expecting extra redundancy
  • Says a dead first server is skipped after the first failure
  • Thinks changing resolv.conf requires restarting a DNS service

context

open as a page

On a Linux host, `dig app.example.com` returns the address you expect, but a program running on that same host resolves the name to something else or fails outright. What explains the difference?

level: middleimportance: must knowfreq 56%

basics

~20 s

They use different code paths. dig speaks DNS straight to a nameserver, while applications call getaddrinfo(), which follows the hosts: line of /etc/nsswitch.conf — static files, local modules, then DNS. A source ahead of DNS can answer first.

open as a page

On a Linux host running systemd-resolved, /etc/resolv.conf contains the single line `nameserver 127.0.0.53`. What is answering at that address, where are the real upstream servers configured, and why does editing that file usually not stick?

level: middleimportance: should knowfreq 47%

basics

~20 s

127.0.0.53 is systemd-resolved's local stub listener, a caching forwarder running on the host itself. The real upstream servers come from the network configuration or resolved.conf and are shown by resolvectl status. The file is a generated symlink, so edits are overwritten.

open as a page

Every name lookup on a Linux server pauses for roughly five seconds and then succeeds. The network is otherwise healthy and the returned addresses are correct. What is the host doing, and what would you change?

level: seniorimportance: should knowfreq 39%

basics

~20 s

Five seconds is the glibc resolver's default per-query timeout, so the host is waiting out a query that never gets answered before falling back. The usual causes are an unreachable first nameserver or one of the parallel A and AAAA queries being dropped.

open as a page

A DNS record was repointed an hour ago. A fresh lookup on a Linux host returns the new address, but a long-running service on that same host keeps connecting to the old one. Where can the stale address still be held?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Not in the C library, which caches nothing between calls. Look at a local caching resolver on the host, an NSS caching daemon, the process's own in-memory address cache, and an already-established connection that was opened to the old address and never closed.

open as a page