skip to content

On a Linux host, `dig app.example.com` returns the address you expect, but a program running on that same host resolves the name to something else or fails outright. What explains the difference?

level: middleimportance: must knowfreq 56%

answer

  1. two different code paths
  2. the app never speaks DNS itself
  3. one file decides the source order
  4. the hosts line in nsswitch.conf
  5. getent, not dig, mirrors the app

basics

~20 s

They use different code paths. dig speaks DNS straight to a nameserver, while applications call getaddrinfo(), which follows the hosts: line of /etc/nsswitch.conf — static files, local modules, then DNS. A source ahead of DNS can answer first.

solid answer

~50 s

`dig` is a DNS query tool: it builds a DNS packet and sends it to a nameserver, and that is all it does. An ordinary application never speaks DNS itself — it calls `getaddrinfo()`, which walks the `hosts:` line in `/etc/nsswitch.conf` in order. A typical line puts `files` (that is, `/etc/hosts`) first, may include modules like `myhostname` or `resolve`, and only then `dns`. So a stale `/etc/hosts` entry, or any NSS module ahead of `dns`, answers before a DNS query is ever sent — and dig, which skips all of that, never sees it. The right diagnostic is `getent hosts <name>` or `getent ahosts <name>`, which goes through the same NSS path the application uses. Also check that the process really shares your `/etc`: a chroot or a separate mount namespace has its own files.

code

bash · 2 lines
bash
getent ahosts app.example.com
dig +short app.example.com

go deeper

for a junior

Recall that /etc/hosts is consulted before DNS on a typical system, so a line there overrides whatever the DNS server says for every program on the host.

for a middle

Explain the mechanism: applications call getaddrinfo(), which walks the hosts: line of /etc/nsswitch.conf module by module, while dig bypasses that entirely. Name getent as the tool that reproduces the application's path.

for a senior

Demonstrate the diagnostic instinct under pressure: compare the NSS path against a direct DNS query first, and if they agree, look at the process's own filesystem view and in-memory cache rather than continuing to debug the nameserver.

for a principal

Own the operational consequence of NSS ordering as a fleet-wide policy: emergency /etc/hosts pins are invisible to DNS-level monitoring and outlive the incident. Argue for how such overrides are recorded, expired and detected across many machines.

## Two different code paths The single most useful mental model for Linux name resolution is that there are **two independent paths**, and the tools you reach for first exercise the wrong one. - `dig`, `host` and `nslookup` are **DNS tools**. They construct a DNS query, send it to a nameserver, and print the reply. They are talking the DNS protocol on purpose, which is what makes them good for debugging DNS itself. - An application calls **`getaddrinfo()`** (or the legacy `gethostbyname()`), a C library function that answers the question "what addresses does this name have?" using *whatever sources this system is configured to use*. DNS is one of them, not the definition of the question. When the two disagree, the answer is almost always that something on the NSS path answered before DNS was consulted. ## The nsswitch hosts line `/etc/nsswitch.conf` maps each kind of lookup (users, groups, hosts, services) to an ordered list of *name service switch* modules. The relevant line looks something like: ``` hosts: files dns ``` or, on a systemd-resolved system, something closer to: ``` hosts: files resolve [!UNAVAIL=return] myhostname dns ``` Each word is a module, consulted left to right: - `files` reads `/etc/hosts`. - `dns` performs the DNS lookup, configured by `/etc/resolv.conf`. - `resolve` (nss-resolve) asks systemd-resolved over its bus API instead of by DNS packet. - `myhostname` synthesises answers for the machine's own hostname, `localhost`, and the gateway. - `mdns4_minimal` answers `.local` names via multicast DNS. The bracketed items are **actions** that control the walk: `[NOTFOUND=return]` or `[!UNAVAIL=return]` stop the sweep instead of falling through to the next module. That is how a `.local` name can be prevented from ever reaching DNS, and how `resolve` can be made authoritative when it is available. The first module that returns a result wins. Nothing merges answers, and nothing prefers the "better" one. ## Concrete causes of the divergence - **A stale `/etc/hosts` entry.** Someone pinned the name during an incident and never removed it. Every application on the host gets the pinned address; dig never looks at the file. - **A module ahead of `dns`.** `myhostname` answering the machine's own name, or an mDNS module claiming a `.local` suffix, produces answers no DNS server ever returned. - **A different `/etc` altogether.** A process running in a chroot, or with its own mount namespace, reads a different `/etc/hosts` and `/etc/resolv.conf` from the one your shell sees. Your view of the filesystem is not necessarily the process's view. - **A per-process cache.** Some runtimes and HTTP clients keep resolved addresses in memory, so the process's *last* lookup, not its next one, is what it is acting on. ## Prove it with the right tool `getent` performs the lookup through NSS, exactly as an application would: ```bash getent hosts app.example.com # single-family lookup via NSS getent ahosts app.example.com # the getaddrinfo path, both families dig +short app.example.com # DNS only, NSS bypassed ``` If `getent` and `dig` agree, the resolution path is consistent and your problem lies elsewhere — in the application's own cache, or in connectivity to the address rather than in resolving it. If they disagree, read the `hosts:` line and then `/etc/hosts`, in that order. One nuance worth knowing so you do not over-claim: on a systemd-resolved system, `dig` pointed at the local stub can *appear* to honour `/etc/hosts`, because resolved reads that file and synthesises answers from it. Querying an upstream server explicitly, for example `dig @10.0.0.53`, removes that ambiguity. ## Why interviewers like this question It separates "DNS is broken" reflexes from an understanding of the host. Candidates who only know `dig` will keep debugging a DNS server that is answering perfectly. The engineer who reaches for `getent` and `/etc/nsswitch.conf` finds a one-line `/etc/hosts` entry in under a minute.

  • What does `[NOTFOUND=return]` mean when it appears in the hosts line?
    It is an action controlling the walk. When the module to its left reports that the name definitively does not exist, the resolver stops there and returns that result instead of falling through to the next source. It is commonly used after an mDNS module so that a `.local` name never leaks out to DNS.
  • You run getent and dig from your shell and both look correct, yet the service still resolves the name wrongly. What now?
    Your shell may not share the service's view. Check whether the process runs in a chroot or its own mount namespace with a different /etc/hosts and /etc/resolv.conf, and whether it runs as a user with different configuration. If the files genuinely match, suspect the process's own in-memory cache and confirm by restarting it.
  • Why does the myhostname NSS module exist at all?
    So that a machine can always resolve its own hostname and localhost even when DNS is unavailable or misconfigured, which many programs assume works. It synthesises those answers locally rather than querying anything, which is also why the machine's own name can resolve to an address that no DNS zone contains.

saying these in an interview costs you the question

  • Assumes dig and the application take the same path
  • Treats /etc/resolv.conf as the whole resolution configuration
  • Never checks /etc/hosts before blaming the DNS server
  • Thinks /etc/nsswitch.conf lists DNS servers
  • Believes DNS overrides a matching /etc/hosts entry

context