Your Go service resolves a hostname on a laptop but not in its new container. How do you diagnose it?
answer
- ask which resolver ran, not why DNS failed
- the package has a debug level
- a laptop and a Linux container rarely match
- does the image even have the file
- the error object names the server that answered
basics
~20 sEstablish which resolver each environment used before theorising. Run with GODEBUG=netdns=go+2 in both places to see the resolver choice and lookup order, then compare the resolver configuration, the hosts file and the C-resolver environment variables between them.
solid answer
~50 sThe first question is not *why does DNS fail* but *which resolver ran*. A laptop very often resolves through the system's C stack while a Linux container defaults to Go's own DNS client, so the two are not running the same code path at all. Set `GODEBUG=netdns=go+2` (or `netdns=2` to observe without forcing) and the `net` package prints which resolver it selected and the lookup order it derived for each name. Then compare the inputs: `/etc/resolv.conf` and `/etc/hosts` inside the container against the laptop's, whether the image contains those files at all, and whether `LOCALDOMAIN`, `RES_OPTIONS` or `HOSTALIASES` are set anywhere. Classify the failure with `*net.DNSError` — `IsNotFound` points at the name or the search list, `IsTimeout` at reachability — and log `DNSError.Server` so you know which nameserver actually answered. Finally, fix the divergence deliberately by pinning the resolver, rather than letting each environment pick.
code
text · 3 lines$ GODEBUG=netdns=go+2 ./svc
go package net: GODEBUG setting forcing use of Go's resolver
go package net: hostLookupOrder(api.internal) = files,dnsgo deeper
Know that a name-resolution failure has a configuration side as well as a network side, and that the error value itself tells you which nameserver was asked.
Be able to name the GODEBUG knob that reveals the resolver decision, and to explain why the same binary can resolve differently on two machines.
Walk the investigation in order: which resolver ran, what configuration fed it, how the error classifies, and only then the network. Say what you would change so the environments stop diverging.
Frame it as environment parity rather than one incident. Decide who guarantees that development, CI and production resolve identically, and what a base-image change is allowed to alter underneath running services.
## Why this failure shape is so common A service that resolves a name fine during development and fails after being moved into a new cluster is rarely a Go bug and rarely a DNS outage. It is almost always that the two environments are not running the same resolver, are not reading the same configuration, or both. Go makes this easy to miss because the API is identical either way: `net.LookupHost` returns the same types whether Go spoke DNS itself or the C library did the work. ## Step 1 — find out which resolver ran Do this before forming any hypothesis. The `net` package has a debug channel built in: - `GODEBUG=netdns=2` — observe only. It prints what the package decided, without changing the decision. - `GODEBUG=netdns=go+2` — force Go's own resolver *and* print. Useful when you want both environments on the same code path so the comparison is apples to apples. - `GODEBUG=netdns=cgo+2` — the mirror image. The output tells you two things: whether a resolver was forced, and the lookup order the package derived for a given hostname — for example the hosts file first, then DNS. Wording differs between releases, but the substance is stable. Run it in both places. Very often the answer is immediately visible: the laptop is using the system resolver, and the container is using Go's own. ## Step 2 — compare the inputs Once you know the code path, compare what feeds it: - **`/etc/resolv.conf` inside the container.** Different nameservers, a different search list, and different options than the laptop's. This is the file Go's resolver reads for its nameservers. - **Does the image contain it at all?** A minimal image built from scratch may have no `/etc/resolv.conf` and no `/etc/hosts`. With no usable configuration, Go's resolver falls back to trying a nameserver on the loopback address, and in a container with nothing listening there, every lookup fails in a way that looks like a network problem. - **`/etc/hosts`.** A name that resolves on the laptop only because of a hosts entry will not resolve anywhere else, and the debug output showing the `files` step succeeding locally makes that obvious. - **Environment variables.** `LOCALDOMAIN`, `RES_OPTIONS` and `HOSTALIASES` are C-resolver knobs whose presence pushes Go onto the C path. If one of them is set in a developer shell profile and not in the deployment, the two environments differ before any DNS packet is sent. - **Is the name reachable from the container's network at all?** A name that only exists in a private zone the laptop reaches over a VPN will not resolve from the cluster, and that is a routing and zone question, not a Go question. ## Step 3 — classify the error, do not read the message Instrument the lookup so the failure is self-describing: ```go var dnsErr *net.DNSError if errors.As(err, &dnsErr) { log.Printf("lookup %s via %s: notfound=%v timeout=%v", dnsErr.Name, dnsErr.Server, dnsErr.IsNotFound, dnsErr.IsTimeout) } ``` `IsNotFound` and `IsTimeout` split the problem space in half. `IsNotFound` means something answered authoritatively: the name is wrong, the zone is wrong, or the query the resolver actually sent was not the one you assumed. `IsTimeout` means nothing answered: the nameserver address is wrong, unreachable, or overloaded. `Server` names which nameserver produced the answer, which resolves arguments about whether the container is even talking to the resolver you think it is. ## Step 4 — remove the divergence rather than patch the symptom Having found the difference, the durable fix is to stop letting each environment choose. Pin the resolver so development, CI and production run the same path — the run-time environment variable if you want it reversible, a build-level choice if you want it guaranteed in the artefact — and make the resolver configuration part of what you review when a base image changes. A one-off fix to the container's configuration leaves the next environment free to diverge again. ## Anti-patterns to avoid saying Jumping straight to packet captures before checking which resolver ran; concluding *DNS is broken* from a single failed lookup with no error classification; adding a retry loop around a lookup that is failing with `IsNotFound`; and hard-coding an IP address to make the symptom go away, which converts a resolution problem into a much worse problem the day that address changes.
- The container's lookup fails with IsNotFound while the laptop succeeds. What do you suspect first?Something answered authoritatively, so the query itself differs. Either the name resolves on the laptop only through a hosts entry or a private zone reachable there, or the container's search list turns the short name into a different fully-qualified query. Compare the hosts file and the resolver configuration between the two, and log `DNSError.Name` to see exactly what was queried.
- What happens if the container image contains no /etc/resolv.conf at all?Go's resolver has no nameservers to read, so it falls back to trying a nameserver on the loopback address. In a minimal image with nothing listening there, every lookup fails, and the symptom looks like a network fault rather than missing configuration. Checking for the file's presence is a fast early step when a scratch-based image is involved.
- Why run with GODEBUG=netdns=2 rather than going straight to netdns=go?Because `netdns=2` observes without changing the decision, and the decision is the evidence you are trying to collect. Forcing the Go resolver first may make the symptom disappear, which tells you the resolvers differ but destroys the information about what the unforced environment was doing and why.
- How would you keep this class of failure from recurring?Stop letting the environment choose. Pin the resolver so development, CI and production use the same path, keep the resolver configuration in scope when a base image changes, and add a start-up self-check that resolves a known name and logs the resolver decision so a divergence is visible in the first log line rather than during an incident.
saying these in an interview costs you the question
- Starts with packet captures before checking which resolver ran
- Assumes the laptop and the Linux container use the same resolver
- Concludes DNS is broken without classifying the error
- Never checks whether the image contains a resolver configuration file
- Hard-codes an IP address to make the symptom disappear