skip to content

Inside a Docker container, a connection to another service fails. How do you prove whether name resolution or the network path is at fault?

level: middleimportance: must knowfreq 71%

answer

  1. Two failures hide behind one sentence
  2. Resolve the name before blaming the path
  3. Refused and timed out mean different things
  4. getent hosts, then connect to the address
  5. 127.0.0.11 in resolv.conf reveals attachment

basics

~20 s

Resolve the name first: getent hosts <name> inside the container. No answer means a resolution problem — read /etc/resolv.conf and check which network the container is on. An answer followed by a refusal or a hang means the path or the peer's listener, not DNS.

solid answer

~50 s

Split the failure before touching anything. **Rung one:** run `getent hosts thumbs-index` inside the container; it goes through the same name-service path the application uses. No output is a resolution failure, and the next step is `cat /etc/resolv.conf` — `nameserver 127.0.0.11` means the container is on a user-defined network where container names are answerable, anything else means it is on the default bridge and the real problem is attachment, confirmed from the host with `docker inspect -f '{{json .NetworkSettings.Networks}}' <name>`. **Rung two:** if it resolves, connect to the address directly and read the error. *Connection refused* means the packet arrived and nothing was listening — wrong port, peer not up yet, or the peer bound to `127.0.0.1` inside its own container. *Timeout* means silence — wrong address, different network, or a filter. Then `ip route` confirms the default gateway matches the network you expect.

code

bash · 5 lines
bash
docker exec thumbs-worker getent hosts thumbs-index
# 172.29.0.7   thumbs-index
docker exec thumbs-worker cat /etc/resolv.conf
# nameserver 127.0.0.11
docker exec thumbs-worker nc -z -w2 172.29.0.7 8422; echo "exit=$?"

go deeper

for a junior

Know the two commands that split the problem: getent hosts <name> to test resolution and a direct connect to the resolved address to test the path. Recall that nameserver 127.0.0.11 in /etc/resolv.conf marks a user-defined Docker network.

for a middle

Explain what each outcome eliminates: no answer means resolution and attachment, refused means the path worked and the listener did not, a timeout means packets were dropped. Be able to read ip route and say which gateway you expected.

for a senior

Demonstrate a repeatable ladder under time pressure, including when the image has no shell and everything must be done from host-side docker inspect and docker port. Say out loud what each rung rules out before running the next one.

for a principal

Own the fact that this ladder gets run at 3 a.m. by someone who did not build the system. Argue for what makes it cheap: consistent network layout, services that bind the wildcard address, readiness gates instead of retry-forever clients, and a documented debugging path for shell-less images.

"It cannot connect" is three different failures wearing one sentence: the name did not resolve, the packet did not arrive, or the packet arrived and nothing was listening. The value of a method is that each rung produces evidence that eliminates a whole class, so you stop guessing at drivers, firewalls and DNS all at once. **Rung 0 — get the exact symptom.** Which name, which port, from which container, and what did the client actually print? "Name or service not known" / "UnknownHostException" is a resolution failure. "Connection refused" is an answered TCP handshake that was rejected. A hang followed by a timeout is silence — packets went nowhere. These three words carry most of the diagnosis. **Rung 1 — resolve the name.** From inside the container, `getent hosts thumbs-index` is the best single probe because it goes through the C library's name-service path, which is the path most runtimes actually use, rather than talking to a DNS server directly. An address means resolution works and you move on. No output means you have a resolution problem and there is no point testing routes yet. **Rung 2 — read `/etc/resolv.conf` inside the container.** `nameserver 127.0.0.11` tells you the container is attached to a user-defined Docker network and its lookups are being served by Docker's embedded resolver, which is where container names are answerable. A resolv.conf that instead lists your host's or corporate resolvers means the container is on the built-in default `bridge`, where other container names are simply not resolvable — a very common cause of "it works in Compose but not with plain `docker run`". At this point the question has changed from "is DNS broken?" to "is this container on the network I think it is?", which is an attachment question you answer from the host with `docker inspect -f '{{json .NetworkSettings.Networks}}' <name>`. **Rung 3 — connect to the address, not the name.** Once you have an address, try the port directly. If the image carries tools, `nc -z -w2 172.29.0.7 8422` or `wget -q -T2 -O- http://172.29.0.7:8422/health` from inside the container is enough. Now read the result: - **Connection refused** — the packet reached a network stack and something actively rejected it. The path is fine. The usual causes are the wrong port, the peer process not started yet, or the peer bound to `127.0.0.1` inside its own container, which is reachable only from that container's own loopback and never from a peer. - **Timeout / hang** — nothing came back. The path is the suspect: the wrong address, the two containers on different networks, an `--internal` network with no outbound route, or a host-level filter. - **Success** — the container network is exonerated and the bug is in the application, its TLS, or its client configuration. **Rung 4 — check the container's own view of the path.** `ip -4 addr show eth0` gives the container's address and prefix; `ip route` shows the default route, which should point at the gateway of the network it is attached to. If the default route's gateway does not match the `Gateway` recorded for that network, or the container has two interfaces and is sending traffic out of the wrong one, you have found a multi-homing problem rather than a DNS one. **Rung 5 — check the peer's listener.** `docker exec thumbs-index ss -ltn` (when the image has the tool) distinguishes `0.0.0.0:8422` from `127.0.0.1:8422` at a glance. This one line resolves more "refused" incidents than anything else. **A worked example.** A photo-thumbnail pipeline runs a Java batch job on a distroless base. Its hourly run fails immediately with `UnknownHostException: thumbs-index`. Rung 1 confirms no answer; rung 2 shows a resolv.conf with the host's upstream resolvers and no `127.0.0.11`, because the job was started with a bare `docker run` and landed on the default bridge while the rest of the stack sits on a user-defined network. Reattaching it fixes resolution — and the very next run fails differently, with "connection refused", because the index process binds `127.0.0.1:8422`. Two rungs, two genuinely different bugs, and neither would have been found by staring at network drivers. **The distroless caveat.** That image has no shell and no `getent`, so `docker exec` gives you nothing. The ladder still works, but it moves to the host side: take the container's address and network membership from `docker inspect`, take the peer's mappings from `docker port`, and run the connectivity probe from a container that *does* have tools on the same network. Recording the address from `docker inspect` before you start is what makes rung 3 possible at all.

  • Why prefer `getent hosts` over `nslookup` as the first probe from inside a container?
    `getent hosts` goes through the C library's name-service path, which is what most application runtimes actually use, so a success or failure there matches what the app experiences. `nslookup` talks to a DNS server directly and bypasses that path, so it can succeed while the application still fails — and many slim images ship neither `nslookup` nor `dig` while `getent` is present.
  • The name resolves and the port is right, yet every attempt is refused. What is the classic cause inside a container?
    The peer process is bound to `127.0.0.1` inside its own container. Loopback is per-network-namespace, so that socket is reachable only from that container itself — no peer on the network, and no published port, can reach it. `docker exec <peer> ss -ltn` shows `127.0.0.1:8422` instead of `0.0.0.0:8422`. The fix is in the application's bind address, not in Docker.
  • The failure is intermittent and only in the first seconds after the stack starts. How does that change your reading?
    An error that is refused early and succeeds on retry is a readiness signal, not a connectivity one — the peer's socket is not open yet. A slow-starting runtime makes the window wide enough to look like a network fault. Treat it as startup ordering: retry with backoff in the client, or gate the dependent process on a real readiness check rather than on the peer's container merely existing.

A failed delivery is either a bad address or a blocked road. Looking up the address first costs one command and tells you which half of the problem you are in.

saying these in an interview costs you the question

  • Jumps to firewall rules before resolving the name
  • Treats connection refused and timeout as the same symptom
  • Says pinging the container name proves the service works
  • Never checks which network the container is actually attached to
  • Assumes localhost inside a container reaches other containers
  • Blames DNS when the address resolved fine

context