Inside a Docker container, a connection to another service fails. How do you prove whether name resolution or the network path is at fault?
answer
- Two failures hide behind one sentence
- Resolve the name before blaming the path
- Refused and timed out mean different things
- getent hosts, then connect to the address
- 127.0.0.11 in resolv.conf reveals attachment
basics
~20 sResolve the name first: getent hosts <name> inside the container. No answer means a resolution problem — read /etc/resolv.conf and check which network the container is on. An answer followed by a refusal or a hang means the path or the peer's listener, not DNS.
solid answer
~50 sSplit the failure before touching anything. **Rung one:** run `getent hosts thumbs-index` inside the container; it goes through the same name-service path the application uses. No output is a resolution failure, and the next step is `cat /etc/resolv.conf` — `nameserver 127.0.0.11` means the container is on a user-defined network where container names are answerable, anything else means it is on the default bridge and the real problem is attachment, confirmed from the host with `docker inspect -f '{{json .NetworkSettings.Networks}}' <name>`. **Rung two:** if it resolves, connect to the address directly and read the error. *Connection refused* means the packet arrived and nothing was listening — wrong port, peer not up yet, or the peer bound to `127.0.0.1` inside its own container. *Timeout* means silence — wrong address, different network, or a filter. Then `ip route` confirms the default gateway matches the network you expect.
code
bash · 5 linesdocker exec thumbs-worker getent hosts thumbs-index
# 172.29.0.7 thumbs-index
docker exec thumbs-worker cat /etc/resolv.conf
# nameserver 127.0.0.11
docker exec thumbs-worker nc -z -w2 172.29.0.7 8422; echo "exit=$?"go deeper
Know the two commands that split the problem: getent hosts <name> to test resolution and a direct connect to the resolved address to test the path. Recall that nameserver 127.0.0.11 in /etc/resolv.conf marks a user-defined Docker network.
Explain what each outcome eliminates: no answer means resolution and attachment, refused means the path worked and the listener did not, a timeout means packets were dropped. Be able to read ip route and say which gateway you expected.
Demonstrate a repeatable ladder under time pressure, including when the image has no shell and everything must be done from host-side docker inspect and docker port. Say out loud what each rung rules out before running the next one.
Own the fact that this ladder gets run at 3 a.m. by someone who did not build the system. Argue for what makes it cheap: consistent network layout, services that bind the wildcard address, readiness gates instead of retry-forever clients, and a documented debugging path for shell-less images.
"It cannot connect" is three different failures wearing one sentence: the name did not resolve, the packet did not arrive, or the packet arrived and nothing was listening. The value of a method is that each rung produces evidence that eliminates a whole class, so you stop guessing at drivers, firewalls and DNS all at once. **Rung 0 — get the exact symptom.** Which name, which port, from which container, and what did the client actually print? "Name or service not known" / "UnknownHostException" is a resolution failure. "Connection refused" is an answered TCP handshake that was rejected. A hang followed by a timeout is silence — packets went nowhere. These three words carry most of the diagnosis. **Rung 1 — resolve the name.** From inside the container, `getent hosts thumbs-index` is the best single probe because it goes through the C library's name-service path, which is the path most runtimes actually use, rather than talking to a DNS server directly. An address means resolution works and you move on. No output means you have a resolution problem and there is no point testing routes yet. **Rung 2 — read `/etc/resolv.conf` inside the container.** `nameserver 127.0.0.11` tells you the container is attached to a user-defined Docker network and its lookups are being served by Docker's embedded resolver, which is where container names are answerable. A resolv.conf that instead lists your host's or corporate resolvers means the container is on the built-in default `bridge`, where other container names are simply not resolvable — a very common cause of "it works in Compose but not with plain `docker run`". At this point the question has changed from "is DNS broken?" to "is this container on the network I think it is?", which is an attachment question you answer from the host with `docker inspect -f '{{json .NetworkSettings.Networks}}' <name>`. **Rung 3 — connect to the address, not the name.** Once you have an address, try the port directly. If the image carries tools, `nc -z -w2 172.29.0.7 8422` or `wget -q -T2 -O- http://172.29.0.7:8422/health` from inside the container is enough. Now read the result: - **Connection refused** — the packet reached a network stack and something actively rejected it. The path is fine. The usual causes are the wrong port, the peer process not started yet, or the peer bound to `127.0.0.1` inside its own container, which is reachable only from that container's own loopback and never from a peer. - **Timeout / hang** — nothing came back. The path is the suspect: the wrong address, the two containers on different networks, an `--internal` network with no outbound route, or a host-level filter. - **Success** — the container network is exonerated and the bug is in the application, its TLS, or its client configuration. **Rung 4 — check the container's own view of the path.** `ip -4 addr show eth0` gives the container's address and prefix; `ip route` shows the default route, which should point at the gateway of the network it is attached to. If the default route's gateway does not match the `Gateway` recorded for that network, or the container has two interfaces and is sending traffic out of the wrong one, you have found a multi-homing problem rather than a DNS one. **Rung 5 — check the peer's listener.** `docker exec thumbs-index ss -ltn` (when the image has the tool) distinguishes `0.0.0.0:8422` from `127.0.0.1:8422` at a glance. This one line resolves more "refused" incidents than anything else. **A worked example.** A photo-thumbnail pipeline runs a Java batch job on a distroless base. Its hourly run fails immediately with `UnknownHostException: thumbs-index`. Rung 1 confirms no answer; rung 2 shows a resolv.conf with the host's upstream resolvers and no `127.0.0.11`, because the job was started with a bare `docker run` and landed on the default bridge while the rest of the stack sits on a user-defined network. Reattaching it fixes resolution — and the very next run fails differently, with "connection refused", because the index process binds `127.0.0.1:8422`. Two rungs, two genuinely different bugs, and neither would have been found by staring at network drivers. **The distroless caveat.** That image has no shell and no `getent`, so `docker exec` gives you nothing. The ladder still works, but it moves to the host side: take the container's address and network membership from `docker inspect`, take the peer's mappings from `docker port`, and run the connectivity probe from a container that *does* have tools on the same network. Recording the address from `docker inspect` before you start is what makes rung 3 possible at all.
- Why prefer `getent hosts` over `nslookup` as the first probe from inside a container?`getent hosts` goes through the C library's name-service path, which is what most application runtimes actually use, so a success or failure there matches what the app experiences. `nslookup` talks to a DNS server directly and bypasses that path, so it can succeed while the application still fails — and many slim images ship neither `nslookup` nor `dig` while `getent` is present.
- The name resolves and the port is right, yet every attempt is refused. What is the classic cause inside a container?The peer process is bound to `127.0.0.1` inside its own container. Loopback is per-network-namespace, so that socket is reachable only from that container itself — no peer on the network, and no published port, can reach it. `docker exec <peer> ss -ltn` shows `127.0.0.1:8422` instead of `0.0.0.0:8422`. The fix is in the application's bind address, not in Docker.
- The failure is intermittent and only in the first seconds after the stack starts. How does that change your reading?An error that is refused early and succeeds on retry is a readiness signal, not a connectivity one — the peer's socket is not open yet. A slow-starting runtime makes the window wide enough to look like a network fault. Treat it as startup ordering: retry with backoff in the client, or gate the dependent process on a real readiness check rather than on the peer's container merely existing.
A failed delivery is either a bad address or a blocked road. Looking up the address first costs one command and tells you which half of the problem you are in.
saying these in an interview costs you the question
- Jumps to firewall rules before resolving the name
- Treats connection refused and timeout as the same symptom
- Says pinging the container name proves the service works
- Never checks which network the container is actually attached to
- Assumes localhost inside a container reaches other containers
- Blames DNS when the address resolved fine