skip to content

Why does a host debugger fail to reach a JVM's JDWP port in a container despite -p 5005:5005?

level: middleimportance: must knowfreq 68%

answer

  1. Two hops, not one
  2. Which interface did the agent pick?
  3. Loopback inside is its own namespace
  4. JDK 9 narrowed the JDWP default
  5. address=*:5005 versus address=5005

basics

~20 s

Because the debug agent is listening on the container's own loopback address, so the published port forwards to nothing. Bind the agent to all interfaces instead - JDWP address=*:5005, debugpy --listen 0.0.0.0:5678 - and the attach succeeds.

solid answer

~40 s

Getting a debugger into a container is two hops, and publishing only fixes the first. `-p 5005:5005` tells the engine to forward host port 5005 into the container's network namespace, but it cannot create a listener: something inside must be bound to an address reachable on the container's interface. A JVM started with `-agentlib:jdwp=transport=dt_socket,server=y,suspend=n,address=5005` on JDK 9 or later binds JDWP to 127.0.0.1 *inside the container*, which is a different loopback from the host's, so the forwarded connection is refused. The fix is `address=*:5005`; the same trap catches debugpy (`--listen 0.0.0.0:5678`, not the default localhost) and delve (`--listen=:2345`). Then publish that port. On a shared machine, bind the host side narrowly with `-p 127.0.0.1:5005:5005`, because JDWP has no authentication.

code

bash · 5 lines
bash
docker run --rm \
  -p 127.0.0.1:5005:5005 \
  -p 127.0.0.1:8080:8080 \
  -e JAVA_TOOL_OPTIONS='-agentlib:jdwp=transport=dt_socket,server=y,suspend=n,address=*:5005' \
  indexer:debug

go deeper

for a junior

Remember the two separate steps: publish the port when you start the container, and start the process so its debug agent listens on 0.0.0.0 rather than localhost. Being able to say why the published port alone is not enough is the whole answer at this level.

for a middle

Be ready to explain that a container has its own network namespace and its own loopback interface, and to spell the flag correctly for at least one runtime - JDWP address=*:5005, debugpy --listen 0.0.0.0:5678. Mentioning that JDK 9 changed the JDWP default earns real credit.

for a senior

An interviewer expects you to diagnose it outside in - is the mapping there, then is anything bound inside - and to raise the security side unprompted: an unauthenticated debug socket is remote code execution, so bind the host side to 127.0.0.1 and keep agent flags out of the shipped image.

for a principal

Own the policy: how a team debugs a running service without any image ever shipping a debug listener by default, who may attach in which environment, and whether the debug path is a separate build target, a tunnel, or simply forbidden outside development.

### The two hops Attaching a debugger from your machine to a process inside a container crosses two boundaries, and each one has to be opened separately. The first hop is the engine's port publishing. `docker run -p 5005:5005` installs a forwarding rule from a port on the host to the same port inside the container's network namespace. That rule exists whether or not anything is listening on the other side - publishing a port is not the same as opening one. The second hop is the process itself. A TCP listener is bound to a specific address, not just a port. If the debug agent binds to 127.0.0.1, it accepts connections that arrive *on the container's own loopback interface*. Traffic forwarded in from the host arrives on the container's ethernet interface (its bridge-network address), not on its loopback, so the kernel inside the container has no socket to hand it to and answers with a reset. The host debugger reports connection refused, which reads exactly like a port-mapping problem and is not one. The crucial point beginners miss: `localhost` inside a container is not the host's `localhost`. Each container has its own network namespace and therefore its own loopback interface, unreachable from anywhere else. Binding a debug agent there makes it reachable only from processes already inside that container. ### The per-runtime spelling **JVM / JDWP.** The agent string is `-agentlib:jdwp=transport=dt_socket,server=y,suspend=n,address=*:5005`. The `*:` prefix means all interfaces. This is the version-sensitive part: before JDK 9 a bare `address=5005` bound to every interface, and a great deal of copied-around advice still shows that form. JDK 9 changed the default to loopback-only for safety (an open JDWP port is unauthenticated remote code execution), and the `*:` form was introduced to opt back in. So the same Dockerfile line that worked on JDK 8 silently stops working on a modern JRE base image. **Python / debugpy.** `python -m debugpy --listen 0.0.0.0:5678 app.py`. The short form `--listen 5678` means localhost only. **Go / delve.** `dlv --headless --listen=:2345 --api-version=2 exec /app/server`. The empty host in `:2345` means all interfaces; `--listen=127.0.0.1:2345` is the trap. A compiled-language service that has no built-in agent needs the debug server binary present in the image, which is usually why the debug image is a separate build stage rather than the one that ships. ### Confirming which hop is broken Work outside in. `docker port <container>` shows whether the mapping exists at all; if it prints nothing, the container was started without `-p` (or with `EXPOSE` only, which is metadata and publishes nothing). If the mapping is there and the connection is still refused, the mapping is not the problem and the agent's bind address is the next suspect - check the command line the process was actually started with, which is where an environment variable that was supposed to add the agent often turns out to be missing. A useful discriminator: if the agent were listening correctly but the app had crashed, you would usually see the container itself exited rather than a refused connection on a running container. ### Network placement Publishing is only needed when the debugger runs on the host. Two containers on the same user-defined bridge network reach each other directly by container or service name on the container port, so a debugger running in a sidecar container needs no `-p` at all - but it still needs the agent bound to something other than loopback, because the sidecar is in a different network namespace. With `--network host` on Linux there is no namespace boundary and no mapping, so a loopback-bound agent *is* reachable from the host; that is why the same image sometimes appears to work on one machine and not another. ### Do not leave it open JDWP, debugpy and delve all grant full control of the process to anyone who connects: read memory, evaluate arbitrary expressions, load classes. None of them authenticates. Publishing 5005 on `0.0.0.0` of a machine with a routable address hands that control to the network. Two habits keep this safe: bind the host side of the publish explicitly (`-p 127.0.0.1:5005:5005`), so only a local client or an SSH tunnel can connect; and keep the agent flags out of the image that ships, supplying them per run through an environment variable or a dedicated debug build target. An image whose default command always opens a debug port is a production incident waiting for a deploy.

  • You publish 5005 on a shared development server so the team can attach. What is wrong with that, and what would you do instead?
    JDWP is unauthenticated: anyone who can reach the port can evaluate arbitrary code as the application user. Publish to the host's loopback only (`-p 127.0.0.1:5005:5005`) and have people reach it over an SSH tunnel, so the network never carries an open debug socket. Keep the agent flags out of the shipped image so a stray deploy cannot open it either.
  • Your debugger runs in a second container on the same user-defined bridge network. Do you still need to publish the debug port?
    No. Publishing maps a host port into the container; container-to-container traffic on a shared user-defined network goes direct, so the debugger connects to `indexer:5005` by name with no `-p` at all. The agent still must not be bound to loopback, because the other container is in a different network namespace and only the container's bridge address is reachable.
  • Why does the same image sometimes attach fine with `--network host` on Linux even though the agent is bound to 127.0.0.1?
    With `--network host` the container shares the host's network namespace, so there is no boundary and no port mapping - the agent's loopback socket *is* the host's loopback socket, and a local debugger reaches it. That is a genuine behaviour difference, not a fluke, and it is a common reason a Linux workflow fails to reproduce elsewhere.

Publishing a port is like having the front desk forward calls to room 512. If the phone in room 512 is unplugged from the wall jack and only wired to an intercom inside the room, the forwarded call still rings nowhere.

saying these in an interview costs you the question

  • Assumes localhost inside the container is the host's localhost
  • Thinks -p makes an app reachable whatever it listens on
  • Confuses EXPOSE in the Dockerfile with actually publishing a port
  • Blames the host firewall before checking the agent's bind address
  • Believes host and container port numbers must match
  • Leaves an unauthenticated debug port published on 0.0.0.0

context