skip to content

How do you make a containerized process wait at startup until your debugger attaches?

level: juniorimportance: should knowfreq 52%

answer

  1. Attaching late misses everything early
  2. Let the agent hold the process
  3. One switch waits, another listens
  4. suspend=y, --wait-for-client
  5. Pass it at run time, never ship it

basics

~10 s

Start the debug agent in suspend mode: JDWP suspend=y, debugpy --wait-for-client, or dlv exec in headless mode. The process then blocks before your code runs, so a debugger attaching afterwards still catches early startup.

solid answer

~40 s

In a container you cannot race the process by hand, so the agent has to hold it for you. With the JVM, `-agentlib:jdwp=transport=dt_socket,server=y,suspend=y,address=*:5005` blocks before `main` until a client connects. Python's debugpy does the same with `--wait-for-client`; delve's headless `dlv exec` holds the program until a client tells it to continue. Supply these per run rather than baking them into the shipped image - `docker run -e JAVA_TOOL_OPTIONS=...` lets the JVM pick the agent up without rebuilding, and a dedicated debug build target works when the flags belong in the command. Two side effects to expect: a suspended process fails its HEALTHCHECK and shows as unhealthy, and a restart policy plus a suspended start can leave a container that looks hung rather than crashed.

code

bash · 4 lines
bash
docker run --rm -it \
  -p 127.0.0.1:5005:5005 \
  -e JAVA_TOOL_OPTIONS='-agentlib:jdwp=transport=dt_socket,server=y,suspend=y,address=*:5005' \
  indexer:1.4.2

go deeper

for a junior

Know that the debug agent can hold the process at startup and name at least one flag - JDWP suspend=y or debugpy --wait-for-client - and know why it matters: a container starts too fast to attach by hand before startup code has run.

for a middle

Explain that server=y and suspend=y control different things, and show how to add the agent to an existing image without rebuilding, typically through an environment variable the runtime already reads or by overriding the container's command for one run.

for a senior

Talk about the blast radius: a suspended container fails health checks, a restart policy makes it re-suspend on every start, and time keeps running while you sit on a breakpoint, so pools and leases can expire under you. Keep debug flags out of shipped images by construction.

for a principal

Decide how debugging is made available at all: whether debug capability is a separate build target, which environments permit an attach, and what stops a debug-enabled image from ever being promoted to production.

### Why suspend exists at all Outside a container you can start a process under a debugger and it is stopped from instruction one. Attaching to something already running is different: whatever happened before you connected is gone. That is fine for a request handler you can trigger again, and useless for a bug in startup - a configuration value read once, a database migration, a cache warmed at boot. Containers make the gap worse. The process starts the moment the container starts, and by the time you notice the container is up, click attach in your editor and the connection completes, several seconds have passed. For a short-lived container - a job that indexes documents and exits - the process may be gone entirely. The answer is to let the debug agent block the program until a client connects. ### The spelling per runtime **JVM.** `-agentlib:jdwp=transport=dt_socket,server=y,suspend=y,address=*:5005`. Two independent switches are often confused. `server=y` means the JVM *listens* and the debugger dials in (the normal container arrangement, because the container has the published port); `server=n` means the JVM dials out to a listening debugger. `suspend=y` is orthogonal: it decides whether the JVM waits before running your code. The startup line printed by the JVM - `Listening for transport dt_socket at address: 5005` - is your confirmation that the agent is up and waiting. **Python.** `python -m debugpy --listen 0.0.0.0:5678 --wait-for-client -m indexer`. Without `--wait-for-client` the module runs immediately and an attach later may miss everything interesting. **Go.** `dlv --headless --listen=:2345 --api-version=2 exec /app/indexer` starts the program under the debugger and holds it, so a client that connects can set breakpoints before anything runs and then continue. ### Getting the flags in without rebuilding The worst version of this workflow is editing the Dockerfile, rebuilding and hoping. Two better routes: Override the command for one run. The image's ENTRYPOINT/CMD is a default; `docker run` can replace the arguments for a single container, so the debug variant of the command is a shell-history line rather than a commit. Use an environment variable the runtime already reads. The JVM reads `JAVA_TOOL_OPTIONS` and prepends whatever it finds to the command line, printing `Picked up JAVA_TOOL_OPTIONS:` as it does. So `docker run -e JAVA_TOOL_OPTIONS='-agentlib:jdwp=...'` turns on debugging in an image built with no knowledge of debugging at all - the single most useful trick in this area, and the reason a Java service rarely needs a separate debug image. When the flags genuinely must live in the image - a Go binary that needs the delve server present, for instance - put them in a dedicated build stage and build that stage explicitly, so the artefact that ships is a different image from the one you debug. ### The side effects a suspended container produces A process stopped before `main` has not opened its service port and is not answering anything. Expect these consequences: - **HEALTHCHECK.** If the image declares one, it will fail while the process is suspended, and after the configured retries the container's health state becomes `unhealthy`. That is correct behaviour, but it can be alarming, and any automation that acts on health state will act. - **Restart policy.** A container that exits and is restarted under `--restart` will suspend again on every start; if you are not watching, you get a container that appears hung instead of one that appears crash-looping. - **Dependent services.** Anything waiting for this process to answer will wait as long as you do. - **Timeouts inside the app.** Once you are stopped on a breakpoint, wall-clock time keeps running. Connection pools, leases and heartbeats can expire while you read a variable, so the state you resume into is not always the state you paused. ### Never ship it `suspend=y` in a production image is a service that does not start: the container comes up, the agent waits for a debugger that will never connect, and every health check fails. It is the single most damaging way to get this wrong, and it is why the flags belong on the run command or in a debug-only build target rather than in the default command of the image you push. The same discipline covers the security problem - an image that never carries the agent by default cannot accidentally expose an unauthenticated debug port in production.

  • What is the difference between `server=y` and `suspend=y` in a JDWP agent string?
    They are independent. `server=y` decides direction: the JVM listens and the debugger connects in, which is what you want in a container with a published port. `suspend=y` decides timing: the JVM blocks before running your code until a client attaches. You can have a listening agent that does not wait (`server=y,suspend=n`) or one that waits, and the two are chosen for different reasons.
  • Your image declares a HEALTHCHECK. What happens while the container sits suspended waiting for you?
    The check keeps failing because the application has not started serving, and after the configured retries the container's health state flips to `unhealthy`. Nothing kills it on a plain engine, but anything watching health - a restart wrapper, a load balancer, an orchestrator probe - will act on that state, so for debug runs it is common to override or lengthen the check.
  • Why prefer `-e JAVA_TOOL_OPTIONS=...` over editing the Dockerfile to add the agent?
    It debugs the exact image that ships, with no rebuild and no risk of the flags leaking into a release. The JVM reads the variable and prepends the options, so a production image gains a debug agent for one container's lifetime and reverts the moment you stop passing it.

saying these in an interview costs you the question

  • Bakes suspend=y into the image that ships to production
  • Confuses server=y with suspend=y
  • Rebuilds the image every time to change debug flags
  • Expects to attach in time to a container that runs for two seconds
  • Forgets the suspended process fails its HEALTHCHECK

context