What does running supervisord inside a Docker container cost you when the container stops or a process crashes?
answer
- Ask who PID 1 actually is now
- A second control plane the engine cannot see through
- Two stop deadlines nested inside one another
- It survives its children by design
- Container Up, exit code generic, logs elsewhere
basics
~20 sThe supervisor becomes PID 1, so Docker only ever sees the supervisor. Stop signals pass through a second, nested timeout; a crashed program leaves the container reporting Up with the supervisor's exit code, so no restart policy fires.
solid answer
~50 sA supervisor inside a container inserts a private control plane between the engine and the application. On stop, `docker stop` sends SIGTERM to the supervisor, which then signals its programs and waits its own configured time; two timeouts are now stacked inside the engine's grace period (10 seconds by default), so the SIGKILL can land while the app is still draining. On crash, the supervisor's job is to survive its children: it retries the program, gives up, and keeps running -- so PID 1 never exits, the container stays `Up`, `--restart on-failure` is never evaluated, and nothing above notices. `State.ExitCode` becomes the supervisor's, not the application's, so the exit status stops naming what died. By default program output also goes to log files inside the container rather than the container's stdout, so `docker logs` shows supervisor chatter instead of the app.
code
dockerfile · 8 linesFROM debian:bookworm-slim
RUN apt-get update \
&& apt-get install -y --no-install-recommends supervisor \
&& rm -rf /var/lib/apt/lists/*
COPY reranker /usr/local/bin/reranker
COPY warmer /usr/local/bin/warmer
COPY supervisord.conf /etc/supervisor/supervisord.conf
CMD ["supervisord", "-n", "-c", "/etc/supervisor/supervisord.conf"]go deeper
Be ready to say that whatever you launch as the container's first process becomes PID 1, and that Docker reports on that process only. Recognising a supervisor in a Dockerfile as a warning sign is enough at this level.
Explain the two concrete failures: a stop signal filtered through a second timeout inside the engine's grace period, and a crashed program that never becomes a container exit, so no restart policy fires and the exit code names the supervisor instead of the app.
Show how you would find this in production -- a container Up for days with a dead service, no exit code, silent logs -- and what you would change, including what a health probe can and cannot rescue.
Own the framing that a supervisor duplicates supervision that the platform already owns. Be ready to say what you would demand of the rare image that keeps one, and why letting each team roll its own erodes the platform's guarantees.
### The shape of the problem `supervisord`, `s6-overlay` and similar supervision suites solve a real problem *outside* containers: on a host you need something to start services, restart them when they die, and keep them running. Inside a container that job already has an owner -- the Docker daemon (and, above it, whatever schedules your containers). Putting a supervisor inside a container inserts a second, private control plane between the engine and your application, and the engine cannot see through it. Concretely, PID 1 in the container becomes the supervisor. Everything the engine reports and everything the engine acts on is therefore about the supervisor, not about your application. ### Stop: two nested deadlines `docker stop` sends SIGTERM to PID 1 and then waits a grace period -- 10 seconds by default, settable with `docker stop -t` or `docker run --stop-timeout` -- before sending SIGKILL to everything left in the container. With a supervisor as PID 1, that SIGTERM is delivered to the supervisor. A well-configured supervisor does forward a stop signal to the programs it manages and then waits its own configured time for them to leave. You now have two timeouts stacked: the engine's grace period, and the supervisor's own wait for each program. If the inner wait is as long as or longer than the outer one -- and the common defaults are uncomfortably close -- the engine's SIGKILL arrives while the application is still draining, and the graceful shutdown you carefully wrote never completes. Worse, a process the supervisor did not start, or a child the application forked into its own process group, may never be signalled at all and is simply killed at the end of the grace period. The failure is quiet. Nothing errors; the container just always takes the full grace period and dies hard, and in-flight requests are lost every deploy. ### Crash: the container stays "up" A supervisor's whole purpose is to survive the death of the things it supervises. Default configurations restart a failed program a few times and then park it in a failed state -- while the supervisor itself keeps running happily. From the engine's point of view nothing happened: PID 1 is alive, so the container is `Up`, so `--restart on-failure` is never evaluated, so no orchestrator above notices either. You get a container that is running and doing nothing. Take a Rust recommendation re-ranker packaged with a cache-warmer under one supervisor. The re-ranker panics 41 hours into the run and the supervisor gives up on it after a few retries. `docker ps` still prints `Up 6 days`; `docker inspect` still shows `"Status": "running"` with no exit code; the only signal anything is wrong is a p99 latency graph on a service that no longer answers. Had the re-ranker been PID 1 of its own container, it would have exited, the exit code would have been recorded, the restart policy would have fired, and a crash loop would have been visible. ### The exit code stops telling you what died `State.ExitCode` is PID 1's status. With a supervisor in front, the container's exit code is the supervisor's -- typically a single generic value however the application died. The distinction between "the application chose to exit", "the application was killed", and "the supervisor could not start anything at all" collapses into one number, and post-mortem tooling that keys on exit status loses its input. ### Logs By default a supervisor writes each program's output to log files *inside* the container rather than to its own stdout, so `docker logs` shows supervisor chatter and none of the application's output. This is fixable -- point each program's stdout and stderr at the container's own stdout/stderr in the supervisor config -- but it is extra configuration you must remember, and interleaving two programs' output on one stream then makes it harder to tell them apart again. ### Everything else that gets murkier Resource limits (`--memory`, `--cpus`, `--pids-limit`) apply to the whole container, so the two programs share one budget and the kernel picks which to kill under pressure. `docker stats` gives you one row for both. A `HEALTHCHECK` can be written to check both, but it is now the only mechanism that notices a dead child -- and an unhealthy container is not restarted by the engine on its own. The image grows to carry the supervisor and its runtime, and its configuration file becomes a second, undocumented deployment descriptor. ### What to do instead Make the application PID 1 of its own container, in exec form, and let the engine do the supervising: one concern per container, a restart policy per container, one stream of logs per container, one exit code per container. If two programs really must ship together, the burden is on that design to restore what the engine lost: forward signals correctly, exit non-zero as soon as any supervised program fails permanently, and route all output to stdout.
- If the supervisor does forward the stop signal to its programs, what is left to go wrong?Timing and coverage. The supervisor's own wait for a program to exit sits inside the engine's grace period, so if the inner wait is as long as the outer one the engine's SIGKILL arrives mid-drain. And a supervisor only signals the programs it started -- children the application forked into their own process group can be missed entirely and are simply killed when the grace period ends.
- How would you make `docker logs` useful again if you are stuck with a supervisor in the image?Configure each supervised program to write its stdout and stderr to the container's own stdout and stderr rather than to log files inside the container, and disable the supervisor's log rotation for those streams so nothing is buffered or truncated. It works, but you have merged two programs onto one stream and taken on configuration the engine would have given you for free.
- Can a HEALTHCHECK rescue the 'container is Up but the app is dead' case?Partly. A probe that actually exercises the application will flip the container to unhealthy, so a human or an outer system can see it. But the engine does not restart an unhealthy container by itself -- something above has to act on that state. It is a detection patch over a supervision problem, not a replacement for the app being PID 1.
The supervisor is a middle manager who reports that everything is fine because he is still at his desk. Head office only ever hears from him, so it never learns the team behind him walked out.
saying these in an interview costs you the question
- Claims a supervisor makes signal handling correct automatically
- Thinks the container exits when a supervised program dies
- Expects the app's exit code to reach docker inspect
- Assumes docker logs picks up supervised programs by default
- Says supervisord is needed because containers cannot restart processes
- Treats the in-container retries as equivalent to a restart policy