skip to content

When would you accept more than one process in a single Docker container?

level: principalimportance: nice to knowfreq 30%

answer

  1. Start by correcting what the rule means
  2. Ask what the exception costs the platform
  3. Lifetime, failure domain, scaling, crossable boundary
  4. Grant it on conditions, not on preference
  5. PID 1 must exit non-zero when a child fails

basics

~20 s

Rarely, and on conditions: when two concerns share a lifetime and a failure domain and the coupling cannot cross a container boundary, or when consuming a vendor image with its own process manager. Then demand signal forwarding, a non-zero exit on child failure, and stdout logging.

solid answer

~40 s

First correct the question: forked workers and runtime threads are one concern, so most apparent exceptions are not exceptions. For genuine cases I apply four tests -- do they share a lifetime, do they share a failure domain, do they scale together, and is the coupling something a container boundary actually breaks? A unix socket or a shared file crosses a boundary fine via a volume; a process that must see another's process table does not. A vendor appliance image, or a CI or sandbox container, is also a legitimate exception. When I grant one I require PID 1 to forward the stop signal, finish shutdown inside the container's stop timeout, and exit non-zero as soon as any supervised program fails permanently; all output on stdout; and a health check covering every process.

go deeper

for a junior

Be ready to say the default is one concern per container and that a process forking its own workers does not break it. You are not expected to adjudicate exceptions, only to recognise the default and why it exists.

for a middle

Explain what is actually lost when two concerns share a container -- one exit code, one restart policy, one set of limits, one log stream -- so you can say concretely what an exception would have to give back.

for a senior

Show the tests you apply in review: shared lifetime, shared failure domain, shared scaling, and whether the coupling can cross a container boundary via a volume or a network. Then state the conditions you attach when you say yes.

for a principal

Own the policy rather than the case. Decide what the supported co-location pattern is, ship it centrally with correct signal and exit behaviour, and be able to articulate which platform guarantees erode when every team writes its own process manager.

### Start by restating the rule correctly The default is one *concern* per container, not one PID. A process that forks a pool of workers, a runtime that runs many threads, an application that shells out for a moment -- all one concern, all fine. The question is only ever about two things that could be operated separately but are being shipped together. ### Why the default is strong Every operational surface the engine offers is anchored on PID 1: the container lives as long as PID 1 lives, `docker stop` signals PID 1, `State.ExitCode` is PID 1's, `--restart` is evaluated when the container exits, `docker logs` replays PID 1's streams, and `--memory`/`--cpus` bound the whole cgroup as one. Co-locating two concerns spends all of that at once: you lose per-concern restart, per-concern exit status, per-concern limits, per-concern scaling and per-concern release. So the burden of proof sits with the exception. ### The tests I apply 1. **Lifetime coupling.** Must the helper die exactly when the application dies, and be meaningless without it? A helper you would ever want to restart alone is a second container. 2. **Failure coupling.** If the helper fails, do I want the whole unit replaced? If the honest answer is "no, the app should keep serving", they are two failure domains and belong apart. 3. **Scaling coupling.** If I need three of one and one of the other, they cannot share a container. 4. **Is the boundary actually crossable?** A shared unix socket or a shared file can be handed across a container boundary with a volume; a shared network endpoint with a user-defined network. If the coupling is something a container boundary genuinely breaks -- one process needing to see another's process table or address space, for instance -- that is a real reason to co-locate. 5. **Whose code is it?** A vendor appliance that ships its own process manager is a legitimate exception: you are consuming an image, not designing one. 6. **Is it even a service?** Build, test and developer-sandbox containers are not long-running services; the rule's whole justification -- operability of a production workload -- barely applies. ### What I require when the exception is granted Co-location is allowed on conditions, because the point is to restore what was lost: * **PID 1 must be honest.** It forwards the stop signal to every child promptly, its own shutdown wait fits *inside* the container's stop timeout, and it exits non-zero as soon as any supervised program fails permanently -- so the container exits, the restart policy fires, and the failure is visible. * **All output on stdout/stderr**, so `docker logs` and log collection still work. * **A health check that covers every process**, not just the one that happens to be listening. Note that an unhealthy container is not itself restarted by the engine -- something above must act on it. * **Documented shared budget.** One `--memory` and one `--cpus` cover both; write down which process is expected to lose under pressure, because otherwise the kernel decides for you. * **A written reason and an expiry.** Exceptions accumulate; each one should name why the split was impossible and what would let it be removed. ### The tradeoff at platform scale The strategic cost of allowing the exception freely is not one awkward image; it is that every team invents its own private supervision layer, and the platform's guarantees -- "a crashed workload restarts", "a failed workload pages someone", "logs are collected", "resource limits are enforced per workload" -- quietly stop holding in ways nobody can enumerate. That is why the useful position for a lead is not "never" but "rarely, and centrally": pick one supported pattern for the genuinely-coupled case, ship it as a base image or a documented recipe with the signal and exit behaviour already correct, and make everything else use two containers. The counter-pressure is real and worth naming: splitting can mean a second build, a second image to keep patched, and a second thing to deploy in lockstep. For a team whose image takes eleven minutes to build, "one more image" sounds expensive. Usually it is not -- the same source tree and the same base can produce both, or the same image can be run twice with different commands -- and the operational clarity is worth more than the build minutes. But a principal should cost it honestly rather than quoting the rule. ### The answer in one breath "Rarely. Forked workers are still one concern, so most apparent exceptions are not exceptions. I accept real co-location when the processes share a lifetime and a failure domain and the boundary genuinely cannot be crossed, or when I am consuming somebody else's appliance image. When I do, I require PID 1 to forward signals and exit non-zero on child failure, all logs on stdout, a health check covering both, and a written reason -- because everything the platform promises is anchored on PID 1."

  • A team argues they cannot split because their image already takes eleven minutes to build and a second image doubles that. How do you respond?
    Usually the premise is wrong: the same source tree and the same base can produce both images from one build, or -- more often -- the same image can simply be run twice with different commands, which costs no extra build at all. If a genuine second build is unavoidable, cost it honestly against the operational loss: no per-concern restart, exit code, limits or scaling. Build minutes are cheap next to a service that dies silently.
  • If you do allow a supervised multi-process container, what single behaviour matters most?
    That PID 1 exits non-zero as soon as any supervised program fails permanently. That one behaviour restores the chain the platform depends on: the container exits, the restart policy is evaluated, an exit code is recorded, and the failure becomes visible. Everything else -- log routing, health probes, signal timing -- is recoverable afterwards; a PID 1 that outlives its dead children is not.
  • What is the organisational cost of granting these exceptions case by case?
    Each exception is a private supervision layer with its own signal, retry and logging semantics, so the platform's guarantees -- crashed workloads restart, failures page someone, logs are collected, limits apply per workload -- stop holding in ways nobody can enumerate. The fix is to make the exception central: one supported pattern, shipped as a base image or documented recipe with the correct behaviour already built in.

saying these in an interview costs you the question

  • Says never, with no view on vendor or sandbox images
  • Counts PIDs instead of concerns when judging a design
  • Allows co-location without requiring signal forwarding
  • Accepts a PID 1 that outlives permanently failed children
  • Ignores that the two processes share one memory and CPU budget
  • Grants exceptions per team with no shared supported pattern

context