Two containers in one co-located group both listen on port 8080 — what happens when the group starts?
answer
- one address, one port space
- two processes on one machine
- the second bind is the one that loses
- address already in use, then restart
- nothing remaps a port inside the group
basics
~20 sOne of them fails. The group has a single address and therefore a single port space, so whichever member binds 8080 first keeps it and the other fails with the port already in use, then restarts into the same failure.
solid answer
~40 sA co-located group is reachable at one address, and one address has one port space, so the two members compete for the same port exactly as two processes on one machine would. Whichever binds first wins; the other gets an address-in-use error, exits, and is restarted by the node agent straight back into the same failure — a restart loop that looks like a broken image but is really a packaging mistake. Nothing remaps it, because there is no second address inside the group to translate onto. The fix is to treat ports as a group-wide resource: give each member a distinct port, make the port configurable instead of baked into the image, or put the second workload in a group of its own.
go deeper
Remember that ports inside a co-located group behave like ports on one machine: two members cannot both take 8080, and the second one to try fails.
Explain why there is nothing to remap — one shared address means one port space, so the platform has no second address to translate the clash onto.
Read the symptom correctly: a member that starts, dies in milliseconds and keeps restarting while its neighbour is healthy points at the bind, not at the image or the workload.
Require every component you allow into a shared group to take its listening port as configuration, so composing two components never depends on their defaults happening to differ.
## One address means one port space The defining property of a co-located group is that its members share a single network address. Everything about ports follows from that one fact. A listening socket is bound to an address and a port, so when two members of the same group both try to listen on 8080, they are asking the operating system for the same socket on the same address. The second request fails, with the ordinary 'address already in use' error that any two processes on one machine would get. This is the most common concrete surprise for someone who has only ever run containers one per address. Separate workloads each get their own address, so both can happily listen on 8080 and nobody notices that the port was ever a shared resource. Inside a group, the port is shared, and the collision is immediate. ## What actually happens, in order 1. The group is placed on a host and its members are started. 2. One member binds 8080 successfully. Which one that is depends on start order among the long-running members, which is not something the model guarantees. 3. The other member's bind fails. A process that treats a failed bind as fatal — most do — logs the error and exits with a non-zero code. 4. The node agent applies the group's restart rule to the failed member and starts it again, usually with a growing backoff. 5. The restarted member binds again, fails again, and the loop continues. Its neighbour stays healthy throughout, because nothing is wrong with it. ## Why nothing remaps the clash When traffic arrives at a group from outside, a platform can map an external port onto a port the group is listening on. That mapping exists because there are two addresses involved — the outside one and the group's. Between members of the same group there is no such pair: there is one address, and translating a port onto the very same address would just be a different collision. So the platform has nothing to do here, and any fix has to happen in the declarations or in the members themselves. | What you see | What it actually is | |---|---| | One member restarting while its neighbour serves | The restarting member lost the bind | | The failure appearing intermittently | Start order among members is not guaranteed | | A published host port change not helping | The clash is inside the group, before publishing matters | | The image being blamed | The image is fine; the group's port space is over-subscribed | ## Reading the symptom The useful diagnostic signal is the pattern, not the message. One member of a group failing within milliseconds of start, repeatedly, while the other members of the same group are healthy, points at something the member could not acquire at start-up — a port, a file, a lock — rather than at a workload problem, which usually takes longer to show up and does not repeat with such precision. Reading the failed member's output from the previous instance rather than the current one is the fastest confirmation, because the error is printed before it exits. ## The fixes, and what each is worth - **Give each member a distinct port.** Simplest and the right default. The group's ports are a small address space that you allocate deliberately, the same way you would on a single machine. - **Make the port configurable.** A component whose port is baked in cannot be composed into a group with anything that shares its default. Anything you intend to co-locate should take its listening port as configuration. - **Separate the groups.** If the two components have no reason to share an address or a scratch area, they should not be in one group at all, and the clash is a signal that they were packed together for convenience. - **Bind on distinct interfaces.** Occasionally two components can be split across different addresses on the host, but inside a group there is only one, so this is not usually available and should not be the plan. One thing that is *not* a fix: retrying. A restart reproduces the same declaration on the same address, so it produces the same failure. Backoff makes the loop quieter, not shorter.
- Which of the two members fails — is that deterministic?Not reliably. Whichever member binds second gets the error, and start order among a group's long-running members is not a guarantee you should lean on, so the same declaration can fail on different members across restarts. That is why the symptom often reads as intermittent rather than as a fixed fault.
- Does changing the published host port fix the clash?No. Publishing maps an external port onto a port the group already listens on; the collision happens inside the group, on its single shared address, before any of that applies. The two members have to agree on different ports first, and only then does what you publish externally matter.
- Why does a group of one member never hit this?Because nothing else is competing for its address. The port space is still group-wide, but with one member there is only one claimant, which is why engineers who have only run single-member groups are surprised the first time they co-locate two components that share a default port.
saying these in an interview costs you the question
- Assuming each member gets its own port space
- Expecting the platform to remap the clashing port
- Blaming the image when the bind fails
- Thinking the same member always loses the race
- Believing enough restarts will clear the clash