When you run a container with the default bridge network, what does the Linux kernel actually set up so it can reach the internet, and why can it still bind port 8080 while another container also uses 8080?
answer
- fresh net ns = loopback only, down
- veth pair: eth0 inside, veth on docker0
- MASQUERADE out, DNAT for -p
- per-namespace port space → 8080 twice is fine
- embedded DNS 127.0.0.11 on user-defined bridges
basics
~20 sThe container gets its own network namespace — a private stack with its own interfaces, routes, iptables rules and full port range. Docker creates a veth pair, puts one end inside as eth0 and attaches the other to the docker0 bridge, then NATs outbound traffic via the host.
solid answer
~50 sA fresh network namespace starts nearly empty: only a down loopback device, no routes, no addresses. Docker then: 1. Creates a **veth pair** — a virtual cable with two ends. 2. Moves one end into the container's namespace, renames it `eth0`, assigns an address from the bridge subnet and adds a default route via the bridge's gateway address. 3. Attaches the other end to the host bridge `docker0`. 4. Adds a **MASQUERADE** NAT rule so packets leaving the host get the host's source address, plus DNAT rules for any published ports. Because the port number space is per-namespace, every container has its own 0–65535 range; two containers binding 8080 never collide. A collision only appears on the host side, when two `-p 8080:...` publishes fight over the host's single port 8080. Containers on the same user-defined bridge reach each other by name via Docker's embedded DNS at 127.0.0.11 inside the namespace.
code
bash · 14 lines# By hand: namespace + veth pair + bridge attachment
ip netns add demo
ip link add veth-h type veth peer name veth-c
ip link set veth-c netns demo
ip link set veth-h master docker0 up
ip netns exec demo ip addr add 172.17.0.99/16 dev veth-c
ip netns exec demo ip link set veth-c name eth0 up
ip netns exec demo ip link set lo up
ip netns exec demo ip route add default via 172.17.0.1
# Inspect a running container's stack from the host
PID=$(docker inspect -f '{{.State.Pid}}' myapp)
nsenter -t $PID -n ip addr
nsenter -t $PID -n ss -ltnpgo deeper
Explain that each container has its own network stack and IP, and that -p maps a host port to a container port.
Describe the veth pair, docker0 bridge, default route, MASQUERADE for egress and DNAT for publishing, and the per-namespace port space.
Cover the alternative modes and their namespace meaning, embedded DNS on user-defined networks, the 127.0.0.1-bind trap, and how to inspect a container's stack with nsenter when the image has no tools.
Weigh bridge-plus-NAT against host networking and higher-performance datapaths for latency-sensitive services, and set the team default for network naming and publishing policy.
## What a network namespace contains The net namespace is the most complete of the namespaces: it virtualises an entire network stack. Inside it live its own network interfaces, IP addresses, routing tables, ARP/neighbour tables, netfilter (iptables/nftables) rules and connection-tracking table, socket table and the full TCP/UDP port number space, plus its own `/proc/net` and sysctl network knobs. A newly created network namespace contains exactly one interface — loopback — and it is administratively down. No routes, no addresses. Nothing can talk to anything, including itself, until you configure it. ## The veth pair Because a physical NIC can only live in one namespace at a time, connectivity for containers is built from **veth** (virtual Ethernet) devices. A veth is created as a *pair*: two interfaces acting as the two ends of a patch cable — a frame written into one end pops out of the other. Docker creates the pair in the host namespace, then moves one end into the container's namespace and renames it `eth0`. The peer stays on the host, visible in `ip link` as something like `veth3a1b2c@if7`, and is enslaved to the `docker0` bridge. `docker0` is a Linux software bridge (a layer-2 switch in the kernel) holding an address such as 172.17.0.1/16, which is the container's default gateway. Every container on the default bridge gets an address from that subnet and a default route pointing at .1. Containers on the same bridge therefore reach each other directly at layer 2 — no NAT involved. ## Getting out to the internet Outbound packets from 172.17.0.x arrive at the host via the bridge. The host has `net.ipv4.ip_forward=1` enabled (Docker sets it) so it routes them onward, and a **MASQUERADE** rule in the `nat` table's POSTROUTING chain rewrites the source address to the host's outbound IP, with conntrack reversing the translation for replies. Without the masquerade, replies would be sent to a private address the wider network cannot route back to. ## Getting in — published ports `-p 8080:80` does two things: it inserts a **DNAT** rule in PREROUTING (and OUTPUT for host-local traffic) that rewrites destination host:8080 to containerIP:80, and it may run a userland `docker-proxy` process for cases the NAT path does not cover, such as localhost-to-published-port on some configurations. Note the direction of the mapping — the *host* port is on the left. Publishing is required precisely because the container's port lives in a different namespace and is otherwise unreachable from outside the bridge subnet. ## Why 8080 twice is fine A socket bind is scoped to the network namespace it happens in. Ten containers can each bind 0.0.0.0:8080 because there are ten independent port spaces. Conflict is a *host* phenomenon: `-p 8080:8080` twice fails with "address already in use" because both publishes claim the host namespace's single port 8080. Map to different host ports, or skip publishing and let a reverse proxy on the same network reach them by container name. ## Name resolution On a **user-defined** bridge, Docker runs an embedded DNS resolver reachable at 127.0.0.11 inside each container's namespace (the container's `/etc/resolv.conf` points there), resolving container names and network aliases to their addresses, and forwarding everything else upstream. The legacy default bridge lacks this automatic DNS, which is the main practical reason to create a user-defined network for anything multi-container. ## Other network modes, and what they mean at the namespace level - `--network host` — **no new net namespace**. The container shares the host's stack: its binds occupy host ports directly, publishing is meaningless, and isolation of the network layer is gone. - `--network none` — a fresh namespace left with only loopback; deliberately unreachable. - `--network container:<other>` — join an existing container's namespace: both see the same interfaces and addresses, talk over 127.0.0.1, and share the port space, so they cannot both bind the same port. This is exactly the model a Kubernetes Pod uses for its containers. ## Inspecting it by hand The container's namespace can be entered with `nsenter -t <hostpid> -n`, after which `ip addr`, `ip route` and `ss -ltnp` describe the container's stack from a host shell — useful when the image lacks networking tools. `ip netns` alone will not list Docker's namespaces because Docker does not create the bind mounts under `/var/run/netns` that the tool expects; bind-mounting `/proc/<pid>/ns/net` there makes them visible. ## Common failure shapes A container that resolves DNS but cannot reach the internet usually means the masquerade rule is missing or a firewall management tool has flushed Docker's chains. A container unreachable from another container on the *default* bridge by name is normally the missing embedded DNS — move both to a user-defined network. And an application that binds `127.0.0.1` inside the container is unreachable through a published port, because that loopback belongs to the container's namespace, not the host's; it must bind 0.0.0.0.
- An app inside a container binds 127.0.0.1:8080 and you publish it with -p 8080:8080, but connections from the host are refused. Why?The loopback address inside the container belongs to that container's own network namespace, so a socket bound there only accepts traffic originating inside the container. Published-port traffic arrives via the veth interface with the container's bridge address as destination, and no socket is listening on it. The fix is to bind 0.0.0.0 inside the container and rely on the publish mapping for exposure control.
- What changes when you run a container with --network host?No new network namespace is created; the process uses the host's stack directly. Its listening sockets occupy host ports, so port conflicts with host services and other host-network containers become real, and -p has no effect. You gain a little performance by removing the bridge and NAT hop, and lose network isolation, so it is usually reserved for things that genuinely need host-level networking.
Each container is an apartment with its own internal phone system; the building's switchboard (NAT) rewrites outgoing calls to the building's number, and only extensions explicitly published get a direct outside line.
saying these in an interview costs you the question
- Saying two containers cannot both listen on 8080 without a mapping.
- Reading -p 8080:80 backwards, as container-port to host-port.
- Claiming --network host still isolates the network stack.
- Expecting container-name DNS to work on the legacy default bridge.
- Thinking the container's eth0 is a physical NIC moved into the namespace rather than one end of a veth pair.