When you publish a container port on a Linux Docker host, what actually happens to a packet arriving at that host port? Describe the role of the bridge interface, the iptables NAT rules Docker installs, and the docker-proxy userland process.
answer
- nat/PREROUTING and OUTPUT -> DOCKER chain -> DNAT to container IP
- FORWARD via DOCKER-USER then Docker chains
- conntrack rewrites replies; client never sees container IP
- POSTROUTING MASQUERADE for outbound
- docker-proxy = userland relay for loopback/hairpin, disable with userland-proxy:false
basics
~20 sDocker adds an iptables DNAT rule in the nat table's DOCKER chain that rewrites the destination to the container's bridge IP and port; the FORWARD chain permits it and conntrack rewrites replies. Outbound container traffic is source-NATed by a MASQUERADE rule. A small docker-proxy process covers cases NAT misses, such as loopback connections.
solid answer
~60 sPublishing is destination NAT. When the container starts, Docker writes an iptables rule into the `DOCKER` chain of the `nat` table, jumped to from PREROUTING and OUTPUT: packets for host port 8080 get their destination rewritten to the container's bridge address, say 172.18.0.2:80. Routing then delivers them over the `docker0`/`br-*` bridge, the `FORWARD` chain (via `DOCKER-USER` and Docker's own chains) accepts them, and conntrack un-rewrites the replies so the client sees answers from the host address. The reverse direction uses source NAT: a `MASQUERADE` rule in POSTROUTING rewrites container source addresses to the host's, which is how containers reach the internet from a private subnet. `docker-proxy` is a small userland process spawned per published port. It accepts connections and copies bytes into the container, covering paths NAT does not, notably host loopback traffic and hairpin cases. It costs a process and a copy per mapping, which is why `userland-proxy: false` in the daemon config is a common tuning, leaving pure iptables. Check it with `iptables -t nat -L DOCKER -n`.
code
bash · 6 linesiptables -t nat -L DOCKER -n -v
iptables -t nat -L POSTROUTING -n -v | grep MASQUERADE
iptables -L DOCKER-USER -n -v
ss -ltnp | grep docker-proxy
conntrack -L 2>/dev/null | headgo deeper
Know that publishing works through NAT rules on the host rather than the container owning a host port.
Trace the path: DOCKER chain DNAT, forward through the bridge, conntrack for replies, MASQUERADE outbound, and name docker-proxy's role.
Show the diagnostic commands, discuss userland-proxy tradeoffs, client-IP preservation, and the fact that Docker rewrites its own chains so custom rules go in DOCKER-USER.
Reason about NAT as a fleet-level constraint: per-port process cost, source-IP loss for security controls, and when host networking or an L7 edge is the better shape.
## The packet path Start from the topology: the container has a private IP on a Linux bridge (`docker0` for the default bridge, `br-<id>` for user-defined ones). That address is not reachable from outside the host, so publishing must translate. When a container is created with `-p 8080:80`, Docker programs iptables: 1. **nat/PREROUTING** jumps to the `DOCKER` chain for packets destined to local addresses. The `DOCKER` chain holds a rule like `-p tcp --dport 8080 -j DNAT --to-destination 172.18.0.2:80`. The packet's destination is rewritten before the final routing decision. 2. **Routing** now sends the packet to the bridge, out of the host stack and into the container's namespace via the veth pair. 3. **filter/FORWARD** must accept it. Docker jumps first to `DOCKER-USER` (left empty for administrators), then to `DOCKER-ISOLATION-STAGE-1/2` and the `DOCKER` filter chain, which contains an accept rule for the published port; established/related traffic is accepted by a conntrack rule. 4. **conntrack** records the translation, so reply packets from 172.18.0.2:80 are automatically rewritten back to the host address and port before leaving. The client never sees the container IP. For traffic originating **on the host** (`curl localhost:8080`), PREROUTING is not traversed — locally generated packets go through nat/OUTPUT, which also jumps to `DOCKER`. ## Outbound: MASQUERADE A container connecting outward has a private source address that the upstream network cannot route back to. A POSTROUTING rule of the form `-s 172.18.0.0/16 ! -o br-xxx -j MASQUERADE` rewrites the source to the outgoing interface's address, and conntrack maps replies back. This is why containers have working internet access with no configuration, and why an external service sees the host's IP rather than the container's. ## What docker-proxy is for For every published port the daemon starts a `docker-proxy` process bound to that host port. It is a plain userland relay: accept a connection, dial the container, copy bytes both ways. Two reasons it exists: - **Loopback and edge cases.** Some paths are awkward for NAT alone — historically connections to 127.0.0.1 and hairpin traffic where a container reaches the host's published address for a container on the same bridge. The proxy makes them behave. - **Port reservation.** Because it actually binds the host port, a conflicting process gets a clear "address already in use" and the mapping visibly owns the port. The costs are real: one process per published port (noticeable when a host publishes hundreds), an extra userland copy on the data path, and the fact that the container sees the proxy's address rather than the true client address in some configurations, which breaks IP-based logging and rate limiting. Setting `"userland-proxy": false` in `/etc/docker/daemon.json` disables it and relies on iptables hairpin NAT; it is a well-trodden setting, with the caveat that some loopback scenarios then behave differently. ## Consequences worth knowing - **Docker owns iptables.** The daemon writes and refreshes these rules, so hand-edited rules inside Docker's chains are overwritten. `DOCKER-USER` is the chain reserved for your rules; it is evaluated before Docker's accept rules. - **NAT means the source address changes.** Applications that log or authorise by client IP need the userland proxy disabled, or a proxy that forwards the original address at the application layer. - **Debugging is concrete.** `iptables -t nat -L DOCKER -n -v` shows every publication and its packet counters, `conntrack -L` shows live translations, and `ss -ltnp` on the host reveals whether docker-proxy holds the port. - **nftables hosts** still work because Docker's rules are applied through the iptables compatibility layer, but the output you inspect may come from `nft list ruleset` instead.
- What changes if you set userland-proxy to false?Publishing then relies purely on iptables NAT, including hairpin rules, so no per-port relay process is started. You save a process and a userland copy per mapping and the container sees the real client address more consistently. The tradeoff is that some loopback and hairpin corner cases behave differently, so it is worth testing local access paths after the change.
- Why does the application inside the container often see the host's address rather than the client's?With the userland proxy in play the connection is made by docker-proxy itself, so the source is the bridge gateway. Even with pure NAT, outbound and hairpin paths are masqueraded. Anything doing IP-based logging, rate limiting or allowlisting needs either the userland proxy disabled or an application-layer forwarded-address header from a real proxy.
DNAT is a switchboard rewriting the extension on an incoming call; MASQUERADE is the office putting its own number on outgoing calls so replies come back.
saying these in an interview costs you the question
- Saying published ports work by the container binding a host port directly
- Hand-editing rules inside Docker's own chains instead of DOCKER-USER
- Thinking docker-proxy carries all published traffic on a modern daemon
- Confusing DNAT (inbound publishing) with MASQUERADE (outbound source NAT)
- Assuming client IPs survive unchanged into the container