skip to content

How do you plan Docker's container address space across a fleet of hosts on a corporate network?

level: principalimportance: should knowfreq 33%

answer

  1. What is actually scarce here
  2. Host-local and NAT'd, so reuse is fine
  3. One block, every host, same file
  4. Count of networks, not count of addresses
  5. Enable the second address family on purpose

basics

~20 s

Agree one private block the organisation does not route, express it as default-address-pools and bip in a daemon.json shipped by configuration management, and reuse it on every host: bridge addresses are host-local and NAT'd, so they need not be unique.

solid answer

~50 s

Start from the insight that decides the whole design: bridge-network addresses are host-local and source-NAT'd on the way out, so the same range can repeat on every host without conflict. Global uniqueness is only required where container addresses are actually routed on the wire. So the fleet needs one agreed block that nothing in the organisation routes — VPNs, peering links and partner networks included — not one block per host. Express it once as `default-address-pools` plus `bip` in a daemon.json managed by configuration management, so every host is addressed identically and a new machine is correct on first boot. Size for network **count**, not address count: a /16 at size 24 gives 256 networks of 254 addresses each. Then decide deliberately about IPv6 rather than drifting into it, and monitor per-host network counts so exhaustion is a warning rather than an outage.

code

json · 6 lines
json
{
  "bip": "10.206.0.1/24",
  "default-address-pools": [
    { "base": "10.207.0.0/16", "size": 24 }
  ]
}

go deeper

for a junior

Understand the starting point: container addresses are private to the host and translated on the way out, so they do not consume corporate address space unless someone deliberately routes them.

for a middle

Be able to write the daemon.json that implements a plan — a bip for the default bridge and default-address-pools entries with a base and size — and explain how many networks a given entry actually provides.

for a senior

Show the operational side: rolling the file out, restarting daemons, recreating networks that still hold old subnets, and monitoring per-host network counts so exhaustion surfaces as an alert rather than a failed deploy.

for a principal

Own the tradeoffs out loud — reuse versus uniqueness, size versus count, dual-stack cost versus need — and justify why a dozen lines of configuration shipped fleet-wide is cheap against an outage whose symptoms look like DNS or the VPN.

### The question behind the question An interviewer asking this wants to see whether you understand what is actually scarce. It is tempting to treat container addresses like datacentre addresses and carve a unique block per host. For the ordinary case that is wasted effort: containers on bridge networks are NAT'd behind the host address on the way out, and nothing outside the host ever addresses them directly. Two hosts can both use 10.207.3.0/24 and never notice each other. Uniqueness only becomes a requirement when container addresses are visible on the wire — a network driver that puts containers directly on the physical LAN, or a routed multi-host scheme — and those are deliberate choices you know you have made. So the design is not "an allocation per host". It is **one identical allocation, applied everywhere**. ### Choosing the block Two properties matter, and only two. **Nothing may route it.** This is the failure that actually happens: a host takes a range that a VPN, a peering link, an acquired company's network or a partner extranet also routes, and every destination in that range becomes unreachable from the host and its containers, because the local connected route beats the remote one. Ask the network team for a block they will keep unrouted, and confirm it against the VPN pools specifically — laptops are where this bites, because a developer creates networks with the VPN down and connects afterwards. **It must have room.** Not room for addresses — a /24 per network is far more than most stacks use — but room for network **count**. A `default-address-pools` entry of `{"base": "10.207.0.0/16", "size": 24}` yields 256 networks. Compare that with the built-in default of roughly 31 and the reason to configure it explicitly is obvious the first time a build host runs many stacks concurrently. The default 172.16.0.0/12 neighbourhood is the single worst place to leave things, because it is the range enterprises most often use themselves. ### Making it stick A correct block that lives in one engineer's notes is not a plan. Ship the configuration: ``` { "bip": "10.206.0.1/24", "default-address-pools": [ { "base": "10.207.0.0/16", "size": 24 } ] } ``` Managed by whatever provisions your hosts, so that a machine built next year is addressed like the machines built today. Two operational details belong in the runbook alongside it. First, the file is read at daemon start, so applying a change means restarting the daemon. Second, and more often missed, the pools govern only **future** allocations: hosts that already carry networks on the old ranges keep them until those networks are removed and recreated, so a migration needs a pass that does exactly that, scheduled when tearing down stacks is acceptable. Developer laptops are part of the fleet even when they are not managed like servers. They are where the VPN collision actually surfaces, so the same file, or a documented equivalent for the desktop engine's settings pane, needs to reach them. ### The IPv6 decision Enable it because you need it, not because it is available. The genuine reasons are an IPv6-only or IPv6-preferred environment, or applications that must be reachable over IPv6 end to end. The costs are real: you need either a delegated prefix or a deliberate unique-local allocation, the daemon's IPv6 firewall and NAT behaviour has changed across engine versions so the semantics differ by version, and every diagnostic habit, firewall rule and monitoring check on the team now has a second address family to cover. Dual-stack is not free, and half-configured dual-stack — addresses assigned but filtering not equivalent — is worse than not enabling it, because it can expose over one family what you carefully restricted on the other. If you enable it, pin the engine version, document the prefix in the same place as the IPv4 pools, and test the filtering path explicitly. ### Living with the decision Add two things to operations. **Monitor the network count per host** and alert well before the pool is exhausted, because exhaustion is not a gentle degradation — the next stack simply fails to start, and the error names address pools while the on-call engineer is looking at disk. **Make automatic creators clean up deterministically**: anything that creates networks per job or per project must remove them in a step that runs even on cancellation, with a periodic sweep as a backstop. The tradeoff to state out loud: this is cheap insurance. Private address space costs nothing, the file is a dozen lines, and the failure it prevents is a whole-machine outage whose cause looks like DNS, firewalls, or the VPN client, and typically consumes a day of the wrong people's time before anyone runs `ip route get` and sees a bridge interface where a tunnel should be.

  • Why can every host in the fleet safely use the same container address block?
    Because bridge-network addresses are host-local and source-NAT'd behind the host's own address on the way out. Nothing off the host ever addresses a container directly, so two hosts using the same range never see each other's. Uniqueness becomes mandatory only when container addresses are exposed on the wire — a driver that puts containers on the physical LAN, or a routed multi-host design — which is a deliberate choice, not the default.
  • How would you migrate an existing fleet onto new pools without an outage?
    Roll the daemon.json out first; it changes nothing until the daemon restarts, and even then existing networks keep their subnets. Then drain each host and restart the daemon so the default bridge picks up the new `bip`, and recreate the remaining networks during a window when tearing their stacks down is acceptable. Verify per host with `docker network inspect` rather than assuming the file implies the state.
  • What would make you enable IPv6 for containers, and what does it cost?
    Enable it when the environment is IPv6-only or IPv6-preferred, or when services must be reachable over IPv6 end to end. It costs a real or unique-local prefix, engine-version-specific firewall and NAT semantics you must verify rather than assume, and duplicated firewall rules, monitoring and diagnostic habits across two address families. Half-configured dual-stack is worse than none, because filtering can differ between the families.
  • What operational signal tells you a host is approaching pool exhaustion?
    The count of allocated networks against the capacity your pools provide. Export `docker network ls -q | wc -l` per host and alert at a fraction of the ceiling. It matters because exhaustion has no gradual phase: everything works until a creation fails outright, and the message names address pools while the responder is usually looking at disk or memory.

saying these in an interview costs you the question

  • Allocates a unique range per host with no routing need
  • Picks a replacement range without asking what is routed
  • Sizes pools for addresses rather than network count
  • Leaves the plan in a wiki instead of config management
  • Enables IPv6 without checking firewall parity
  • Forgets developer laptops, where VPN collisions actually happen

context