skip to content

Docker reports "could not find an available, non-overlapping IPv4 address pool" on a CI host. Why, and how do you fix it?

level: seniorimportance: should knowfreq 40%

answer

  1. The message names pools, not disk
  2. Count the networks nothing is attached to
  3. Cancelled jobs leave objects behind
  4. About thirty-one, then nothing
  5. Reclaim today, raise the ceiling for tomorrow

basics

~10 s

The daemon's address-pool list is empty: roughly 31 networks fit in the built-in defaults and leaked per-project networks consumed them all. Prune the unused networks, fix the teardown, then raise the ceiling with default-address-pools.

solid answer

~50 s

That message comes from IPAM at **network creation**: every candidate block in the daemon's pool list is already taken by an existing Docker network or overlaps a host route, so there is nothing left to allocate. The built-in list holds about 31 networks, which a CI host burns through quickly when each pipeline run creates a project network and cancelled or crashed jobs never tear theirs down. Confirm with `docker network ls`, which will show dozens of stale networks with no containers attached. Reclaim them with `docker network prune -f`, then fix the cause: run the teardown in a step that always executes, and give each run a distinct project name so runs cannot collide. Finally raise the ceiling in `/etc/docker/daemon.json` with a `default-address-pools` entry — a /16 base at size 24 gives 256 networks.

code

bash · 4 lines
bash
docker network ls -q | wc -l
docker network inspect $(docker network ls -q) \
  -f '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}} containers={{len .Containers}}'
docker network prune -f

go deeper

for a junior

Learn to read the error literally: it says address pool, so the problem is addresses, not disk or memory. docker network ls and docker network prune are the first two commands to reach for.

for a middle

Explain why the ceiling exists — a finite built-in pool list of about 31 networks — and how per-project networks accumulate when a stack is started automatically but never torn down.

for a senior

Show the whole loop: diagnose by counting networks with no attachments, reclaim safely, fix the teardown so cleanup runs even when a job is cancelled, and raise the pools to match real peak concurrency.

for a principal

Treat host-local networks as a leakable shared resource with a quota, like ports or disk. Anything that creates them automatically needs deterministic cleanup, a sweeper as a backstop, and provisioning sized for burst concurrency.

### What the message actually means The error is raised by the daemon's IPAM driver when it is asked to create a network and has no block to give it. It has walked its list of candidate pools, found each one either already assigned to an existing Docker network or overlapping a route on the host, and given up. Nothing is wrong with disk, memory, the driver or the container image — it is purely an address bookkeeping failure, and it always happens at network **creation** time, never while an existing network is running. On a default host the list is small: 172.17.0.0/16 through 172.31.0.0/16 plus 192.168.0.0/16 in /20s, about 31 networks, with one already spent on `docker0`. That is generous for a laptop and thin for a build machine. ### The CI shape of it A typical runner builds two things per pipeline: an nginx-fronted static bundle and an order-checkout API, each brought up as a small stack for integration tests. Every stack creates its own network. When a job finishes cleanly, teardown removes it. When a job is cancelled, times out, or the agent is killed mid-run, the containers may be reaped but the **network object survives**, because it is a separate resource with its own lifecycle and nothing owns it once the job process is gone. So the failure arrives gradually and then all at once. `docker network ls` on the sick host shows something like 34 leftover networks with names ending in `_default`, none of them referenced by a running container, and the next pipeline that tries to create the 32nd network fails before a single test runs. The distraction on such a host is usually disk. Somebody notices a 92 GB build-cache directory, runs a cleanup, frees the disk, and the pipelines still fail — because disk was never the problem. Read the error text: it names address pools, and only address pools. ### Diagnosis in three commands ``` docker network ls docker network ls -q | wc -l docker network inspect $(docker network ls -q) -f '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}} containers={{len .Containers}}' ``` The third line is the useful one: it prints every network with the subnet it holds and how many containers are attached. A long tail of networks with `containers=0` is the diagnosis, and the subnets shown will march neatly up through 172.18, 172.19, 172.20 and onward until they run out. ### Reclaiming ``` docker network prune -f ``` removes every network with nothing attached. `docker system prune` does the same as part of a wider sweep that also clears stopped containers, dangling images and build cache — fine on a CI host, dangerous to reach for reflexively on a machine with anything you care about. Neither can touch a network that still has a container attached, so a job holding a stack open blocks its own network from being reclaimed; stop the container first. ### Stopping it recurring Pruning is the fix for today. Three changes stop tomorrow: 1. **Always-run teardown.** Put the stack shutdown in a step the pipeline executes even when the job fails or is cancelled — the equivalent of a `trap` in a shell script or an `always()` cleanup stage. Include the flag that removes orphaned containers so a partially-started stack is fully cleaned. 2. **Distinct project names per run.** Deriving the project name from the build number keeps concurrent runs from sharing or stealing each other's networks, and makes leaked ones trivially identifiable by name. 3. **A scheduled sweep.** A periodic `docker network prune -f` on the runner is a cheap safety net for the leaks you did not anticipate; it can only remove networks nothing is attached to. ### Raising the ceiling Even with clean teardown, a busy runner with many concurrent jobs can want more than 31 networks at once. Configure the pools explicitly: ``` { "default-address-pools": [ { "base": "10.207.0.0/16", "size": 24 } ] } ``` `size` is the prefix length of each carved network, so this single entry supplies 256 networks of 254 usable addresses each — ample for a build host, where per-network address counts are tiny and network **count** is what runs out. Restart the daemon to apply it, and remember the pools govern only future allocations: existing networks keep the subnets they already hold. ### The sentence that closes the answer Treat networks as a finite, leakable resource on shared hosts, exactly like ports and disk. Anything that creates one automatically must remove it deterministically, and the host should be provisioned with enough address space that a burst of concurrency is not an outage.

  • Roughly how many networks does a default Docker host support, and why that number?
    About 31. The built-in pool list is 172.17.0.0/16 through 172.31.0.0/16 — fifteen blocks — plus 192.168.0.0/16 carved into sixteen /20s. One is already spent on the default bridge. There is no other limit at play: not interfaces, not memory. Once the list is exhausted the next `docker network create` fails with the non-overlapping-pool error.
  • Will `docker network prune` remove a network that still has a stopped container attached?
    No. Prune only removes networks with nothing attached, and a stopped container still counts as attached until it is removed. That is why a host can look full of unused networks that refuse to prune. Remove the containers first — `docker container prune` then `docker network prune` — or use `docker system prune`, which sweeps containers before networks in the same pass.
  • Would enlarging the pools alone have been an acceptable fix?
    Only as a stopgap. Bigger pools raise the ceiling from about 31 to hundreds, which buys time, but a genuine leak still climbs until it hits the new ceiling — later and more confusingly. Fix the teardown so networks are removed deterministically, then size the pools for real peak concurrency. Do both: one addresses the leak, the other addresses legitimate load.

saying these in an interview costs you the question

  • Blames disk usage or the build cache for the error
  • Restarts the daemon and calls it fixed
  • Thinks the limit is on containers, not networks
  • Assumes prune removes networks with containers attached
  • Raises the pools without fixing the leak
  • Suggests reinstalling Docker to clear the state

context