Docker reports "could not find an available, non-overlapping IPv4 address pool" on a CI host. Why, and how do you fix it?
answer
- The message names pools, not disk
- Count the networks nothing is attached to
- Cancelled jobs leave objects behind
- About thirty-one, then nothing
- Reclaim today, raise the ceiling for tomorrow
basics
~10 sThe daemon's address-pool list is empty: roughly 31 networks fit in the built-in defaults and leaked per-project networks consumed them all. Prune the unused networks, fix the teardown, then raise the ceiling with default-address-pools.
solid answer
~50 sThat message comes from IPAM at **network creation**: every candidate block in the daemon's pool list is already taken by an existing Docker network or overlaps a host route, so there is nothing left to allocate. The built-in list holds about 31 networks, which a CI host burns through quickly when each pipeline run creates a project network and cancelled or crashed jobs never tear theirs down. Confirm with `docker network ls`, which will show dozens of stale networks with no containers attached. Reclaim them with `docker network prune -f`, then fix the cause: run the teardown in a step that always executes, and give each run a distinct project name so runs cannot collide. Finally raise the ceiling in `/etc/docker/daemon.json` with a `default-address-pools` entry — a /16 base at size 24 gives 256 networks.
code
bash · 4 linesdocker network ls -q | wc -l
docker network inspect $(docker network ls -q) \
-f '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}} containers={{len .Containers}}'
docker network prune -fgo deeper
Learn to read the error literally: it says address pool, so the problem is addresses, not disk or memory. docker network ls and docker network prune are the first two commands to reach for.
Explain why the ceiling exists — a finite built-in pool list of about 31 networks — and how per-project networks accumulate when a stack is started automatically but never torn down.
Show the whole loop: diagnose by counting networks with no attachments, reclaim safely, fix the teardown so cleanup runs even when a job is cancelled, and raise the pools to match real peak concurrency.
Treat host-local networks as a leakable shared resource with a quota, like ports or disk. Anything that creates them automatically needs deterministic cleanup, a sweeper as a backstop, and provisioning sized for burst concurrency.
### What the message actually means The error is raised by the daemon's IPAM driver when it is asked to create a network and has no block to give it. It has walked its list of candidate pools, found each one either already assigned to an existing Docker network or overlapping a route on the host, and given up. Nothing is wrong with disk, memory, the driver or the container image — it is purely an address bookkeeping failure, and it always happens at network **creation** time, never while an existing network is running. On a default host the list is small: 172.17.0.0/16 through 172.31.0.0/16 plus 192.168.0.0/16 in /20s, about 31 networks, with one already spent on `docker0`. That is generous for a laptop and thin for a build machine. ### The CI shape of it A typical runner builds two things per pipeline: an nginx-fronted static bundle and an order-checkout API, each brought up as a small stack for integration tests. Every stack creates its own network. When a job finishes cleanly, teardown removes it. When a job is cancelled, times out, or the agent is killed mid-run, the containers may be reaped but the **network object survives**, because it is a separate resource with its own lifecycle and nothing owns it once the job process is gone. So the failure arrives gradually and then all at once. `docker network ls` on the sick host shows something like 34 leftover networks with names ending in `_default`, none of them referenced by a running container, and the next pipeline that tries to create the 32nd network fails before a single test runs. The distraction on such a host is usually disk. Somebody notices a 92 GB build-cache directory, runs a cleanup, frees the disk, and the pipelines still fail — because disk was never the problem. Read the error text: it names address pools, and only address pools. ### Diagnosis in three commands ``` docker network ls docker network ls -q | wc -l docker network inspect $(docker network ls -q) -f '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}} containers={{len .Containers}}' ``` The third line is the useful one: it prints every network with the subnet it holds and how many containers are attached. A long tail of networks with `containers=0` is the diagnosis, and the subnets shown will march neatly up through 172.18, 172.19, 172.20 and onward until they run out. ### Reclaiming ``` docker network prune -f ``` removes every network with nothing attached. `docker system prune` does the same as part of a wider sweep that also clears stopped containers, dangling images and build cache — fine on a CI host, dangerous to reach for reflexively on a machine with anything you care about. Neither can touch a network that still has a container attached, so a job holding a stack open blocks its own network from being reclaimed; stop the container first. ### Stopping it recurring Pruning is the fix for today. Three changes stop tomorrow: 1. **Always-run teardown.** Put the stack shutdown in a step the pipeline executes even when the job fails or is cancelled — the equivalent of a `trap` in a shell script or an `always()` cleanup stage. Include the flag that removes orphaned containers so a partially-started stack is fully cleaned. 2. **Distinct project names per run.** Deriving the project name from the build number keeps concurrent runs from sharing or stealing each other's networks, and makes leaked ones trivially identifiable by name. 3. **A scheduled sweep.** A periodic `docker network prune -f` on the runner is a cheap safety net for the leaks you did not anticipate; it can only remove networks nothing is attached to. ### Raising the ceiling Even with clean teardown, a busy runner with many concurrent jobs can want more than 31 networks at once. Configure the pools explicitly: ``` { "default-address-pools": [ { "base": "10.207.0.0/16", "size": 24 } ] } ``` `size` is the prefix length of each carved network, so this single entry supplies 256 networks of 254 usable addresses each — ample for a build host, where per-network address counts are tiny and network **count** is what runs out. Restart the daemon to apply it, and remember the pools govern only future allocations: existing networks keep the subnets they already hold. ### The sentence that closes the answer Treat networks as a finite, leakable resource on shared hosts, exactly like ports and disk. Anything that creates one automatically must remove it deterministically, and the host should be provisioned with enough address space that a burst of concurrency is not an outage.
- Roughly how many networks does a default Docker host support, and why that number?About 31. The built-in pool list is 172.17.0.0/16 through 172.31.0.0/16 — fifteen blocks — plus 192.168.0.0/16 carved into sixteen /20s. One is already spent on the default bridge. There is no other limit at play: not interfaces, not memory. Once the list is exhausted the next `docker network create` fails with the non-overlapping-pool error.
- Will `docker network prune` remove a network that still has a stopped container attached?No. Prune only removes networks with nothing attached, and a stopped container still counts as attached until it is removed. That is why a host can look full of unused networks that refuse to prune. Remove the containers first — `docker container prune` then `docker network prune` — or use `docker system prune`, which sweeps containers before networks in the same pass.
- Would enlarging the pools alone have been an acceptable fix?Only as a stopgap. Bigger pools raise the ceiling from about 31 to hundreds, which buys time, but a genuine leak still climbs until it hits the new ceiling — later and more confusingly. Fix the teardown so networks are removed deterministically, then size the pools for real peak concurrency. Do both: one addresses the leak, the other addresses legitimate load.
saying these in an interview costs you the question
- Blames disk usage or the build cache for the error
- Restarts the daemon and calls it fixed
- Thinks the limit is on containers, not networks
- Assumes prune removes networks with containers attached
- Raises the pools without fixing the leak
- Suggests reinstalling Docker to clear the state