skip to content

A host with idle processors accepts no more workloads because its address range is full - why, and what fixes it?

level: seniorimportance: should knowfreq 44%

answer

  1. a third capacity dimension
  2. one pool, fixed blocks
  3. two ceilings from one number
  4. released is not free yet
  5. held set against live set

basics

~20 s

Addresses are a capacity dimension of their own. One pool is carved into fixed per-host ranges, so a host whose range is used up places nothing more, whatever its spare processor. The fix is re-planning the ranges or enlarging the pool.

solid answer

~50 s

Every workload consumes an address, and those addresses come from one pool that was carved into a fixed range per host when the fleet was planned. That range is a placement limit exactly like processor and memory, and it is the one nobody watches - so a host sits half idle and simply stops accepting work. Churn makes it worse than the arithmetic suggests: released addresses are usually held for a cool-down before reuse, so in-flight traffic does not reach an unrelated new workload, and with thousands of short-lived replicas the held set can dwarf the live set. Fixes, in ascending cost: shrink the per-host range so the pool spreads over more hosts, hand out ranges in blocks on demand instead of one fixed block per host, shorten the hold window, or enlarge the pool - which needs agreement from whoever owns the surrounding network.

code

yaml · 13 lines
yaml
addressPlan:
  poolAddresses: 65536         # one pool for the whole fleet
  perHostRange: 256            # fixed block carved for each host
  reservedPerRange: 2          # unusable inside every block

  maxHosts: 256                # poolAddresses / perHostRange
  maxWorkloadsPerHost: 254     # perHostRange - reservedPerRange

  releasedAddressHoldSeconds: 120   # kept back before reuse

# churn check on one host:
#   3 starts/second x 120s hold = 360 addresses held
#   360 > 254 -> the range is exhausted while few replicas are alive

go deeper

for a junior

Recall that a workload needs an address to start, and that those addresses come from a fixed block given to each host - so a host can run out of addresses before it runs out of anything else.

for a middle

Explain both ceilings the plan sets at once: pool divided by range gives the maximum host count, and the range itself caps workloads per host. Show the trade between them.

for a senior

Bring in churn. Released addresses are held before reuse, so on a high-turnover workload the held set can exceed the range while few replicas are alive - and name the fixes with what each one costs.

for a principal

Own the address plan as capacity planning. Decide the pool, the range size and the hold window against projected fleet growth and churn, and get the number watched before it becomes a migration under pressure.

## Addresses are capacity, and nobody dashboards them A telemetry ingester scales to thousands of short-lived replicas across a large host fleet. Placement fails on a host that has free processor and free memory. The reason is the third resource: **every workload consumes an address**, and addresses are finite in a way that is decided once, at planning time, and then forgotten. The usual arrangement is a single pool carved into a **fixed range per host**. When a host joins the fleet it is handed one block; every workload it runs takes one address out of that block; when the block is empty, that host places nothing more. The scheduler is not confused and nothing is broken - the host genuinely has no address to give. ## The arithmetic sets two ceilings at once 1. **How many hosts the fleet can have** - the pool size divided by the per-host range. 2. **How many workloads each host can run** - the per-host range, less the handful of addresses reserved inside each range. With a pool of 65,536 addresses and a range of 256 per host, that is at most **256 hosts** and at most **254 workloads on any one of them**. Both numbers are fixed by one decision made long before the incident, and they trade directly against each other: halving the range to 128 doubles the fleet ceiling to 512 hosts while cutting each host to 126 workloads. ## Churn makes the real ceiling lower than the arithmetic When a replica ends, its address is not normally handed straight to the next one. Platforms hold it for a cool-down, because traffic still in flight to the old workload would otherwise arrive at an unrelated new one - a failure mode far nastier than a placement error. The consequence is that **the addresses in use are the live ones plus the recently released ones**. Work it through on one host: a 254-address range, a 120-second hold, and a workload churning three starts per second. The held set alone is 360 addresses - more than the range holds - so the host exhausts its range while only a few dozen replicas are actually alive at any moment. Short-lived workloads at scale are the case that finds this, and the live-replica count on a dashboard says nothing about it. ## The fixes, in ascending order of pain | fix | what it buys | what it costs | |---|---|---| | shorten the hold window | usable addresses back on churning hosts | a narrower safety margin for in-flight traffic | | shrink the per-host range | more hosts fit in the same pool | fewer workloads per host, fleet-wide | | allocate ranges in blocks on demand | capacity follows actual density | more moving parts; the allocator becomes critical | | group co-located containers on one address | several containers, one address | only applies where they genuinely belong together | | enlarge the pool | removes both ceilings | needs the network's owners to agree, and a migration | | move to a larger address family | removes the scarcity | a fleet-wide change touching everything | ## How it presents, and why it is misdiagnosed - Workloads sit unplaced while capacity dashboards show plenty of processor and memory free. - Adding hosts does not help once the **pool** is the binding constraint: the new hosts get no range at all. - It is intermittent under churn - the same workload places fine a minute later, when held addresses are returned - so it gets written off as a transient. - The diagnostic evidence lives with whatever allocates addresses on the host, not in the application's own output, so the first hour is usually spent in the wrong place. The habit worth taking from this is to treat address capacity as a first-class number: watch used-against-range per host and used-against-pool for the fleet, and set the alarm well before either runs out, because every fix on the list takes longer than the incident does.

  • Why can shrinking every per-host range make things worse for some hosts?
    Because the two ceilings trade against each other. A smaller range means the same pool stretches over more hosts, which fixes a fleet that has run out of ranges - but it lowers the cap on how many workloads any single host can run. A fleet of large hosts packing hundreds of small replicas each is exactly where that cure is worse than the disease, and it is a fleet-wide change either way.
  • Thousands of short-lived replicas churn on one host. Why does its range fill faster than the live count suggests?
    Because a released address is held back before it is reused, so that traffic still in flight to the ended workload cannot reach an unrelated new one. Under high churn the held set can be several times the live set - on a 254-address range, three starts a second against a 120-second hold reserves 360 addresses. The host exhausts its range with only a few dozen replicas actually running.

A car park that assigns each floor a fixed block of numbered bays. A floor can look half empty and still turn drivers away, because every one of its bays is either occupied or still spoken for by someone who just left.

saying these in an interview costs you the question

  • Assumes capacity means only processor and memory.
  • Thinks a released address is available for reuse immediately.
  • Believes enlarging the pool is a local change on one host.
  • Says shrinking the per-host range costs nothing.
  • Thinks adding hosts always adds capacity, even with the pool exhausted.