Your zonal subnet stops accepting new instances at peak although capacity and quota are fine - why?
answer
- one zone fails, the others do not
- counted per subnet, never per network
- interfaces, not instances
- the platform holds back a few
- rolling deployment overlaps old and new
basics
~20 sThe subnet has no free addresses left. Every attached interface consumes one, including interfaces you did not create, the platform holds back a few per subnet, and a rolling deployment runs old and new instances together - so peak demand exceeds the steady-state count.
solid answer
~40 sScale-out needs an address as much as it needs capacity, and addresses are counted per subnet, not per network. Three things add up: every network interface in the subnet holds one, including those created for managed components you placed inside it; the platform reserves a handful of addresses in each subnet for its own use; and a rolling replacement holds old and new instances at the same time, so the peak demand is well above the steady-state fleet. Because the prefix is per subnet and the subnet is per zone, the symptom is one zone failing to scale while the others are fine. The fix is a new subnet out of unallocated space added to that tier's placement set, and then sizing prefixes for surge rather than steady state.
code
pseudocode · 12 linesfunction canScaleOut(subnet, newInstances, rollingDeploy):
usable = subnet.totalAddresses - subnet.platformReserved
inUse = count(interfaces attached in subnet) // not instances: interfaces
demand = newInstances
if rollingDeploy:
demand = demand + subnet.currentInstances // old and new coexist
if (usable - inUse) < demand:
return "blocked: subnet " + subnet.name + " in zone " + subnet.zone
return "ok"
// free addresses are read per subnet; a network-wide total hides thisgo deeper
Know that launching a workload consumes an address from its subnet, and that a subnet can run out even when machines and quota are available.
Explain the accounting: interfaces rather than instances, a platform reservation per subnet, and the overlap a rolling replacement creates.
Use the one-zone symptom to name the resource before measuring, then reach for the additional subnet rather than trying to widen a live prefix.
Turn it into a standard: per-tier prefixes sized for surge, free-address alarms per subnet, and reserved space so the fast remedy exists during an incident.
## The symptom and what it rules out A scale-out event fails, the platform reports that it could not place the instances, and every obvious explanation checks out: capacity is available, the account's compute quota is not reached, the workload definition is unchanged. And it is failing in **one zone**, while the same tier scales normally in the others. That asymmetry is the tell. Quotas are usually account or region scoped; capacity shortages rarely land on exactly one tier. **A per-subnet resource is exhausted, and the per-subnet resource that runs out first is addresses.** ## Where the addresses went - **Every network interface holds one.** That includes interfaces you never explicitly created: managed components placed inside your subnets each attach one, sometimes one per zone they serve. - **The platform reserves some.** Every subnet loses a handful of addresses to the platform's own use - the exact number is set by the platform - so the usable count is below the prefix's raw size. - **Deleted workloads can leave interfaces behind.** An orphaned interface is invisible in the compute inventory and still holds its address. - **Rolling replacements double-count.** During a deployment, replacement instances are running before the originals are removed, so the tier momentarily needs far more addresses than its steady-state count. - **Other tiers share the prefix.** If a subnet was cut for one tier and quietly reused by three, the first tier to scale loses. ## The diagnosis order 1. **Read free addresses per subnet**, not for the network. A network with thousands of free addresses tells you nothing about the one prefix that is full. 2. **Compare zones.** If only one zone fails, the per-subnet explanation is already the strongest. 3. **Count interfaces, not instances.** The gap between the two is where the surprise usually lives. 4. **Check the deployment in flight.** A scale-out during a rolling replacement is competing with itself for addresses. 5. **Look for orphans** left by deleted workloads before adding capacity you may not need. ## The fixes, in the order you would actually reach for them | Move | Speed | What it costs | |---|---|---| | Release orphaned interfaces | Immediate | A careful check that nothing owns them | | Pause the rolling deployment | Immediate | The deployment waits; the incident does not worsen | | Add a subnet from unallocated space and register it with the tier | Fast, if space was reserved | A change review | | Move a co-tenant tier out of the prefix | Slower | A rolling replacement for that tier | | Re-cut the tier onto a larger prefix | Slowest | A full replacement of the tier | Notice that nothing on that list widens the existing prefix - that is not on offer while it carries workloads. ## What this leaves behind as a rule - **Size prefixes for surge, not steady state.** The peak is the steady-state fleet, plus the failover share it may absorb, plus the deployment overlap - not the number of instances you usually run. - **Alarm on free addresses per subnet**, per zone, with a threshold that fires before the next deployment rather than during it. - **Count interfaces in the plan**, including the ones managed components will attach. - **Give each tier its own prefix** so one tier's growth cannot silently starve another's. - **Keep unallocated space in the network** so the fast fix stays available at three in the morning. The interview-grade version of this answer is not `we ran out of IPs`; it is the reasoning chain. The failure is scoped to a subnet, the subnet is scoped to a zone, so the one-zone symptom identifies the resource before anything is measured - and the remedy is another subnet rather than a bigger one, because a prefix carrying workloads does not grow.
- Why does the failure appear in one zone only, on a platform with zonal subnets?Because the prefix belongs to the subnet and the subnet belongs to one zone. Addresses are therefore a per-zone pool for that tier, and the zone whose prefix filled first stops placing instances while the others, drawing on their own prefixes, carry on.
- The subnet reports plenty of free addresses but placement still fails. What else would you check?Whether the tier is even allowed to place there - a subnet that is not registered with the workload's placement set is not used - and whether the shortage is really in another subnet the same deployment touches. Then look past the network at capacity and quota again.
- How would you monitor for this rather than discover it during an incident?Alarm on free addresses per subnet and per zone, with a threshold set above the largest deployment surge that tier performs, so the alert fires while there is still room to act. A network-wide free-address metric will stay green through the entire incident.
saying these in an interview costs you the question
- Reads the network's total free addresses instead of the subnet's
- Assumes only the instances they launched consume addresses
- Thinks a rolling deployment needs no more addresses than steady state
- Expects the platform to spill into a neighbouring subnet automatically
- Believes every address in the prefix is usable by workloads
- Plans to widen the full prefix rather than add a subnet