skip to content

In capacity planning, what do N+1 and N+2 redundancy mean, and what peak utilization ceiling does N+1 impose on a service spread evenly across three availability zones?

level: seniorimportance: should knowfreq 50%

answer

  1. spare capacity, not spare hardware
  2. how many units can you lose at peak?
  3. (K - s) / K is the ceiling
  4. two zones means half your fleet idles
  5. more, smaller units make redundancy cheaper

basics

~20 s

N is the capacity that serves peak load; N+1 adds one spare unit and N+2 adds two. Across three equal zones, N+1 caps peak utilization near 67%, so the two surviving zones can carry the entire peak.

solid answer

~50 s

N is the number of units — zones, cells, machines, whatever failure domain you are insuring against — needed to serve peak load. N+1 means you can lose one and still serve peak; N+2 means you can lose two. The consequence people miss is that this sets a utilization ceiling, not just a machine count. With K equal zones, surviving the loss of one means each zone must be sized for `peak / (K - 1)`, so at normal peak each zone runs at `(K - 1) / K` of its capacity. Three zones gives about 67%; two zones gives 50%, meaning each zone idles half the time. N+2 across three zones is 33%, which is why teams add zones or cells instead — five zones puts N+2 back at 60%. And the redundancy is only as good as the tightest dependency: stateless headroom means nothing if the database primary sits in one zone.

code

python · 13 lines
python
def peak_utilization_ceiling(units, spares):
    """Highest utilization each unit may reach at peak while still
    serving full peak load after losing `spares` units."""
    surviving = units - spares
    return surviving / units if surviving > 0 else 0.0

for k in (2, 3, 4, 5, 10):
    print(k,
          f"N+1 {peak_utilization_ceiling(k, 1):.0%}",
          f"N+2 {peak_utilization_ceiling(k, 2):.0%}")
# 2  N+1 50%  N+2 0%
# 3  N+1 67%  N+2 33%
# 5  N+1 80%  N+2 60%

go deeper

for a junior

Know that N is the capacity needed to serve peak and that the +1 is spare capacity for losing a unit, not a spare machine sitting in a cupboard.

for a middle

Be able to derive the ceiling: with K equal units, surviving one loss means each is sized for peak/(K-1), so peak utilization cannot exceed (K-1)/K — 67% at three zones, 50% at two.

for a senior

Demonstrate that you have hit the traps: uneven per-zone load, cold caches and connection limits after failover, per-zone quota, and the dependency whose redundancy is worse than the tier you sized.

for a principal

Own the shape decision — how many failure domains to run at all — because more, smaller cells buy the same tolerance at much higher utilization, and justify N+1 versus N+2 against measured failure rates and the cost of a shortfall.

## What N actually counts N is not a server count. It is the amount of capacity required to serve peak load, measured in whatever unit you are willing to lose all at once. That unit is the **failure domain**: a machine, a rack, a cell, an availability zone, a region. Choosing it is the first decision, because N+1 across machines and N+1 across zones are wildly different amounts of money and protect against wildly different events. - **N** — serves peak with nothing to spare. Any loss is a shortfall. - **N+1** — survives the loss of one unit at peak. - **N+2** — survives the loss of two: typically one unit intentionally out (maintenance, a rollout, a drain) *and* one lost unexpectedly during that window. That second point is the whole argument for N+2. If you plan N+1 and then take a zone out for a scheduled upgrade, you are running at N for the duration — and upgrade windows are exactly when unexpected things happen. ## The utilization ceiling, derived Spread load evenly across K equal units, each of capacity C. Total capacity is K x C. To serve peak load P with one unit gone: ``` (K - 1) * C >= P => C >= P / (K - 1) ``` At normal peak each unit carries P / K, so its utilization is ``` (P / K) / (P / (K - 1)) = (K - 1) / K ``` For N+s spares generally the ceiling is `(K - s) / K`: | Units K | N+1 ceiling | N+2 ceiling | |---|---|---| | 2 | 50% | not possible | | 3 | 67% | 33% | | 4 | 75% | 50% | | 5 | 80% | 60% | | 10 | 90% | 80% | This table is the single most useful thing in the topic. It says redundancy gets cheaper as the units get smaller and more numerous, and it says a two-zone deployment is inherently expensive insurance — you buy twice your peak capacity to be N+1. It also explains why large operators slice a region into many cells rather than three big ones: the same failure tolerance at far higher utilization. ## The ceiling is usually the binding constraint In practice, redundancy sets the utilization target more often than performance does. If a team says "we run at 60% because latency degrades above that," and they also run three zones, the redundancy math was already going to cap them near 67%. The number you should carry is the *minimum* of the redundancy ceiling and whatever operational ceiling you have measured for the service. ## Traps that make the headroom fake **Aggregate capacity is not usable capacity.** A fleet at 60% overall can have one zone at 95% if traffic is not evenly distributed — sticky sessions, uneven shard placement, a client library that prefers the nearest zone. Measure utilization per failure domain, never only in aggregate. **Capacity is not fungible.** Replicas in the surviving zones can only take the load if the data is there, the caches are warm, the connection pools are large enough, and quota exists in that zone. Cold caches after a failover can multiply database load exactly when a third of your fleet just vanished. **The tightest dependency defines your real redundancy.** A stateless tier at N+2 in front of a single-zone database primary is N+2 for nothing that matters. Redundancy is a per-dependency property, and the service inherits the worst one. **Untested headroom is a hypothesis.** The way you find out whether the surviving zones absorb the load inside your latency targets is to take a zone out on purpose and watch — deliberately, in daylight, with a way to put it back. Per-zone quota limits, connection ceilings and cold caches all show up there and nowhere else. ## Deciding between N+1 and N+2 It is a cost decision with a stated risk. Ask: how often does a unit go out for planned work, how long is it out, what is the historical rate of unplanned loss, and what does a shortfall cost during the overlap? For a service with monthly maintenance windows and a serious cost of downtime, N+2 is easy to justify. For a batch tier that can be late, N+1 or even N is defensible. What is not defensible is asserting N+2 as a universal standard without knowing what it costs, or claiming N+1 while running three zones at 90%.

  • Why do many teams provision to N+2 rather than N+1?
    Because one unit is often intentionally out — maintenance, an upgrade, a drain, a rollout. During that window an N+1 plan is running at N, with zero margin for an unplanned loss, and upgrade windows are precisely when unplanned losses happen. N+2 buys tolerance for one planned plus one unplanned removal at the same time.
  • Your stateless tier is N+2 across zones but the database primary lives in one zone. What is your real redundancy?
    Whatever the database's is. Redundancy is a per-dependency property and the service inherits the weakest one; extra application replicas cannot serve requests whose writes have nowhere to go. Before buying more stateless headroom, fix the single-zone dependency or state explicitly that the plan protects against a different failure than the one you are exposed to.
  • How do you find out whether the headroom is real rather than arithmetic?
    Remove a zone deliberately and measure whether the survivors carry peak inside the latency target. That is where per-zone quota caps, connection-pool ceilings, uneven traffic distribution and cold caches surface — none of them appear in a spreadsheet, and all of them turn a calculated N+1 into an actual shortfall.

saying these in an interview costs you the question

  • Thinking N+1 means keeping one spare server somewhere
  • Running three zones near 90% and still calling it N+1
  • Reading aggregate fleet utilization instead of per-zone utilization
  • Assuming survivors absorb the load without ever measuring it
  • Adding stateless headroom while the database stays single-zone

context