skip to content

Capacity Planning & Load Management

Making sure services have headroom before they need it, and shedding load gracefully when they don't. Interviewers pair this with system design: quantitative forecasting plus defined overload behavior is what separates senior answers.

on this pageshow

questions

15

In performance testing, what is the difference between a load test, a stress test, and a soak (endurance) test, and what does each one find that the others miss?

level: juniorimportance: must knowfreq 68%

answer

  1. three shapes, three questions
  2. expected peak versus past it
  3. some bugs need hours
  4. restart interval hides the leak

basics

~20 s

A load test applies expected peak traffic to confirm latency and error targets hold. A stress test pushes past that limit to find where and how the service breaks. A soak test runs moderate load for hours to expose slow leaks.

solid answer

~50 s

They differ in the traffic profile applied and the question being answered. A **load test** holds the traffic you actually expect — usually forecast peak — and asks whether latency, errors, and saturation stay inside the service's targets during a sustained hold. A **stress test** deliberately goes past that point to learn two things: the rate at which the service leaves its acceptable band, and the failure mode when it does — graceful shedding versus unbounded queues, cascading timeouts, or a pool that stays wedged after load drops. A **soak test** runs realistic moderate load for many hours to catch defects that are a function of elapsed time rather than rate: heap and file-descriptor leaks, unbounded caches, disks filling with logs. The practical rule is that the soak must outlast whatever normally restarts the process, or the restart masks the leak in the test exactly as it does in production.

go deeper

for a junior

Be able to name all three and say what each is for in one sentence: expected peak, past the limit, and over a long duration. Interviewers use this as a quick check that you have run a performance test at all.

for a middle

Explain what makes each run valid — discarding warm-up and judging the plateau, using a production-like request mix, and matching soak duration to the restart interval. Be ready to say which shape you would pick for a specific upcoming risk.

for a senior

Show that the stress test's real output is the failure mode and the recovery behaviour, not just the breaking number. Talk about what you observed when a service could not recover after the load dropped, and what you changed as a result.

for a principal

Own the question of which shapes are mandatory for which class of service and who pays for the environment and the time. Argue where a standing soak in a pre-production tier earns its cost versus where production canaries and real-traffic replay give a better answer more cheaply.

## Why the distinction matters "Performance testing" is not one activity. The three classic shapes apply different traffic profiles and answer different questions. Picking the wrong shape gets you a green result that says nothing about the risk you were actually worried about — which is worse than no test, because it buys false confidence before a launch. ## Load test: does it meet its targets at expected peak? A load test applies the traffic you expect — the peak of a normal day, or a forecast peak for a known event — and holds it there. Pass criteria come from the service's own targets: p99 latency under the threshold, error rate under the threshold, and no saturation signal trending upward across the hold. Two details decide whether the number means anything. First, **hold long enough to reach steady state**. Runtime warm-up, cache fill, connection pools growing to their working size, and autoscaling all move the numbers for the first several minutes; discard the warm-up and judge the plateau, not the average over the whole run. Second, **the request mix must resemble production**. A test that hammers one cheap cached endpoint measures that endpoint, not the service — real traffic has a mix of expensive writes, cold reads, and pathological large payloads, and the expensive tail is usually what saturates first. ## Stress test: where and how does it break? A stress test goes deliberately past expected peak until something gives. You are not measuring compliance, you are collecting two facts. The first is the breaking point: the offered rate at which latency or errors leave the acceptable band. The second, and more valuable, is the **failure mode**. Does the service shed the excess and stay up with elevated errors, or does it collapse — queues growing without bound, request timeouts cascading into callers that retry, out-of-memory kills, a connection pool wedged so hard it does not recover when load falls? A service that degrades predictably at 3x and recovers on its own is in far better shape than one with a slightly higher ceiling that needs a restart afterwards. That last point is why a stress test should always include the ramp back **down**. Return to normal load at the end and watch whether the service returns to baseline latency unaided. Many systems have a hysteresis problem — they cannot dig themselves out of a queue backlog even once the offered rate is survivable — and that is a much more dangerous property than the ceiling itself. ## Soak test: what only breaks with time? A soak or endurance test runs moderate, realistic load for hours or days. It finds the class of defect that is a function of elapsed time or cumulative request count rather than of rate, and is therefore invisible in a fifteen-minute run: - heap that grows slowly until an out-of-memory kill - sockets or file descriptors never closed, hitting a per-process limit - an in-memory cache with no bound or eviction policy - logs or temp files filling a disk - a metric whose label cardinality grows with unique inputs - credential, token, or connection expiry paths that only execute after hours - memory fragmentation and allocator drift The useful rule: **a soak must outlast the natural cleanup interval.** If the platform restarts instances every six hours, a two-hour soak proves nothing about a leak that takes a day — and it should also prompt the question of whether production is quietly relying on those restarts to stay alive. ## Related shapes worth naming A **spike test** jumps to high load instantly rather than ramping. It finds cold caches, autoscaler reaction lag, connection-pool warm-up cost, and thundering-herd behaviour that a smooth ramp completely hides. A **capacity or ramp test** steps load upward specifically to locate the point where latency turns, which is a different exercise from confirming a target at a fixed rate. ## Choosing, before a specific event The choice has a real cost, so tie it to the risk you are carrying: - A launch with a fixed date and a forecast multiplier wants a load test at forecast peak, plus a stress test so you know how much slack exists if the forecast is wrong. - A service that just added a cache or a new client library wants a soak, because that is where leaks live. - A service fronting a scheduled on-sale wants a spike test, because the risk is arrival *shape*, not total volume. ## Common mistakes Reporting a stress-test breaking point as "capacity" — the breaking point is where it fails, not where it is safe to run. Running against stubbed dependencies that reply in a millisecond, which moves the bottleneck entirely. Not resetting state between runs, so run two inherits run one's warmed caches and looks faster. And treating zero errors at peak as proof of headroom when you never measured how far past peak the cliff sits.

  • How long should a soak test actually run to be worth doing?
    Long enough to outlast whatever resets state in production. If instances are redeployed or restarted every six hours, a soak shorter than that cannot reveal a leak with a one-day time constant — and it hints that production is relying on restarts. Also cover at least one full traffic cycle, so scheduled jobs and daily batch work run inside the window.
  • Where does a spike test fit alongside these three?
    A spike test jumps straight to high load instead of ramping, so it isolates arrival shape rather than volume. It exposes cold caches, autoscaler reaction lag, connection-pool warm-up, and thundering-herd effects that a gradual ramp lets the system absorb. Run it whenever real traffic arrives instantly — a scheduled on-sale, a push notification, or failover of another region's traffic onto yours.
  • Why should a stress test always ramp back down at the end?
    To test recovery, which is a separate property from the ceiling. Some services never dig out of a queue backlog even once the offered rate is survivable again, so they need a restart to recover. Knowing that in a test is the difference between a five-minute incident and an hour-long one.

saying these in an interview costs you the question

  • Calls any traffic generation a load test regardless of goal
  • Reports the stress-test breaking point as the service's capacity
  • Believes a fifteen-minute run can reveal a memory leak
  • Treats zero errors at peak as proof of headroom
  • Stops the stress test at the peak, never testing recovery

context

open as a page

Your capacity plan has to cover both steady user growth and a marketing launch that goes live next month. How does forecasting organic growth differ from forecasting launch-driven demand, and how do you provision for each?

level: middleimportance: must knowfreq 62%

basics

~20 s

Organic growth is extrapolated from your own traffic history and arrives smoothly. Launch-driven demand has no history to extrapolate, so its size and timing must come from business inputs and be provisioned before the event, not corrected after it.

open as a page

During a ramp load test you plot achieved throughput and p99 latency against the offered request rate. How do you identify the knee of the latency curve, and what number do you take away as the service's usable limit?

level: middleimportance: must knowfreq 55%

basics

~20 s

The knee is where achieved throughput stops tracking the offered rate and p99 latency starts climbing super-linearly — queues are no longer draining. The usable limit is the highest sustained rate that still meets the latency target, which sits below the knee, not at peak throughput.

open as a page

Your retail service hits its yearly peak on Black Friday, eight weeks away, and the forecast is roughly 4x your normal daily peak. Which parts of that capacity have lead times you cannot compress, and how would you sequence the eight weeks?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Compute is usually the easy part. The long poles are cloud quota increases and capacity reservations for specific instance types and zones, stateful work like resharding or index builds, and third-party rate limits. Start those first and leave the final weeks for verification and a freeze.

open as a page

Your service is over capacity and must reject some fraction of its traffic. Why is shedding a random 10% of requests worse than shedding a deliberately chosen 10%, and what does a system need in place to be able to choose?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Random shedding drops checkout traffic and background prefetches at the same rate, so it damages revenue-bearing work to save cheap work. Choosing requires a criticality label attached at the entry point and propagated to every downstream call, with each server shedding lowest tier first.

open as a page

Requests hitting a backend service triple within a minute while its success rate falls, and after you restart it, it saturates again within seconds. How do you confirm this is a retry storm rather than a genuine traffic surge, and what do you do to break it?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Retry storms show load rising as success falls, with internal call volume spiking while edge traffic stays flat — organic surges raise both together. Breaking one requires cutting offered load below the degraded service's reduced capacity, then restoring in steps.

open as a page

Capacity plans usually size a fleet so peak load lands well below 100% of its capacity — often somewhere near 60%. Why leave that much unused, and what do you gain and lose by raising the target?

level: juniorimportance: should knowfreq 52%

basics

~20 s

The gap absorbs what the plan cannot: a lost failure domain, forecast error, the minutes it takes to add capacity, and the extra load of a rollout or a retry surge. Raising the target cuts spend proportionally and removes that cushion.

open as a page

To make sure no request is ever rejected, a team makes their service's own inbound request queue unbounded. During the next traffic spike, what actually happens to latency, memory and success rate, and what queue design would you use instead?

level: middleimportance: should knowfreq 45%

basics

~20 s

An unbounded inbound queue converts an overload into unbounded latency plus a memory incident: the queue grows without limit, every request waits longer than its client's timeout, and the process eventually dies. A bounded, deadline-aware queue rejects immediately instead.

open as a page

A service sized for about 10,000 requests per second is suddenly offered 30,000. Instead of serving roughly 10,000 of them successfully and failing the rest, nearly every request now times out. Explain what is happening inside the server, and what changes if it sheds the excess load instead.

level: middleimportance: should knowfreq 58%

basics

~20 s

Accepting more work than it can finish makes a server spend its capacity on requests whose callers have already timed out, so goodput — useful completed work — collapses toward zero. Shedding excess load early and cheaply keeps the fraction it does serve healthy.

open as a page

In capacity planning, what do N+1 and N+2 redundancy mean, and what peak utilization ceiling does N+1 impose on a service spread evenly across three availability zones?

level: seniorimportance: should knowfreq 50%

basics

~20 s

N is the capacity that serves peak load; N+1 adds one spare unit and N+2 adds two. Across three equal zones, N+1 caps peak utilization near 67%, so the two surviving zones can carry the entire peak.

open as a page

Your team validates capacity in a staging environment that is a one-tenth-scale copy of production. What makes the resulting throughput number untrustworthy, and what would you do to get a number you can actually act on?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Scaled-down environments do not scale linearly: smaller datasets inflate cache hit rates, shared tiers do not shrink with the app tier, stubbed dependencies answer far too fast, and synthetic traffic lacks real skew. Fix it by testing the shared tier separately, replaying real traffic, or measuring capacity in production under controlled conditions.

open as a page

In a load test, your service shows CPU around 45% across the fleet, yet p99 latency has tripled and total throughput has stopped rising. Which saturation signals do you look at, and what would each one tell you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Utilisation is not saturation. Look for evidence of waiting rather than of being busy: request queue depth and time-in-queue, thread-pool and connection-pool wait, downstream call latency, lock contention, and per-instance skew. Each names a different bottleneck, and each implies a different fix.

open as a page

A full capacity ramp test takes hours, so it cannot run on every commit. How would you catch a capacity regression — a change that quietly cuts maximum sustainable throughput by 20% — before it reaches production?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Stop measuring the maximum. Run a short test at a fixed request rate well below the limit and compare resource cost per request and latency against a stored baseline. Cost per request drifts detectably long before the ceiling moves, and it needs only minutes of stable load.

open as a page

A platform lead argues that because every service now scales automatically with load, the company can drop its capacity planning process entirely. You are accountable for both reliability and infrastructure spend — how do you answer?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Automatic scaling allocates capacity that already exists; it does not create it. Account and region quotas, physical zone capacity, stateful tiers and vendor limits all set ceilings, scale-up takes minutes, and buying multi-year commitments requires a demand forecast in the first place.

open as a page

One of your three availability zones fails and the surviving capacity can serve only about 70% of peak traffic. Would you let all users experience a slow, partly broken service, or deliberately serve 70% of them fully and reject the rest? How would you decide, and who has to agree to that beforehand?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Deliberate shedding usually wins: partially served users retry and consume more capacity than rejected ones, so uniform degradation costs more than it saves. Shed by session rather than per request, rank traffic by pre-agreed business criticality, and get that ladder signed off before the incident.

open as a page