skip to content

Load Testing & Saturation Signals

Finding a system's real limits before production traffic does. Interviewers ask how you would validate that a service handles 10x — load-test types and reading saturation signals are the substance of that answer.

on this pageshow

questions

5

In performance testing, what is the difference between a load test, a stress test, and a soak (endurance) test, and what does each one find that the others miss?

level: juniorimportance: must knowfreq 68%

answer

  1. three shapes, three questions
  2. expected peak versus past it
  3. some bugs need hours
  4. restart interval hides the leak

basics

~20 s

A load test applies expected peak traffic to confirm latency and error targets hold. A stress test pushes past that limit to find where and how the service breaks. A soak test runs moderate load for hours to expose slow leaks.

solid answer

~50 s

They differ in the traffic profile applied and the question being answered. A **load test** holds the traffic you actually expect — usually forecast peak — and asks whether latency, errors, and saturation stay inside the service's targets during a sustained hold. A **stress test** deliberately goes past that point to learn two things: the rate at which the service leaves its acceptable band, and the failure mode when it does — graceful shedding versus unbounded queues, cascading timeouts, or a pool that stays wedged after load drops. A **soak test** runs realistic moderate load for many hours to catch defects that are a function of elapsed time rather than rate: heap and file-descriptor leaks, unbounded caches, disks filling with logs. The practical rule is that the soak must outlast whatever normally restarts the process, or the restart masks the leak in the test exactly as it does in production.

go deeper

for a junior

Be able to name all three and say what each is for in one sentence: expected peak, past the limit, and over a long duration. Interviewers use this as a quick check that you have run a performance test at all.

for a middle

Explain what makes each run valid — discarding warm-up and judging the plateau, using a production-like request mix, and matching soak duration to the restart interval. Be ready to say which shape you would pick for a specific upcoming risk.

for a senior

Show that the stress test's real output is the failure mode and the recovery behaviour, not just the breaking number. Talk about what you observed when a service could not recover after the load dropped, and what you changed as a result.

for a principal

Own the question of which shapes are mandatory for which class of service and who pays for the environment and the time. Argue where a standing soak in a pre-production tier earns its cost versus where production canaries and real-traffic replay give a better answer more cheaply.

## Why the distinction matters "Performance testing" is not one activity. The three classic shapes apply different traffic profiles and answer different questions. Picking the wrong shape gets you a green result that says nothing about the risk you were actually worried about — which is worse than no test, because it buys false confidence before a launch. ## Load test: does it meet its targets at expected peak? A load test applies the traffic you expect — the peak of a normal day, or a forecast peak for a known event — and holds it there. Pass criteria come from the service's own targets: p99 latency under the threshold, error rate under the threshold, and no saturation signal trending upward across the hold. Two details decide whether the number means anything. First, **hold long enough to reach steady state**. Runtime warm-up, cache fill, connection pools growing to their working size, and autoscaling all move the numbers for the first several minutes; discard the warm-up and judge the plateau, not the average over the whole run. Second, **the request mix must resemble production**. A test that hammers one cheap cached endpoint measures that endpoint, not the service — real traffic has a mix of expensive writes, cold reads, and pathological large payloads, and the expensive tail is usually what saturates first. ## Stress test: where and how does it break? A stress test goes deliberately past expected peak until something gives. You are not measuring compliance, you are collecting two facts. The first is the breaking point: the offered rate at which latency or errors leave the acceptable band. The second, and more valuable, is the **failure mode**. Does the service shed the excess and stay up with elevated errors, or does it collapse — queues growing without bound, request timeouts cascading into callers that retry, out-of-memory kills, a connection pool wedged so hard it does not recover when load falls? A service that degrades predictably at 3x and recovers on its own is in far better shape than one with a slightly higher ceiling that needs a restart afterwards. That last point is why a stress test should always include the ramp back **down**. Return to normal load at the end and watch whether the service returns to baseline latency unaided. Many systems have a hysteresis problem — they cannot dig themselves out of a queue backlog even once the offered rate is survivable — and that is a much more dangerous property than the ceiling itself. ## Soak test: what only breaks with time? A soak or endurance test runs moderate, realistic load for hours or days. It finds the class of defect that is a function of elapsed time or cumulative request count rather than of rate, and is therefore invisible in a fifteen-minute run: - heap that grows slowly until an out-of-memory kill - sockets or file descriptors never closed, hitting a per-process limit - an in-memory cache with no bound or eviction policy - logs or temp files filling a disk - a metric whose label cardinality grows with unique inputs - credential, token, or connection expiry paths that only execute after hours - memory fragmentation and allocator drift The useful rule: **a soak must outlast the natural cleanup interval.** If the platform restarts instances every six hours, a two-hour soak proves nothing about a leak that takes a day — and it should also prompt the question of whether production is quietly relying on those restarts to stay alive. ## Related shapes worth naming A **spike test** jumps to high load instantly rather than ramping. It finds cold caches, autoscaler reaction lag, connection-pool warm-up cost, and thundering-herd behaviour that a smooth ramp completely hides. A **capacity or ramp test** steps load upward specifically to locate the point where latency turns, which is a different exercise from confirming a target at a fixed rate. ## Choosing, before a specific event The choice has a real cost, so tie it to the risk you are carrying: - A launch with a fixed date and a forecast multiplier wants a load test at forecast peak, plus a stress test so you know how much slack exists if the forecast is wrong. - A service that just added a cache or a new client library wants a soak, because that is where leaks live. - A service fronting a scheduled on-sale wants a spike test, because the risk is arrival *shape*, not total volume. ## Common mistakes Reporting a stress-test breaking point as "capacity" — the breaking point is where it fails, not where it is safe to run. Running against stubbed dependencies that reply in a millisecond, which moves the bottleneck entirely. Not resetting state between runs, so run two inherits run one's warmed caches and looks faster. And treating zero errors at peak as proof of headroom when you never measured how far past peak the cliff sits.

  • How long should a soak test actually run to be worth doing?
    Long enough to outlast whatever resets state in production. If instances are redeployed or restarted every six hours, a soak shorter than that cannot reveal a leak with a one-day time constant — and it hints that production is relying on restarts. Also cover at least one full traffic cycle, so scheduled jobs and daily batch work run inside the window.
  • Where does a spike test fit alongside these three?
    A spike test jumps straight to high load instead of ramping, so it isolates arrival shape rather than volume. It exposes cold caches, autoscaler reaction lag, connection-pool warm-up, and thundering-herd effects that a gradual ramp lets the system absorb. Run it whenever real traffic arrives instantly — a scheduled on-sale, a push notification, or failover of another region's traffic onto yours.
  • Why should a stress test always ramp back down at the end?
    To test recovery, which is a separate property from the ceiling. Some services never dig out of a queue backlog even once the offered rate is survivable again, so they need a restart to recover. Knowing that in a test is the difference between a five-minute incident and an hour-long one.

saying these in an interview costs you the question

  • Calls any traffic generation a load test regardless of goal
  • Reports the stress-test breaking point as the service's capacity
  • Believes a fifteen-minute run can reveal a memory leak
  • Treats zero errors at peak as proof of headroom
  • Stops the stress test at the peak, never testing recovery

context

open as a page

During a ramp load test you plot achieved throughput and p99 latency against the offered request rate. How do you identify the knee of the latency curve, and what number do you take away as the service's usable limit?

level: middleimportance: must knowfreq 55%

basics

~20 s

The knee is where achieved throughput stops tracking the offered rate and p99 latency starts climbing super-linearly — queues are no longer draining. The usable limit is the highest sustained rate that still meets the latency target, which sits below the knee, not at peak throughput.

open as a page

Your team validates capacity in a staging environment that is a one-tenth-scale copy of production. What makes the resulting throughput number untrustworthy, and what would you do to get a number you can actually act on?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Scaled-down environments do not scale linearly: smaller datasets inflate cache hit rates, shared tiers do not shrink with the app tier, stubbed dependencies answer far too fast, and synthetic traffic lacks real skew. Fix it by testing the shared tier separately, replaying real traffic, or measuring capacity in production under controlled conditions.

open as a page

In a load test, your service shows CPU around 45% across the fleet, yet p99 latency has tripled and total throughput has stopped rising. Which saturation signals do you look at, and what would each one tell you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Utilisation is not saturation. Look for evidence of waiting rather than of being busy: request queue depth and time-in-queue, thread-pool and connection-pool wait, downstream call latency, lock contention, and per-instance skew. Each names a different bottleneck, and each implies a different fix.

open as a page

A full capacity ramp test takes hours, so it cannot run on every commit. How would you catch a capacity regression — a change that quietly cuts maximum sustainable throughput by 20% — before it reaches production?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Stop measuring the maximum. Run a short test at a fixed request rate well below the limit and compare resource cost per request and latency against a stored baseline. Cost per request drifts detectably long before the ceiling moves, and it needs only minutes of stable load.

open as a page