skip to content

What is volume testing, and how does it differ from a load test that raises request rate?

level: juniorimportance: must knowfreq 62%

answer

  1. Which variable is deliberately changed
  2. Rate held still, stored data grows
  3. Cost that rises with dataset size
  4. Development-sized data makes everything cheap
  5. Unbounded results and jobs outgrowing windows

basics

~20 s

Volume testing keeps the request rate fixed and grows the stored dataset instead: more rows, longer collections, larger files. It exposes defects that appear only at size, such as unbounded queries, memory that scales with result size, and jobs that outgrow their window.

solid answer

~40 s

A load run varies the arrival rate and asks how many concurrent requests the system can serve. A volume run holds the workload constant and varies how much data the system already holds, asking what breaks as the dataset grows by an order of magnitude. They find different defects. Load finds contention: queue growth, pool exhaustion, saturation. Volume finds size-sensitivity: a query whose cost rises with table size, a response that returns an entire collection, an export that materialises every row in memory, a nightly job whose runtime crosses its window. Volume-sensitive defects are usually invisible under load, because a development dataset of a few thousand rows makes every one of those operations cheap. Mature teams run both: fix the dataset and vary the rate, then fix the rate and vary the dataset.

code

pseudocode · 7 lines
pseudocode
for size in [47_000, 750_000, 12_400_000]:
    seed_parcels(count = size, shape = production_profile)
    warm_up(duration = 5.minutes)
    result = run_fixed_workload(rate = 30.per_second, duration = 20.minutes)
    record(size, result.p95_latency, result.peak_memory, result.bytes_per_response)

report_growth_shape(records)  # is cost rising faster than size?

go deeper

for a junior

Be ready to state the one-line difference: rate is the variable in a load run, stored data is the variable in a volume run. Have two concrete examples of a defect only volume finds, such as an endpoint returning an entire collection.

for a middle

Explain where size enters request cost — data examined, items returned, memory held — and why a small seeded database makes all three look cheap. An interviewer expects you to name the target size and say how you derived it from growth.

for a senior

Show the operating judgement: hold the workload constant so size is the only variable, measure at several sizes rather than one, and assert on memory and response size alongside latency. Be able to describe a real degradation you found this way.

for a principal

Own the tradeoff of what the organisation pays for. Maintaining a large, realistically shaped dataset costs storage, refresh time and engineering attention, so decide which services justify it, how often it is regenerated, and what the alternative signal is for the rest.

## Two different independent variables Every non-functional run has one thing you deliberately change and everything else you try to hold still. In a load run the changing thing is **arrival rate**: how many requests per second reach the system, or how many concurrent users are active. In a **volume run** the changing thing is the **stored dataset**: the number of rows in a table, the number of items in a collection an endpoint returns, the number of files in a bucket, the depth of a history, the size of an export. That single difference decides which defects each run can find. A load run at a fixed, small dataset can hammer a system for hours without ever touching a code path whose cost depends on how much data exists. A volume run at one request per second, against a dataset a hundred times larger, will find it immediately. ## What a volume-sensitive defect looks like The recurring families are worth memorising, because interviewers ask for examples: - **Unbounded reads.** An endpoint that returns *all* of something. With 40 rows per account it looks fine; with 90,000 it returns a multi-megabyte response and the serialiser dominates the request. - **Cost that rises with table size.** A lookup that was resolved cheaply on a small table degrades as the table grows, because the amount of data examined per request grows with it. The symptom is a latency curve that bends upward while the request rate is flat. - **Memory proportional to result size.** A report or export that builds the whole result in memory before writing it. Peak memory then tracks the dataset, and the failure mode is not slowness but an out-of-memory kill. - **Deep offsets.** Paging works on page 2 and collapses on page 4,000, because reaching a deep page costs more than reaching a shallow one. - **Jobs that outgrow a window.** A nightly reconciliation that took 12 minutes at launch takes 3 hours two years later. Nothing failed; the runtime simply crossed the window it was allowed. - **Long-collection rendering.** A screen or payload that is fine for a typical entity and unusable for the largest one. ## Why development-sized data hides all of it A seeded development database is usually two or three orders of magnitude smaller than production and, worse, evenly shaped. Everything fits in cache, every scan is short, every collection is small, and every one of the defects above measures as fast. This is the core reason the discipline exists as a named activity: correctness testing and load testing both pass, and the system still degrades in production, because neither run ever changed the variable that mattered. Consider a parcel-tracking gateway. Its team ran a healthy load profile every week against a seeded database of about 47,000 parcels, and the tracking-history endpoint sat at a p95 of 180 ms. In production, after eighteen months, the same endpoint at the same request rate sat at 4.3 seconds. No release caused it. The dataset had grown to 12.4 million parcels and nothing in the test ever asked what happens at that size. ## How a volume run is actually built 1. **Choose a target size.** Usually the projected dataset one to two years out, derived from the current growth rate, not a round number picked for comfort. 2. **Seed to that size with realistic shape**, not uniform filler, so worst-case entities exist. 3. **Hold the workload constant.** Same request mix, same modest rate, same warm state. If the rate moves too, you cannot attribute a change to size. 4. **Measure at more than one size.** A single pass/fail at one size tells you almost nothing about the trend; two or three sizes tell you whether cost grows in step with data or faster. 5. **Assert on more than elapsed time.** Peak memory, response size, and a work-done proxy such as rows examined per request all reveal size-sensitivity earlier than the clock does. ## Where it sits among the other runs Volume testing is a peer of, not a replacement for, the rate-driven runs. It also overlaps with, but is not the same as, an endurance run: an endurance run varies **time** and finds drift and leaks; a volume run varies **stored size** and finds cost that scales with it. In many systems the two interact, because a long-running system accumulates data. Keeping the vocabulary straight is exactly what an interviewer is checking when they ask for the difference.

  • Where does dataset size actually enter the cost of a single request?
    In three places. The amount of stored data examined to satisfy a lookup, which grows if the access path degrades as the table grows. The number of items returned, which drives serialisation, transfer and client rendering. And the amount held in memory at once, when a handler materialises a whole result before emitting it. A useful habit is to ask, for each endpoint, which of the three is proportional to stored size, since anything proportional is a candidate defect.
  • If building a full volume run is expensive, what is the cheapest first version worth having?
    Seed one representative table to the projected size, then run your existing functional regression pack against it and record per-case timings and peak memory. You are not measuring throughput; you are looking for the handful of cases that go from milliseconds to seconds. That single afternoon usually surfaces the unbounded read and the in-memory export, which are the two most common findings, and it justifies building the proper multi-size run afterwards.

A load test asks how many customers can be served at once; a volume test asks what happens when the filing cabinet behind the counter grows from one drawer to a warehouse, even with a single customer at the window.

saying these in an interview costs you the question

  • Says volume testing is just load testing with more users
  • Assumes a few thousand development rows are representative
  • Measures only elapsed time, never memory or response size
  • Believes faster hardware removes cost that scales with data
  • Treats a pass at today's size as proof for next year
  • Confuses growing the dataset with running for longer

context