How do you design a volume run that reveals super-linear cost growth before production reaches that size?
answer
- One measurement cannot show a trend
- Three sizes, everything else held constant
- Compare the shape of the curve
- Work per request warns before the clock does
- Convert the bend into a calendar date
basics
~20 sRun the same fixed workload against several dataset sizes — today's, four times, sixteen times — and compare the shape of the curve rather than one pass or fail. Anything whose cost rises faster than the data is the finding.
solid answer
~50 sA single measurement at a single size tells you almost nothing, because you cannot tell a cost that grows in step with the data from one that grows faster. So measure at three or more sizes with everything else held constant: same request mix, same modest rate, same hardware, same warm-up excluded. For each size record latency percentiles per endpoint, peak memory, response size, and a work-done proxy such as records examined per request. Then read the shape: cost proportional to size is expected and can be planned for, cost rising faster than size is a defect and will arrive on a date you can calculate from the growth rate. Report the bend, not the absolute number, and assert separately that memory stays bounded regardless of result size, because a memory failure is a crash rather than a slowdown.
code
pseudocode · 10 linessizes = [47_000, 750_000, 12_400_000]
measurements = [measure_fixed_workload(size) for size in sizes]
for a, b in consecutive_pairs(measurements):
size_factor = b.size / a.size
cost_factor = b.records_examined_per_request / a.records_examined_per_request
assert cost_factor <= size_factor, \
"work per request grows faster than the dataset"
assert max(m.peak_memory for m in measurements) < memory_limit * 0.7go deeper
Be ready to explain why one measurement at one dataset size cannot tell you whether cost grows in step with the data or faster, and name at least two sizes you would compare.
Explain the controls that make the comparison valid — same workload, same hardware, same warm-up rule, same shape at every size — and name what you record beyond latency, especially peak memory and response size.
Demonstrate reading the curve rather than a threshold: identify sub-linear, linear and super-linear growth, use work per request as the early signal, and translate the bend into a date from the measured growth rate.
Own the economics and the standard. Decide which services get a multi-size run, what runs per merge versus on a schedule, how snapshots are stored and refreshed, and how a growth-curve finding enters roadmap planning rather than dying in a report.
## The question a volume run should answer "Is it fast enough at ten million rows?" is a weak question, because the answer expires. The strong question is **"how does cost grow as the data grows, and on what date does that curve cross the limit?"** Answering it requires more than one measurement, which is the single most common gap in volume testing practice. ## Holding everything else still Size must be the only thing that changes between runs, or the comparison is worthless. In practice that means: - **Same request mix and same rate.** Deliberately modest — you are not looking for saturation, you are looking for size-sensitivity. A rate low enough that queueing is negligible keeps the measurement clean. - **Same hardware and same resource limits** across sizes, including memory limits, because a larger dataset changes cache behaviour and you want that effect to be real rather than an artefact of a different machine. - **Same warm state.** Define a fixed warm-up and exclude it from the measurement window, and apply the identical rule at each size. - **Same shape profile at every size.** If the small dataset is uniform and the large one is skewed, you have changed two variables and the curve is meaningless. - **Repeat each size** at least twice. A single outlier run at one size can invent or hide a bend. ## What to record Four families, and the last two are the ones teams forget: 1. **Latency percentiles per endpoint**, never a single aggregate across the whole workload — one degrading endpoint disappears inside a healthy average. 2. **Peak memory** for the process and, if you can get it, per-request allocation. This is what catches a handler that materialises an entire result. 3. **Response size**, which catches unbounded reads directly: if bytes returned grow with the dataset, the endpoint has no bound on it. 4. **A work-done proxy**, such as the number of stored records examined per request. This is the earliest signal of all, because work per request can double while wall-clock latency still looks fine thanks to caching, and it tells you the problem is coming before it hurts. ## Reading the shape With three sizes, say 1x, 4x and 16x, the arithmetic is easy enough to do in your head: - Cost flat as size grows: bounded work. Ideal. - Cost rising roughly with the logarithm of size: a healthy indexed lookup shape. - Cost rising in proportion to size: linear. Often acceptable for a batch job, rarely acceptable for an interactive request, and always something to plan capacity around. - Cost rising faster than size — doubling the data more than doubles the cost: this is the finding. It is the shape of a nested scan, a per-item lookup inside a loop over a growing collection, or a degraded access path. Then convert the bend into a date. If the dataset grows at a measured rate, and the curve crosses the acceptable threshold at some size, you can say "this endpoint stops meeting its target in roughly seven months" — a statement a product owner can act on, which a raw millisecond figure is not. ## The usual offenders Three patterns account for most super-linear findings, and it is worth naming them for an interviewer: - **Deep offsets.** Requesting a far-in page costs more than requesting an early one, because reaching the page is itself work. Test page depth explicitly — first page, middle page, last page — rather than always paging from the start, which is what a naive script does and why deep-page cost stays invisible. - **Streamed versus materialised exports.** An export must be asserted on **peak memory**, not just duration. The correct claim is that memory stays flat while output size grows; if it tracks the row count, the export is materialising and it will eventually be killed rather than merely slow. - **Per-item work inside a growing collection.** A loop that issues one lookup per element is invisible at 40 elements and fatal at 90,000. ## A worked example A parcel-tracking gateway ran this at 47,000 / 750,000 / 12,400,000 parcels with a fixed 30-requests-per-second mix. Tracking-history p95 went 180 ms, 640 ms, 4,310 ms. Between the first two sizes the data grew about 16x and latency about 3.6x — sub-linear, healthy. Between the second and third the data grew about 16.5x and latency 6.7x — still sub-linear, but the work-done proxy told a different story: records examined per request went 14, 51, 2,940, which is faster than the data grew. The endpoint was fine on the clock and already broken in its shape; six months later, at production scale, it was the top-latency endpoint. The export path told the cleaner story: peak memory 310 MB, 1.2 GB, and an out-of-memory kill at the third size, which is a pass/fail no percentile could soften. ## Making it affordable The honest objection is cost: three shaped datasets are storage plus generation time. Two mitigations work. First, do not run all sizes on every change — run the largest size on a schedule and only the smallest on each merge, comparing the small size against its own history to catch a change in work per request. Second, keep the sizes as restorable snapshots so a run starts from a fixed state rather than regenerating. For an 11-person team, one scheduled multi-size run per week plus a per-merge work-per-request check is a realistic standing cost.
- Why measure records examined per request when you already have latency?Because latency lags the defect. Caching, a warm working set and spare hardware can keep wall-clock time acceptable long after the work per request has started growing faster than the dataset. The work proxy is monotonic in the thing you care about and is far less noisy than timing, so it detects the bend earlier and gives a stable signal you can gate a merge on. When the proxy and the clock disagree, believe the proxy about the trend and the clock about today.
- How do you keep a multi-size volume run affordable for a small team?Split it. Run the full multi-size comparison on a schedule — weekly is usually enough, since dataset growth is slow — and on every merge run only the smallest shaped dataset, comparing work per request against its own recorded history. Keep each size as a restorable snapshot so runs start from a fixed state instead of regenerating for hours. That gives fast per-merge protection against a shape regression and a slower, thorough answer about the growth curve.
- What makes a comparison between two sizes invalid?Any second variable. Different hardware or memory limits, a different request mix or rate, a different warm-up rule, a different shape profile between the small and large datasets, or a code change between the two runs. Also a single unrepeated run per size, since one noisy result can manufacture or conceal a bend. Record the configuration alongside the measurements so a later reader can tell whether two numbers are actually comparable.
saying these in an interview costs you the question
- Measures at one dataset size and calls it a trend
- Reports one aggregate latency across all endpoints
- Always pages from the first page, never a deep one
- Ignores peak memory because the run did not time out
- Changes hardware or workload between size comparisons
- Reports milliseconds instead of when the limit is crossed