skip to content

Your team validates capacity in a staging environment that is a one-tenth-scale copy of production. What makes the resulting throughput number untrustworthy, and what would you do to get a number you can actually act on?

level: seniorimportance: should knowfreq 42%

answer

  1. a scale factor is not a multiplier
  2. shared tiers do not shrink
  3. dataset size changes cache hit rate
  4. real traffic or real production

basics

~20 s

Scaled-down environments do not scale linearly: smaller datasets inflate cache hit rates, shared tiers do not shrink with the app tier, stubbed dependencies answer far too fast, and synthetic traffic lacks real skew. Fix it by testing the shared tier separately, replaying real traffic, or measuring capacity in production under controlled conditions.

solid answer

~50 s

The core problem is that dividing an environment by ten does not divide the workload's behaviour by ten. The dataset is smaller, so cache hit rates are unrealistically high and index lookups cheaper — the per-request cost you measure is simply lower than production's. Shared tiers usually do not scale down at all: one database, one cache cluster, one gateway, and whichever one saturates first in production may not even be exercised in staging. Stubbed dependencies replying in a millisecond move the bottleneck somewhere else entirely. And synthetic traffic is too uniform: no hot keys, no oversized payloads, no bot traffic, no cold-cache long tail. To get a usable number I would size the shear-able application tier from a per-instance test, test each shared tier against its own load separately, and then get the real answer from production — replaying or shadowing real requests, or draining part of the fleet during a low-traffic window so real traffic concentrates on fewer instances, with a defined abort threshold.

go deeper

for a junior

Know that a smaller test environment does not simply produce a proportionally smaller version of production behaviour, and that a result from staging should not be multiplied up into a capacity figure.

for a middle

Be able to name the specific fidelity gaps — dataset size and cache hit rate, shared tiers that do not shear, fast stubs, uniform traffic — and explain which direction each one biases the result.

for a senior

Show the decomposition into per-instance and shared-tier questions, and describe a real higher-fidelity technique you have used: traffic replay, shadowing, or concentrating live traffic by draining capacity, including the safeguards you put around it.

for a principal

Own the tradeoff between the standing cost of a high-fidelity pre-production environment and the organisational risk appetite for measuring in production, and set the rules under which teams are allowed to do the latter.

## Why a scale factor is not a multiplier A one-tenth environment is only trustworthy if every dimension of the system scales by the same factor, and none of them do. The extrapolation `staging_throughput × 10` fails for several independent reasons, each of which can be worth a factor of two or more on its own. **Dataset size.** This is the biggest and most under-appreciated one. A tenth-size database has a smaller working set, so a far larger share of it fits in the cache and in the buffer pool; index B-trees are shallower; table scans that are catastrophic in production are cheap. The per-request cost you measure is not the production per-request cost, and the difference is usually in staging's favour. Query plans can also differ outright, because the planner chooses based on cardinality statistics. **Shared tiers do not shrink.** Application instances shear neatly; databases, caches, queues, service meshes, DNS, and gateways generally do not. If production runs one primary database and staging runs one primary database, staging is testing one-tenth of the load against 100% of that tier — so it will never reveal that the database saturates first. Conversely, if staging's database is a tiny instance, it saturates immediately and you conclude the app is slower than it is. **Stubbed or scaled-down dependencies.** A mock returning in 1 ms when production's dependency takes 40 ms changes concurrency, connection-pool pressure, and thread-pool behaviour completely. You are testing a different system. **Traffic realism.** Synthetic load is uniform: even key distribution, uniform payload size, no bots, no retry storms, no long tail of cold reads, and no correlation between requests. Real traffic has hot keys, heavy accounts, and users who trigger the expensive path. Uniform traffic hides hotspots — which is precisely the failure mode you were trying to find. **Environmental differences.** Instance families, kernel and runtime versions, whether the environment spans multiple availability zones (cross-AZ latency on every hop is real), network bandwidth allocation and burst credits, and the absence of noisy neighbours all shift the number. ## What to do instead, in ascending order of cost and fidelity **1. Split the question.** Measure the shear-able application tier as a per-instance figure — requests per second per instance at target latency, with realistic dependency latency injected. Separately, load-test each shared tier against its own expected aggregate load. This gives you two defensible numbers instead of one bogus one, and it tells you which tier is the actual constraint. **2. Fix the fidelity that is cheap to fix.** Restore a production-sized (anonymised) dataset even if the compute stays small — dataset realism buys more accuracy per dollar than anything else. Inject realistic dependency latency into stubs rather than letting them reply instantly. Derive the request mix and key distribution from real access logs rather than inventing it. **3. Replay real traffic.** Capture production requests and replay them at a controlled multiple against a test fleet. This solves traffic realism and shape at once. The caveats are real: replaying writes is dangerous, so most teams replay reads only or write to a shadow store; and personal data in captured payloads must be handled properly. **4. Shadow / mirror traffic.** Duplicate live requests to a parallel fleet whose responses are discarded. Fidelity is high and users are unaffected, but only for idempotent read paths, and you must ensure the shadow fleet cannot write to production stores or call third parties that charge or send email. **5. Measure capacity in production.** The highest-fidelity answer, and the one mature SRE organisations rely on. Two shapes: add synthetic load on top of real traffic during a low-traffic window, or — often better — *remove* capacity, draining instances or a zone so that real traffic concentrates on fewer servers and drives them toward saturation. The second requires no synthetic traffic and is inherently realistic. ## Testing in production is a decision with a cost It is also the fastest route to a self-inflicted outage, so it is only defensible with controls in place: a low-traffic window, an explicit abort threshold defined in advance on a user-visible signal ("stop if p99 exceeds X or error rate exceeds Y"), an operator watching in real time, a one-action way to stop or restore drained capacity, and prior announcement so nobody starts an incident response over your test. Ramp in steps and stop at the first sign of user impact rather than proceeding to the cliff — you are looking for the trend, not the breaking point. ## The honest framing in an interview The strong answer is not "never use staging". Staging catches functional regressions, gross inefficiencies, and relative changes between builds, all cheaply. What it cannot produce is an absolute capacity number you would size a fleet from. Say which question you are answering, be explicit about the fidelity gap, and describe the production measurement that anchors the absolute figure.

  • If you could improve only one thing about the staging environment, what would it be?
    A production-sized, anonymised dataset. Data volume drives cache hit rate, buffer-pool residency, index depth, and query plan choice, so a tiny dataset makes every request cheaper than reality in a way no amount of extra load generation corrects. Compute can often be scaled down honestly; data cannot.
  • What controls would you insist on before running a capacity test against production traffic?
    A low-traffic window, an abort threshold agreed in advance on a user-visible signal such as p99 latency or error rate, a named operator watching live, a single action that stops the test or restores drained capacity within seconds, and advance notice to on-call so nobody declares an incident over it. Ramp in steps and stop at first user impact, not at the cliff.
  • Why is draining capacity sometimes better than generating synthetic load in production?
    Because the traffic is genuinely real — correct mix, correct key skew, correct payloads, correct client retry behaviour — so nothing has to be simulated. You concentrate existing load onto fewer instances and watch them approach saturation. It also fails safe: restoring the drained capacity immediately reverses the experiment.

saying these in an interview costs you the question

  • Multiplies a one-tenth staging result by ten
  • Assumes stubbed dependencies do not change the bottleneck
  • Tests with a tiny dataset and trusts the per-request cost
  • Generates perfectly uniform traffic and calls it realistic
  • Runs load against production with no abort criteria

context