skip to content

How would you build a workload model from real production traffic and keep it representative over time?

level: principalimportance: should knowfreq 42%

answer

  1. Evidence before opinion
  2. A model, reviewed, not a folder of scripts
  3. Which day are you reproducing
  4. Real data is skewed, synthetic data is flat
  5. Models rot as the product changes

basics

~20 s

Derive the transaction mix, arrival shape and data profile from observed traffic rather than opinion, state the reference period, and keep the model a versioned, owned artefact that is re-derived on a schedule instead of ageing into fiction.

solid answer

~50 s

Start from evidence: counts per operation, arrival pattern over the day, session structure, payload sizes and the distribution of the data the requests touch. Turn that into an explicit model — a weighted transaction mix, an arrival process matching how requests really arrive, think or pacing values drawn from observed gaps, and a data pool with realistic skew rather than uniform synthetic rows. Choose the reference period deliberately and say which one: a representative busy day, a seasonal peak, or a named incident. Then treat the model as a product artefact — versioned, owned, reviewed when a feature ships, and re-derived on a cadence — because the fastest way to a useless performance suite is a mix frozen at the shape the product had two years ago. Pass criteria are agreed alongside it, per run shape, in terms the product understands, and a run whose workload has drifted from production is reported as invalid rather than green.

go deeper

for a junior

Be ready to say that the mix of operations and the data behind them should come from what real traffic does, and that a test built on guesses can pass while real users suffer.

for a middle

Explain the model's components — mix, arrival shape, session structure, think time, data profile — and where each one comes from in observed traffic rather than in assumption.

for a senior

Show how you validate the model against production statistics, choose and state a reference period, and refuse to sign off a run whose workload no longer matches what the system receives.

for a principal

Own the model as a governed artefact: ownership, versioning alongside results, a re-derivation cadence, drift treated as invalidity, agreed pass criteria per run shape in business terms, and the privacy discipline behind any model derived from real traffic.

### A model, not a script A performance run is a simulation, and its credibility rests entirely on how well the simulated demand resembles real demand. The deliverable is therefore not a script but a **workload model**: an explicit, reviewable description of what the system is asked to do, from which scripts are generated. Writing it down separately is what makes it arguable — a stakeholder can dispute a mix table, but nobody meaningfully reviews a folder of recorded scripts. ### What the model has to contain **Transaction mix.** The relative frequency of each business operation. Derived from counts over a stated period, not from a workshop. Watch for the trap where the most-discussed operation is rare and a boring one dominates: the ratio between a cheap read and an expensive write often decides everything about the result. **Arrival shape.** How requests arrive, not merely how many: the daily and weekly profile, the burstiness within an hour, and whether arrivals are independent or clustered into sessions. A flat rate equal to the daily mean can be an order of magnitude gentler than the busiest ten minutes. **Session structure.** Which operations occur together and in what order, since a realistic sequence exercises caches, sessions and locks the way one-shot requests never will. **Think time and pacing.** Drawn from the observed distribution of gaps between actions in the same session, with spread. Constant pauses make a simulated population march in lockstep and produce artificial waves. **Data profile.** The distribution of the entities the requests touch. Real data is skewed — a few accounts enormous, most tiny; a few products hot, most cold. Uniform synthetic data hides both the pathological large case and the cache behaviour that real skew produces. **Environmental context.** Concurrent batch work, scheduled jobs, and the state of the datastore, since a run against an empty datastore is not a run against production behaviour. ### Choosing the reference period, and saying which Every model implies a moment in time. A representative busy weekday supports the routine question; a seasonal or campaign peak supports the readiness question; a reconstructed incident supports the regression question. These are different runs with different pass criteria, and the mistake is not choosing one — it is failing to state which one a green result refers to. Anonymise or synthesise the data behind any model derived from real traffic; a workload model must never become a route by which production personal data ends up in a test environment. ### Keeping it alive Models rot. Features ship, a channel takes off, a client library starts batching, an integration partner triples its calls. A mix that was right eighteen months ago will pass runs while the real risk moves elsewhere. Practical countermeasures: - **Own it.** A named owner per model, in the same review path as the service it exercises. - **Re-derive on a cadence.** Recompute from fresh traffic on a schedule, and diff against the current model. A large diff is itself a finding worth discussing. - **Tie it to change.** A new user-facing operation is not done until the model accounts for it — the same discipline as updating a contract or a dashboard. - **Version it.** Store the model with the results it produced, so a comparison across releases can prove that the workload, not just the code, was constant. - **Report drift as invalidity.** A run whose model no longer matches observed traffic is reported as stale, not as passing. ### Pass criteria that survive contact with stakeholders Each run shape needs its own criterion, agreed in advance and expressed in business terms: a percentile of user-visible latency for a named operation and an error-rate limit under the expected mix; a defined *breaking mode* for a run pushed past the target — degraded but recovering, rather than data loss or a failure to recover; bounded drift over a long run for memory, latency and error rate; and a recovery-time expectation after a sudden surge. Criteria stated as a percentile and a limit are debatable and improvable. Criteria stated as *looks fine* are neither. ### A worked example An insurance quote engine has been exercised nightly for a year by a 6-hour run built from a launch-era model: 82% single-driver motor quotes, 18% multi-vehicle. Traffic has since shifted — a partner channel now submits multi-vehicle quotes with an average of 3.4 vehicles each, at 41% of volume, and each one fans out to several rating calls. The nightly run stays green while real evening latency degrades, because the suite is faithfully reproducing a product that no longer exists. Re-deriving the mix from a recent busy week, re-seeding the data pool with the observed distribution of vehicles per quote, and rebuilding the arrival shape from the real evening profile reproduces the degradation on the second run. The subsequent review also caught an off-by-one boundary in the data pool generator, which had capped the vehicles-per-quote field one below the maximum the product allows — so the single most expensive case in production had never once been exercised. Both defects have the same root: the workload was an artefact nobody owned.

  • How would you validate that a workload model actually resembles production?
    Compare the run's own output against production statistics over the reference period: operation mix as executed, achieved rate profile, payload size distribution, and downstream call counts per business transaction. The last one is the sharpest check, since a model with the right request mix but the wrong data can produce a very different amount of downstream work. Publish the comparison with the results.
  • Two teams disagree about the transaction mix. How do you settle it?
    By measurement, and if the data genuinely supports both readings, by running both as named scenarios rather than by averaging them into a compromise nobody believes. A blended mix often represents no real hour of the day. Two clearly labelled models with their own criteria give a decision-maker more than one number that split the difference.
  • What do you do when production traffic cannot be observed in enough detail to derive a model?
    State the assumption explicitly and make the model falsifiable: build from whatever partial evidence exists, record every guess as a named parameter, and run sensitivity checks on the parameters that matter most. Then invest in the telemetry gap, because a workload model is only as good as the observation behind it and no amount of test engineering substitutes for knowing what the system actually receives.
  • How do you keep long performance runs affordable without letting the model degrade?
    Tier them: a short, cheap run on a reduced but proportionally faithful mix for routine feedback, and a full-shape run on a schedule for the criteria that need duration. What must never be cut to save time is the model's fidelity — shrinking the mix to the two easiest operations turns a cheap run into a misleading one, whereas shortening the steady state merely narrows what it can claim.

A workload model is a weather forecast for your system: derived from observation, stated for a named period, and worthless the moment nobody updates it.

saying these in an interview costs you the question

  • Deriving the transaction mix from a workshop instead of from traffic
  • Using a flat daily-average rate as the peak workload
  • Seeding uniform synthetic data with none of production's skew
  • Leaving the model unowned and unchanged for years
  • Reporting a green run without saying which period it reproduces
  • Copying production records into a test environment unanonymised

context