skip to content

Why does uniformly generated seed data understate how a system behaves at volume?

level: middleimportance: should knowfreq 54%

answer

  1. Size is not the same as shape
  2. Flat distributions have no worst case
  3. Hot keys, long tails, rare variants
  4. Derive a profile from aggregate statistics
  5. Plant named worst-case fixtures, assert on them

basics

~20 s

Generated rows tend to be uniform: equal group sizes, same-width text, no gaps, one format. Real data is skewed, with hot keys, long tails and rare variants. Uniform seeds hide the worst case, so the run measures an average production never has.

solid answer

~40 s

Generating ten million rows is easy; generating ten million rows shaped like production is the actual work. A naive generator spreads records evenly across keys, gives every text field a similar length, leaves no gaps and emits one canonical format. Production does none of that: a few keys hold a large share of the rows, some entities have enormous collections while most have three, and a minority of records carry an odd variant that takes a slower path. Uniform seeding hides the two defects that matter most — the worst-case entity, which nobody ever renders or exports, and the rare-variant path, never hot enough in the test to show its cost. Derive a shape profile from aggregate production statistics, generate to it, and assert against the extreme entity rather than the median one.

code

pseudocode · 12 lines
pseudocode
profile = {
  parcels_per_merchant: { p50: 120, p95: 9_400, p999: 610_000, max: 4_720_000 },
  history_events_per_parcel: { p50: 6, p95: 41, max: 1_180 },
  format_variants: { canonical: 0.968, locale_dependent: 0.032 },
  absent_delivery_window_rate: 0.147,
  created_at_spread: 62.months
}

seed_from(profile, total_parcels = 12_400_000)
plant_named_fixture("worst_case_merchant")   # largest parcel count
plant_named_fixture("worst_case_history")    # longest event history
plant_named_fixture("locale_dependent_row")  # rare-variant slow path

go deeper

for a junior

Be ready to say that ten million identical-looking rows are not ten million realistic rows, and give one example of a shape that matters, such as one key holding a large share of the records.

for a middle

An interviewer expects you to list several shape axes — key skew, collection length, value length, absent values, variant mix, record age — and explain what each hides when it is flat. Say how you would measure them.

for a senior

Demonstrate the working method: a checked-in shape profile built from aggregate statistics, planted worst-case fixtures with names, and assertions written against the extreme entity rather than the mean. Mention refreshing the profile as production drifts.

for a principal

Own the standard and its cost. Decide who maintains the profile, how often it is refreshed, and how the organisation avoids the failure mode where a single frozen worst case gets optimised for while the real tail moves past it.

## Volume without shape is only half the variable When a team decides to test at scale, the first instinct is a loop that inserts rows until the count is right. That gets the size honestly and the *shape* completely wrong, and shape is where most volume-sensitive defects live. A dataset has at least six shape axes worth reasoning about: 1. **Key skew (cardinality and hot keys).** How rows distribute across the values you group and filter by. Real distributions are lopsided; generated ones are flat. 2. **Collection-length distribution.** How many children an entity has. Typically a long tail: most entities tiny, a handful enormous. 3. **Value-length distribution.** Text fields that are short on average and occasionally very long. 4. **Absent and default values.** Real columns have gaps, defaults and legacy placeholders; generators fill everything. 5. **Variant mix.** A minority of records carrying an older or region-specific representation that the code handles on a separate path. 6. **Temporal distribution.** When rows were created — a decade of history versus everything inserted in the last hour, which changes clustering and how much of the data is actually warm. ## What each one hides when it is flat **Flat key skew** makes every grouped operation cost the same. If the truth is that one key owns a large share of the rows, then the request that touches that key does far more work than the average, and that request is the one your users complain about. A uniform seed has no such request, so the measured distribution of latencies is narrow and comfortable and completely unlike production, where the tail is driven by the hot key, not by the network. **A flat collection-length distribution** removes the worst-case entity. If your generator gives every account exactly 40 items, no code path ever handles the account with 90,000. That account is where the unbounded response, the quadratic rendering loop and the memory spike live. **A flat variant mix** removes the slow path entirely. This one is subtle and it is worth an example. A parcel-tracking gateway seeded 12.4 million parcels with a single canonical weight-and-date representation, and its volume run looked clean. In production, roughly 3.2% of inbound records arrived with a locale-dependent numeric and date format, and the gateway parsed those through a fallback that constructed a formatter per record. At development size the fallback ran a few hundred times a day and cost nothing measurable. At production volume it ran on hundreds of thousands of records per hour during ingest, and the ingest job — previously 12 minutes — began missing its schedule. The defect was not the parser; it was the seed data, which had no variants in it at all, so the path was never sampled. **Flat temporal distribution** distorts warmth. If everything was written moments ago, an unrealistic share of it is in cache and the run measures memory-speed access where production would go to storage. ## Building a profile instead of a count The practical method is to describe the shape as a small set of numbers and generate against it: - rows per key at the 50th, 95th and 99.9th percentile, plus the single largest; - the length distribution of the child collections, again by percentile; - the fraction of records in each format or status variant; - absent-value rates per field; - the age spread of records. Those are aggregates, so they can be shared without exposing anything sensitive, and they change slowly enough that a quarterly refresh is usually sufficient. (Where the underlying records come from, and how they are masked or subsetted, is a separate concern with its own handling; what matters here is the shape you are aiming at.) Then make the extremes explicit rather than hoping the random draw produces them. A good generator plants **named worst-case fixtures** — the account with the largest collection, the key holding the biggest share, the record in the rare format — so a case can assert against them by name and fail loudly when a change makes the worst case worse. Randomly hoping to hit the tail is how tail defects escape. ## What to assert once the shape is right Assert against the extreme, not the median. "The tracking history renders in under a second for the parcel with the longest history" is a useful oracle; "average render time is 90 ms" is not, because the average is dominated by tiny entities that were never at risk. Equally, report which entity produced the worst measurement, so the number is traceable to a fixture someone can reproduce. A small team can afford this. An 11-person team does not need a data-engineering function to keep five percentile figures and a variant mix in a checked-in profile file and regenerate a shaped dataset on a schedule; what it cannot afford is discovering the shape for the first time from a production incident.

  • How would you discover the shape of production data without extracting the records themselves?
    Query for aggregates only: counts grouped by the key of interest reduced to percentiles, length percentiles for the collections and text fields, the proportion of rows in each status or format variant, absent-value rates, and the spread of creation dates. Those outputs are a handful of numbers, carry no record content, and are cheap to refresh on a schedule. Store them as a checked-in profile so the generator and the reviewers are looking at the same target.
  • Is there a risk in seeding the worst case as a named fixture rather than drawing it randomly?
    Yes, and it is worth naming: a fixed worst case can be optimised for specifically while the real tail moves past it. Mitigate it by refreshing the profile against production periodically, and by keeping some randomised generation alongside the planted fixtures so new shapes can appear. The planted fixture gives you a stable, reproducible assertion; the randomised remainder gives you the chance of discovering something you did not think to plant.

saying these in an interview costs you the question

  • Thinks a correct row count makes the dataset realistic
  • Spreads records evenly across every key by default
  • Reports average measurements when the tail is the risk
  • Never seeds rare variants, so slow paths stay unsampled
  • Assumes randomly generated data will hit the extremes
  • Fills every field, so absent-value paths are untested

context