skip to content

What does a generated dataset with exactly two children per parent fail to exercise?

level: seniorimportance: should knowfreq 44%

answer

  1. Two per parent is nobody's real data
  2. Both ends of the count distribution
  3. Many parents have none, one has thousands
  4. Quantiles, not an average child count
  5. A constant multiplier hides arithmetic bugs

basics

~20 s

Everything at the ends of the distribution: parents with no children at all, and the rare parent with thousands. A flat two-per-parent set never reaches empty-list handling, paging past the first page, chunking, truncation or tie-heavy ordering.

solid answer

~50 s

Real relationship counts are heavy-tailed: a large share of parents have **no** children, most have one or a few, and a thin tail runs to thousands. A uniform two-per-parent set sits entirely in the middle of that curve, so three families of behaviour never run. The **zero** case - an empty list rendered, an aggregate over no rows, a combining query that legitimately drops the parent - is where most empty-state defects live. The **extreme** case reaches paging past the first page, chunked processing, truncation and size limits. And the uniformity itself hides arithmetic: when every parent has exactly two children, a query that wrongly duplicates rows returns a plausible small number nobody questions. Generate to measured quantiles - the share at zero, the median, a high quantile, the maximum - and pin the extreme parents to stable identifiers so tests can address them.

code

pseudocode · 17 lines
pseudocode
fanOut = distribution {
    shareWithNone: 0.34            // parents that never acquire a child
    typical:       quantiles { p50: 1, p75: 3, p90: 11 }
    tail:          { p99: 240, max: 18400 }
}

for parent in parents:
    childCount = fanOut.sample()
    generate(childCount, "child", { parentId: parent.id })

// anchors tests can address by name, stable across builds
forceChildCount(parent = "PARENT-NONE",    count = 0)
forceChildCount(parent = "PARENT-EXTREME", count = fanOut.tail.max)

// assert the shape produced, not the shape configured
failBuildUnless(abs(shareOf(childCount == 0) - fanOut.shareWithNone) < 0.02)
failBuildUnless(max(childCount) >= fanOut.tail.max)

go deeper

for a junior

Know what fan-out means - how many child rows hang off one parent - and that real counts are uneven: many parents have none and a few have very many. Uniform test data is the exception, not the norm.

for a middle

Explain which code paths only the ends reach: empty-list handling at one end; paging, chunking and truncation at the other. Be able to say how you would specify a shape rather than a single average child count.

for a senior

An interviewer expects the subtler cost too: a constant fan-out makes arithmetic bugs return plausible numbers. Show how you take quantiles from production counts, pin the extreme parents to stable identifiers, and assert the shape actually produced.

for a principal

Own the standard for shared generated datasets: which relationship distributions every environment must reproduce, where those figures come from, and how they are refreshed as the real distribution moves, so each team is not separately guessing a shape.

## What fan-out is, and why its shape is the point **Fan-out** is how many child rows hang off one parent: orders per customer, comments per post, line items per invoice, sessions per device. A generator that produces a fixed small number for every parent — the classic two — creates a set whose relationship counts are all identical. Real relationship counts almost never are. They are **heavy-tailed**: a large share of parents have none at all, most of the rest have one or a handful, and a thin tail runs to hundreds or thousands. That difference is not cosmetic. A uniform set sits entirely in the middle of the curve, so whole families of behaviour never run. This is a question about the **shape** of the generated set, not about how much of it there is. ## The three regions, and what only each one reaches | Region | What real data looks like there | What only this region exercises | |---|---|---| | Zero | A large share of parents never acquire a child | Empty-list rendering, aggregates over no rows, a combining query that legitimately drops the parent, the "nothing here yet" copy | | Typical | One to a handful | The ordinary path — and the only one a uniform set covers | | Tail of the fan-out distribution | A thin set of parents with hundreds or thousands | Paging past the first page, chunked processing, output truncation, size limits on a payload or a field, ordering that is only unstable among many ties | The zero region is the one that costs teams the most, because "a parent with no children" is not an exotic state — it is the state every parent starts in, and it is where the majority of empty-state defects live. It is also the region a naive generator is least likely to produce, because a loop that runs "for each parent, make some children" tends to have a floor of one. ## Skew lives on more than one axis - **Degree on the other side.** In a many-to-many relationship, most children belong to one parent and a few belong to hundreds. A generator that fixes the count on one side only still produces a uniform set on the other. - **Reuse skew in referenced values.** A handful of lookup values claim most of the rows: the popular product, the default category, the one country. Grouping, caching and deduplication behave differently when one value dominates, and a set that spreads references evenly never shows it. - **Depth in a hierarchy.** Most chains are two levels deep and one is fifteen. Recursive walks, depth limits and cycle guards only run against the deep one. Note the boundary: a parent with zero children is a legitimate point on the distribution, not a broken row. Manufacturing rows that are deliberately malformed is a separate exercise with a separate purpose. ## Why uniformity hides arithmetic bugs This is the effect people miss. When every parent has exactly two children, a query that accidentally duplicates rows returns exactly twice the right answer — and twice a small number is still a plausible small number. A wrong grouping key, a combining step applied one level too high, a total summed over the wrong set: each produces a value wrong by a small constant factor, which is the least visible kind of wrong. Against a varied set the same bug produces a number that matches nothing anybody expected, and a single assertion catches it. ## How to specify the shape you want 1. **Take quantiles from production counts, never a mean.** The mean of a heavy-tailed count is a number no parent actually has. What you want is the share at zero, the median, a high quantile and the maximum. 2. **Fit the generator to hit those figures**, including the mass at zero, which is the one most configurations quietly omit. 3. **Pin the anchors.** Give the parent with no children and the parent with the largest count stable, well-known identifiers that do not change between builds, so a test can address them directly instead of scanning for one and hoping. 4. **Assert the realised shape, not the configured one.** After generation, measure the share at zero and the maximum actually produced, and check them against the specification. A generator's configuration and its output part company quietly — a filter applied downstream, a cap somebody added, a rounding rule — and nobody notices until the assertions that depended on the tail stop meaning anything. ## What good enough looks like You do not need to reproduce the real distribution faithfully. You need every region represented: a meaningful population of parents with none, the ordinary middle, and at least one parent far enough into the tail of the fan-out distribution to trip paging, chunking and truncation. Getting those three represented is most of the value; matching the exact curve is refinement, and it is worth doing only where a feature's behaviour genuinely turns on the curve rather than on the extremes.

  • Where do the target fan-out numbers come from if you cannot copy production rows into a test dataset?
    Counts are not the rows. Aggregate the relationship counts in production into a few figures - the share of parents with none, the median, a high quantile and the maximum - and carry only those numbers across. That is enough to fit a generator, and it describes no individual. Refresh the figures occasionally, because a distribution measured two years ago is a guess about today.
  • How do you stop a skewed generated set from making tests non-deterministic?
    Skew does not cause it; randomly reassigning the skew on every build does. If which parent receives the large count changes each time, a test that picks a parent sees a different count each run. Pin the anchors: the parent with none and the parent with the largest count keep stable identifiers across builds, and only the unremarkable middle is allowed to vary.

A photograph of a crowd in which everyone is exactly average height tells you nothing about whether the doorway is tall enough.

saying these in an interview costs you the question

  • Says two children per parent is a realistic average
  • Never generates a parent with zero children
  • Dismisses the largest fan-out as an edge nobody hits
  • Fits the generator to a mean instead of quantiles
  • Assumes the configured distribution equals the one produced