skip to content

How do authored generation rules and a generator fitted to real records differ in what the test data guarantees?

level: middleimportance: must knowfreq 48%

answer

  1. Two families of manufactured test data
  2. One is written down, one is learned
  3. Ask what each can promise per row
  4. Unstated structure comes from only one

basics

~20 s

Authored rules guarantee exactly the constraints someone wrote down, and nothing else. A generator fitted to real records reproduces that data's shape, including correlations nobody stated, but guarantees no particular constraint holds in every produced row.

solid answer

~50 s

Both manufacture data instead of copying it, but the guarantee sits in different places. An **authored rule set** states ranges, allowed values and cross-field invariants directly: an age between 18 and 120, a closing date never earlier than an opening date. Every produced row satisfies exactly those statements, so the dataset is predictable and any row is explainable by pointing at a rule. It is also only as rich as what the author thought to write, so realistic skew, unstated correlations and odd-but-real values are simply absent. A **generator fitted to a real extract** learns that extract's shape: value frequencies, field lengths, which fields move together. The output resembles production and carries structure nobody specified, which is where it earns its keep. But nothing about it is promised: an invariant the product depends on may be violated in some fraction of rows.

code

pseudocode · 16 lines
pseudocode
# authored rule set - the guarantee is the text itself
rules for customer_record:
    age         in 18..120
    country     one of ALLOWED_COUNTRIES
    opened_at   <= today
    invariant   closed_at is null or closed_at >= opened_at
    invariant   balance >= 0

row = produce(rules)
assert every rule holds            # true by construction
assert absence_rate(middle_name)   # nothing was stated, nothing to assert

# fitted generator - the guarantee is resemblance to the extract
generator = fit(shape_of(extract_of_real_customers))
row = generator.sample()
assert closed_at >= opened_at      # holds only as often as it did in the extract

go deeper

for a junior

Be ready to say what it means to manufacture test data instead of copying it, and to name the two ways it is produced: rules a person writes, or a generator fitted to a sample of real records.

for a middle

Explain where each guarantee lives. Authored rules hold exactly what was written and nothing more; a generator fitted to real records promises resemblance in aggregate and promises nothing about any single produced row.

for a senior

Show you have been bitten. Describe a dataset that satisfied every stated rule and still missed a defect, and an intermittent failure caused by a produced row nobody had constrained.

for a principal

Own the composition. Argue for fitting for shape and asserting invariants over the produced rows, and be able to say what that costs in ownership, review effort and protected-data scope.

## Two different promises Manufactured test data comes from one of two families, and they promise different things. An **authored rule set** is a specification a person writes: field ranges, allowed value sets, format rules, and invariants that hold across fields — a closing date is never earlier than an opening date; a settled amount never exceeds the authorised amount. A producer walks the rules and emits rows that satisfy them. The guarantee is exact and narrow: *every row satisfies the stated rules, and the dataset says nothing at all about anything you did not state.* A **generator fitted to a real extract** works the other way round. You take an extract of real records, fit a generator to its shape — how often each value occurs, how long text fields run, how many records sit in each state, which fields move together — then sample new rows from that fit. The guarantee is statistical and broad: *the output resembles the extract in aggregate, and it carries structure nobody wrote down.* What it does not promise is that any particular row obeys any particular rule. ## What each one gives you, and what it costs | Dimension | Authored rules | Generator fitted to real records | |---|---|---| | Where the guarantee lives | In the text of the rules | In resemblance to the source extract | | Per-row promise | Every stated invariant holds | None; an invariant holds only as often as it did in the source | | Unstated structure | Absent — nothing you did not write | Present — correlations, skew, long tails in value frequency and text length | | Explaining one row | Point at the rule that produced it | Inspect the row; there is no rule to point at | | Review | Read and diff the rules | Review the extract, the fitting step and the output | | Privacy exposure | None beyond invented values | Derived from protected data | | Setup cost | Write rules; grows with the domain | Build a fitting pipeline; needs a standing owner | | Rare combinations | Only the ones you wrote | Under-represented; the fit pulls toward the common case | ## The failure mode of each Authored rules fail by **flattering the product**. The author writes what the system expects, so the data is exactly the data the system already handles: every optional field is populated because nobody wrote a rule about absence, name fields are short and plain because that is what the author typed, and values are spread evenly because an even spread was the easiest thing to write. The suite goes green on data that resembles nothing a real user will send. That is how a rule-built dataset can pass for two years and still miss a defect that appears on the first day of real traffic. A fitted generator fails by **being unaccountable**. Nothing states what the output must satisfy, so a constraint the product silently depends on can be violated in a small fraction of rows, and a test run then fails intermittently with no readable cause. Because a fit smooths, the unusual combinations that actually break code are the first thing it loses: the generator learns the middle of the data very well and the edges hardly at all. And a refit on a newer extract changes the produced data without changing a single reviewable line anywhere. ## Using them together In practice the two compose, and the composition is what a strong candidate reaches for: 1. **Fit for shape.** Draw rows from a generator fitted to a real extract so that volume, skew and field lengths resemble production. 2. **Assert with rules.** Run the authored rule set over the produced rows as an acceptance filter rather than as a producer: reject or repair any row that violates an invariant the product depends on. 3. **Record what you checked.** Keep a summary of the produced dataset — per-field ranges, absence rates, category frequencies, counts per state — so the next refit can be compared with the last one instead of accepted on trust. That layering keeps what each family is genuinely good at. The fit supplies realism nobody could have specified; the rules supply the invariants the product depends on, in a form a person can read. ## If you can only have one Prefer authored rules when the dataset's job is to verify stated behaviour deterministically, when the invariants matter more than the shape, or when nobody may lawfully touch an extract of real records. Prefer a fitted generator when the dataset's job is to make the system behave as it does under real traffic — realistic volume, skewed value frequencies, plausible text lengths — and when someone will own the extract, the refits and the privacy question that arrives with them. The framing that survives an interview is short: **authored rules guarantee intent, a fitted generator reproduces reality, and neither guarantees the other.**

  • Where do authored rules most often fail in practice?
    They encode what the author believed rather than what arrives. Optional fields are present far more often than in real traffic, text is short and plain, values are spread evenly, and states the product rarely reaches never appear. The data satisfies every stated rule and still fails to exercise the product.
  • What can a fitted generator produce that a rule author would never have written?
    Joint behaviour: which fields move together, how skewed the value frequencies are, how long the longest text really gets, and what proportion of records sit in unusual states. Those are the properties nobody writes down and often the ones that break code — though when one of them fails a test, there is no rule to point at, so reproduction takes longer.

Authored rules are a recipe; a generator fitted to real records is a photograph of the finished dish. One tells you what must be true, the other only what it looked like.

saying these in an interview costs you the question

  • Says a fitted generator needs no constraints because the output looks realistic
  • Assumes an authored rule set produces realistic-looking data by default
  • Treats the two approaches as interchangeable names for one technique
  • Cannot name anything an authored rule set fails to express
  • Believes wide ranges in a rule set cover rare field combinations