What can a reviewer check in an authored test-data rule set that a generator fitted to real records cannot show?
answer
- Ask what a reviewer can actually read
- Rules are text; a fit is parameters
- A rule change shows in the diff
- Nothing readable means review moves elsewhere
basics
~20 sIntent. A rule set states in readable text what the data must satisfy and what was deliberately left unconstrained, and a diff shows what a change altered. A fitted generator exposes only its output, so review moves there.
solid answer
~50 sAn authored rule set can be read line by line. A reviewer sees which constraints are asserted, which fields are deliberately unconstrained and, when each non-obvious rule records why it exists, can trace a bound to what motivated it. When it changes, the diff names exactly what changed, and the effect can be reasoned about before any data is produced. A generator fitted to a real extract offers none of that. Its behaviour lives in fitted parameters that carry no stated intent, so there is nothing to read, and a refit on a newer extract quietly changes the produced data with nothing visible to review. The response is to move review to the output: assert explicit checks over the produced dataset and record a summary — per-field ranges, absence rates, category frequencies — so each refit is compared with the previous one rather than accepted on trust.
code
pseudocode · 17 lines# review moves from the producer to the produced dataset
checks over produced_dataset:
require every row: age in 18..120
require every row: closed_at is null or closed_at >= opened_at
require absence_rate(middle_name) between 0.10 and 0.60
require distinct(country) >= 12
summary = {
rows: count(produced_dataset),
per_field_range: ranges(produced_dataset),
absence_rates: absence(produced_dataset),
category_freq: top_frequencies(produced_dataset, n = 20)
}
record summary as dataset_summary(fit_id = current_fit)
# the reviewable event is this comparison, not the refit itself
report diff(summary, dataset_summary(fit_id = previous_fit))go deeper
Know that the rules for manufactured test data can be a file in the repository, read and changed like any other code, and that reading them tells you what the data is allowed to look like.
Explain what a reviewer actually does with a rule set: read each constraint, notice which fields are deliberately unconstrained, and read a diff when a bound moves. Then explain why a fitted generator offers none of those.
Demonstrate the move to output review. Describe the checks you would assert over produced rows and the summary you would record, so a refit is compared with the previous one instead of accepted because it ran.
Set the standard others follow: what must be reviewable before a produced dataset is trusted across teams, who approves a refit, and what evidence is retained when someone outside the team asks what the dataset represents.
## What "reviewable" means for manufactured test data When many teams draw from one manufactured dataset, somebody has to be able to answer *why is the data like this?* — usually a reviewer approving a change, sometimes an engineer explaining a failing test run, occasionally someone from outside the team asking what the dataset represents. Review has only two possible subjects: **the producer** or **the produced data**. Which of the two is available is the whole difference between the authored and the fitted approach. ## What an authored rule set exposes A rule set is text in the repository, so all of the following are directly inspectable: - **The constraints themselves.** A reader sees the allowed ranges, the value sets and the cross-field invariants, and can decide whether each is right. - **The deliberate silences.** Fields with no rule are visible as fields with no rule. That is real information: it tells a reviewer that absence, length and casing of those fields were never considered. - **The reason for a bound**, when the rule set is written to carry it. The difference between an age bound of 120 that came from a domain constraint and one that came from a guess is invisible unless somebody wrote it down. - **The change.** A diff shows precisely which constraint moved, and the consequence can be argued about before a single row is produced. - **The absence of a privacy question.** Nothing in the rules came from a real person, so review does not have to consider what the dataset discloses. ## What a fitted generator does not expose A generator fitted to an extract of real records has behaviour that lives in fitted parameters. Those parameters are numbers; they carry no statement of intent, and no reviewer can look at them and say whether they are right. Three consequences follow. First, **there is no diff worth reading**. A refit on a newer extract may change value frequencies, absence rates and the length of the longest text field, while the repository shows only that a fit was rerun. Second, **the one genuinely reviewable input is often the one you may not show**. The source extract explains everything about the generator's behaviour, and it is real customer data, so a wide review audience cannot see it. Third, **the failure explanation is weaker**. When a produced row breaks a test, an authored rule set answers "this rule produced it"; a fit answers "it was sampled", which is not an answer anyone can act on. | Question a reviewer asks | Authored rule set | Fitted generator | |---|---|---| | What must every row satisfy? | Read the rules | Nothing is stated | | What was deliberately ignored? | Visible as unconstrained fields | Unknowable | | Why is this bound what it is? | Recorded beside the rule, if written | No bound was chosen | | What did the last change do? | Read the diff | Compare outputs before and after | | Can a wide audience inspect the input? | Yes, it is invented | Usually not; it is real data | ## Moving review to the output Because the producer cannot be read, review shifts to the produced dataset, and it has to be made repeatable rather than done once by eye: 1. **Assert explicit checks over the produced rows** — the invariants the product depends on, plus bounds on absence rates and category counts, so that "the fit went strange" fails loudly instead of silently. 2. **Record a summary of each produced dataset**: row count, per-field ranges, absence rates, the most frequent categories, counts per state. This is the thing a reviewer actually reads. 3. **Compare each refit against the previous recorded summary**, and require an explanation for the differences. A jump in one field's absence rate is a reviewable event; a rerun fitting step is not. 4. **Keep the comparison with the change**, so the question "what did this refit do to the shared dataset?" has a durable answer months later. ## The honest limit on both sides Reviewability is a practice, not a property of the technique. An authored rule set nobody reads, with unexplained bounds and no reasons recorded, is barely more accountable than a fit — the reader sees a number and cannot tell a domain constraint from a typing accident. Conversely, a fitted generator wrapped in asserted checks and recorded summaries can be more trustworthy than a neglected rule set, because at least the produced data is being examined every time. The defensible position is therefore not "authored rules are auditable". It is: *authored rules make intent available for review, a fit does not, and if you choose the fit you owe the dataset a review path over its output instead.*
- A generator is refitted on a newer extract and the suite starts failing intermittently. How does that recorded summary help?It localises the change before anyone reads a failing test. Comparing the new summary with the previous one shows the field whose absence rate jumped, the category that became dominant or the range that widened. Without a recorded summary the only evidence is an unstable suite, and the investigation starts from nothing.
- Does an authored rule set really carry intent, or only constraints?Only if it is written to. A bound with no recorded reason is nearly as opaque as a fitted parameter: a reader sees 120 and cannot tell a genuine domain limit from a guess that has since become load-bearing. Intent survives when each non-obvious rule names what motivated it.
- Who should sign off a refit of a dataset shared across several teams?Whoever owns the dataset, on the evidence of the summary comparison rather than on the fact that the fitting step succeeded. The teams drawing from it need notice that shape changed, because their tests were calibrated against the previous shape.
saying these in an interview costs you the question
- Claims a fitted generator can be reviewed by reading its parameters
- Treats a green suite as proof the produced dataset is correct
- Says the reasons live elsewhere, so rules need no recorded motivation
- Assumes a refit is safe because no code changed
- Reviews the produced rows once and never checks again