Teams share one manufactured test dataset with conflicting needs - how do you pick between authored rules and a fitted generator?
answer
- Ask what the shared dataset is for
- Different teams want different guarantees
- Try layering before allowing a fork
- Two producers means two of everything
basics
~20 sDecide from the dataset's job, not the technique: authored rules when stated invariants and explainability dominate, a generator fitted to real records when production-like shape does. Layer them before forking, and run two producers only with a named owner.
solid answer
~50 sStart from what the shared dataset is for. If most teams verify stated behaviour and need every failure explainable, one authored rule set serves them: any team can add a constraint, the diff shows what changed, and no protected-data question arises. If most teams need the system to behave as it does under real traffic — realistic volume, skew, value lengths — a generator fitted to an extract earns its cost, provided someone owns the extract, the refits and the privacy assessment. Where the needs genuinely conflict, **layer rather than fork**: fit for shape, then apply each team's invariants as an acceptance filter over the produced rows. Running two independent producers is the expensive answer — two review paths, two protected-data scopes, and a question at every failure about which producer made the row. Take it only with a named owner and a stated condition for collapsing back to one.
code
yaml · 13 linesshared_dataset:
producer: fitted_generator
fitted_to: quarterly_extract_of_real_records
acceptance_rules: union_of_team_rules # any team may add a constraint
on_violation: repair_or_resample
owner: platform_data_team
refit_cadence: quarterly
refit_review: compare_summary_with_previous
privacy_scope: derived_from_protected
fork_policy: named_owner_and_stated_reason_required
exit_condition: |
if extract access is withdrawn, fall back to authored rules only
and accept the loss of realistic value skewgo deeper
Know that a team can either write the rules its test data must satisfy or fit a generator to a sample of real records, and that the choice has consequences well beyond convenience.
Explain the trade: stated invariants and explainability on one side, production-like shape on the other. Be able to name what each approach costs to set up and what it costs to keep running.
Argue from the dataset's job rather than the technique, and describe layering — fit for shape, then assert each team's invariants over the produced rows — as the alternative to forking the dataset.
Own the decision and its exit: the named owner, the evidence required before a change lands, the protected-data scope of the produced rows, when a second producer is justified, and what happens when a team forks anyway.
## Start from the dataset's job, not the technique The argument usually arrives as a technique preference — one group wants rules they can read, another wants data that behaves like real traffic — and it is unresolvable in that form because both are right about their own work. The question that resolves it is: *what is this shared dataset for?* A dataset whose job is to verify stated behaviour deterministically wants authored rules. A dataset whose job is to make the system behave as it does under real traffic wants a fit. A dataset asked to do both is being asked for two things, and the honest options are to layer them, to split them, or to say which one loses. ## Conditions that pick each | Condition | Points to authored rules | Points to a fitted generator | |---|---|---| | Primary use of the data | Verifying stated behaviour | Making the system behave realistically | | What a failure must yield | An explanation someone can read | A realistic reproduction | | Access to an extract of real records | None permitted, or none available | Available and lawful to fit against | | Standing ownership | Nobody can own a pipeline | A team will own extract, refit and assessment | | What the teams disagree about | Which invariants must hold | How closely the shape must match | | Audience for the result | Includes people outside the team | Internal engineering only | | Value distribution matters | Not really | Decisively — skew changes behaviour | Two conditions dominate in practice. **Lawful access to real records** is binary: without it the fitted option does not exist, however much anyone wants realism. **Standing ownership** is the one teams underestimate: a fit is not a one-off task but a recurring obligation — refits, comparisons against the previous output, and a privacy assessment that expires each time the source changes. An unowned fitting pipeline degrades into a dataset nobody can explain and nobody dares change. ## The layered option, which is usually the answer Before accepting the split, try composing them. Fit a generator for shape, then run each team's invariants as an acceptance filter over the produced rows: reject or repair rows that violate a constraint some team depends on. Every team keeps what it needed — realistic shape from the fit, its own guarantees from the rules — and the constraints stay readable, which restores most of what the fit gave away. This is also how a conflict gets tested honestly. When two teams claim incompatible needs, express both as acceptance rules and see whether they actually collide. Most do not: one team wants a value present, another wants it varied, and both are satisfiable in one dataset. A genuine collision — one team requires a state another requires never to occur — is a real signal for a second dataset. An unresolved preference is not. ## What running both actually costs If two producers really are justified, price the decision honestly rather than discovering the cost later: 1. **Two review paths.** Rule diffs on one side, output-summary comparisons on the other, and reviewers who have to understand both. 2. **Two protected-data scopes.** The fitted branch drags the source extract's handling rules along with it; keeping the authored branch clean requires them to stay genuinely separate, and any pipeline that mixes rows contaminates the simpler one. 3. **A provenance question at every failure.** "Which producer made this row?" is now a step in every investigation, and it is the step people forget. 4. **Alignment drift between the two datasets.** A constraint added on the authored side and not applied on the fitted side means the same test passes on one dataset and fails on the other, which erodes trust in both. 5. **Two things to keep alive.** Both need an owner. In practice one becomes the real dataset and the other rots quietly while tests still depend on it. ## The decision a lead actually owns The technique is the smaller half of the decision. The parts that belong to a lead are: **who owns the shared dataset**, **what evidence is required before a change to it lands**, **what the protected-data scope of the produced rows is**, and **what the exit condition is** — the circumstance under which a second producer is retired, or under which the fitted branch falls back to authored rules because access to the extract was withdrawn. And there is the failure mode to plan against: if the shared dataset does not serve a team, that team quietly forks it, and the organisation ends up with several manufactured datasets, no owner and no shared invariants — the outcome the debate was supposed to prevent. Make forking a decision with a name attached rather than something that happens by default, and the technique argument mostly settles itself.
- One team says the shared dataset's invariants are wrong for it. What do you do before letting it fork?Express the team's requirement as an acceptance rule over the produced rows and see whether it genuinely collides with an existing one. Most claimed conflicts are satisfiable together. A real collision — one team needs a state another needs excluded — justifies a second dataset; a preference does not.
- What is the strongest argument for authored rules in an organisation that could afford a fitting pipeline?Explainability that survives leaving the team. When a failure has to be explained to someone outside engineering, a readable constraint that produced the row is evidence; "the generator sampled it" is not. The authored branch also stays entirely outside the protected-data question, which removes a recurring assessment.
- How would you know a year later that the choice was wrong?Watch for teams building private datasets beside the shared one, for defects that reach real traffic on value shapes the dataset never contained, and for refits nobody reviews. Any of those means the dataset stopped serving its job, and the review is of its job rather than of the technique.
saying these in an interview costs you the question
- Picks the technique before asking what the dataset is for
- Assumes a fitted generator is simply the more advanced option
- Runs two producers with no owner named for either
- Lets each team fork the dataset to end the argument
- Ignores that the fitted branch carries protected-data handling along