skip to content

How do you decide how tightly to constrain the generators in a property suite — matching the production input distribution versus maximising boundary coverage — and how do you keep those decisions reviewable as a codebase grows?

level: principalimportance: nice to knowfreq 15%

answer

  1. generate the contract, not the traffic
  2. every constraint traceable to a rule
  3. too narrow is the dangerous failure
  4. named generators beside domain types
  5. enumerate finite, bound expensive

basics

~20 s

Constrain to the documented contract, not to production frequency. Generators should span every input the system promises to accept — including boundaries and absent values — and exclude only what a stated rule forbids. Share them as named, reviewed generators per domain type.

solid answer

~60 s

My rule is that a generator encodes the **contract**, not the traffic. Production distribution is the wrong target: real traffic is dominated by the easy middle of the domain, and the whole reason to generate inputs is to reach the parts traffic does not. So the domain of each generator should be exactly what the system says it accepts — the validated range, the documented size limits, the character set, and whether absence is legal — and every constraint should be traceable to that rule. A constraint added because a property went red is a defect in review. Because an `Arb` mixes declared boundary values into its stream, expressing the true bounds is also what gets you boundary coverage of the *right* boundaries. Wider is not automatically better: an unbounded generator tests inputs the API rejects at the edge, and unbounded collection sizes make properties the slowest thing in the suite. Operationally I keep named generators next to the domain types they describe, so they are reviewed like any contract, reused across properties, and updated in the same change as a validation rule.

go deeper

for a junior

Recall the principle: the generator should produce the inputs the system says it accepts, including boundaries and empty/absent values.

for a middle

Contrast the two failure modes — too wide producing false failures, too narrow hiding real ones — and show how bounds, sizes, character sets and optionality come from the contract.

for a senior

Add the operational practice: named reusable generators, updating them with validation changes, enumerating finite dimensions, and bounding the expensive ones for runtime.

for a principal

Argue generators as executable specification with a review standard (every constraint cites a rule), the organisational drift problem as the system grows, and the narrow exception where input distribution genuinely is part of the claim.

## The wrong question: what does production look like? It is tempting to build generators that mirror observed traffic — realistic name lengths, typical order sizes, the currencies that actually appear. That instinct is misplaced for property testing. Production traffic is concentrated in the easy middle of the input domain; that region is already covered by every integration test and every day of real usage. The value of a generator is that it reaches the parts of the domain that traffic reaches once a quarter, at 3am, from one misbehaving client. Distribution-matching also cannot be delivered by the tool: `Arb` deliberately injects boundary values into its stream, so the value sequence is intentionally biased and is not a fair sample of anything. Trying to make it distributional fights the design. ## The right question: what does the contract accept? The useful target is the **contract**: the set of inputs the system promises to handle. For each generated field, three decisions fall out of it. 1. **Bounds.** What range is validated or documented? Generate exactly that, so the injected edge cases land on your real boundaries rather than on the type's. 2. **Shape and size.** What size limits does the schema, the API or the storage layer impose? Bound collections and strings there. This is simultaneously the correctness decision and the runtime decision. 3. **Optionality.** Can this value legitimately be absent? If so it must be generated as absent sometimes; if not, generating it as absent tests a case the system was never asked to handle. A generator built this way reads as an executable sentence: *"a name of 1 to 20 alphanumeric characters, a quantity from 1 to 99, an optional note, and one of the supported currencies"*. That sentence is reviewable by someone who knows the domain but not the framework, which is the property you want. ## Two failure modes, symmetric and both common **Too wide.** The generator produces inputs the contract excludes — an empty identifier, a 10,000-element batch, control characters in a code field. Properties fail on inputs the system rightly rejects, and the team learns to distrust property tests. The correct response is a constraint that points at the rule, not a failure tolerance and not a guard in the test body. **Too narrow.** The generator has been trimmed until everything passes. This is the more dangerous one because the suite is green: the property now asserts something about a sanitised sliver of the domain, and the interesting inputs — empty, boundary, absent, unusual unicode — have been legislated away. The tell is a constraint nobody can justify by pointing at a validation rule. Hence the review standard: **every constraint must be traceable to a stated rule**, and "it made the test pass" is not one. Conversely, the absence of a constraint should be deliberate too: an unbounded string in a field that accepts arbitrary user text is a *decision*, and a good one, because that is where encoding and length bugs live. ## Keeping the decisions reviewable at scale Ad-hoc inline generators are fine for one property and corrosive across a hundred. As the suite grows: - **Name generators after domain types** and keep them beside those types, so `orderArb`, `usernameArb`, `isoCurrencyExhaustive` are single definitions the team reuses. A property then reads as a claim about the domain, not a wall of generator plumbing. - **Change the generator in the same commit as the validation rule.** When a field's max length changes, the generator is part of the contract change. If generators live near their types, that is a natural diff; if they are scattered inline, it never happens and the suite silently drifts out of contract. - **Enumerate the small categorical dimensions.** Anything genuinely finite — statuses, supported verbs, currencies — should be exhaustively enumerated rather than sampled, so adding a constant automatically extends coverage rather than requiring someone to remember. - **Bound the expensive dimensions explicitly.** State the size limits that keep the property fast, and treat a property that dominates suite runtime as a generator problem first, an iteration-count problem second. - **Review generators like API definitions.** They are the executable statement of what the system claims to accept; a wrong generator is a wrong specification, not a test-hygiene nit. ## Where distribution *does* matter One honest exception: when the property is about performance or about a probabilistic guarantee, the shape of the input distribution is part of the claim, and a generator biased toward boundaries will mislead. Those cases need a purpose-built harness rather than the property generators, and it is worth saying so explicitly rather than quietly reusing a generator whose bias undermines the measurement. ## The one-sentence version Generate the contract, enumerate what is finite, bound what is expensive, model absence explicitly, and make every constraint point at a rule — then keep the generators next to the types they describe so those decisions stay visible as the system changes.

  • Why not build generators that match observed production input distributions?
    Because the middle of the distribution is the part already covered by real usage and integration tests, while the value of generation lies in reaching rare regions. Kotest's `Arb` also injects boundary values deliberately, so its stream is intentionally biased and cannot be treated as a faithful sample anyway. Distribution-shaped generation belongs in a purpose-built performance or statistical harness.
  • How do you catch a generator that has drifted out of line with the code's validation rules?
    Keep generators beside the domain types they describe so that a change to a validation rule touches the generator in the same diff, and hold the review standard that every constraint cites the rule that justifies it. Constraints with no justification, and nullable fields never generated as null, are the two drift signals worth scanning for.
  • A property is the slowest test in the suite. Where do you look first?
    At the generator's size and shape parameters before the iteration count. Unbounded collection and string sizes multiply the cost of every iteration, so bounding them to the documented limits usually recovers most of the runtime while keeping coverage honest. Cutting iterations is the second lever, and it has a real cost in exploration.

saying these in an interview costs you the question

  • Tuning generators toward production frequency instead of the accepted contract
  • Narrowing a generator until the suite is green, with no rule to cite
  • Treating an unconstrained generator as automatically more thorough
  • Leaving nullable domain fields never generated as absent
  • Duplicating ad-hoc inline generators instead of naming and sharing them per domain type

context