Your team's Kotest property tests each declare their own generators inline, and the same domain objects are generated five slightly different ways across the suite. How would you organise generators for a real domain model, and what conventions would you set?
answer
- one named arb per domain type, composed upward
- test-fixtures source set = single source of truth
- legal by construction; filter is a smell
- purity → the printed seed still reproduces
- weight the mix; shrinkers only where counterexamples are read
basics
~20 sTreat generators as a shared fixture library: one named Arb per domain type in a test-fixtures source set, composed upward with Arb.bind, constrained so every sample is legal by construction, with per-test narrowing done by composition rather than by copying and filtering.
solid answer
~60 sMake generators a first-class, reviewed artefact: - **One named `Arb` per domain type**, defined once (a `Generators.kt` object or a test-fixtures module) and composed upward with `Arb.bind` — `moneyArb` into `lineArb` into `orderArb`. Duplicated inline generators drift; a shared one gets fixed once. - **Legal by construction.** Constrain the component generators so the type's invariants hold without filtering — bounded ranges, non-empty strings, derived dependent fields via `arbitrary { }`. Rejection sampling is the smell that a constraint is in the wrong place. - **Narrow by composition, not by copy.** A test needing only EU orders writes `orderArb.map { it.copy(region = EU) }` or takes a parameterised generator (`fun orderArb(region: Region = ...)`), rather than forking a private variant. - **Keep them pure.** No clocks, no `Random`, no I/O — otherwise the printed seed no longer reproduces failures. - **Watch the distribution.** Weighted `Arb.choose` where the real input mix matters; shrinkers supplied for aggregates that show up in counterexamples. And review generators like production code — a wrong generator makes a green suite meaningless.
code
kotlin · 14 linesobject Gen {
val money: Arb<Money> = Arb.long(0..1_000_000).map(::Money)
val line: Arb<OrderLine> = Arb.bind(Arb.string(3..12), Arb.int(1..10), money, ::OrderLine)
fun order(status: Arb<Status> = Arb.enum<Status>()): Arb<Order> = arbitrary {
val lines = Arb.list(line, 1..5).bind()
Order(
id = Arb.uuid().bind().toString(),
lines = lines,
total = Money(lines.sumOf { it.amount.cents }),
status = status.bind(),
)
}
}go deeper
Say that generators should be defined once per domain type and reused, not re-written inline in every spec.
Add composition (small arbs combined with Arb.bind), constraining rather than filtering, and narrowing at the call site.
Bring in purity and seed reproducibility, distribution weighting, and where shrinking investment pays off.
Frame it as specification governance: generators reviewed like production code, one source of truth in test fixtures, explicit valid-vs-hostile policies at boundaries, and distribution choices justified rather than defaulted.
## Why generator sprawl matters In a property suite the generator *is* the specification of the input space. When five specs each build an `Order` slightly differently, you have five undocumented, contradictory specifications, and the resulting green build tells you less than it appears to. Worse, the drift is invisible: nothing fails when one spec's generator forgets that quantities are positive — the tests just stop covering the interesting cases, or start reporting counterexamples that could never occur in production. ## The structure **One arb per type, composed upward.** Mirror the domain model: ```kotlin object Gen { val money = Arb.long(0..1_000_000).map(::Money) val sku = Arb.stringPattern-free() // build from primitives you control val line = Arb.bind(sku, Arb.int(1..10), money) { s, q, m -> OrderLine(s, q, m) } val order = arbitrary { val lines = Arb.list(line, 1..5).bind() Order(id = Arb.uuid().bind().toString(), lines = lines, total = lines.sumOf { it.amount }) } } ``` Each level is small, named, and reusable. `Arb.bind` handles independent fields; the `arbitrary { }` builder handles derived ones. Aggregates get built from the parts, exactly like the production model. **Where it lives.** A dedicated test-fixtures source set or module, so both unit and integration specs can use the same generators without one test module depending on another's test code. The key property is that there is exactly one definition to change when the domain changes. ## The conventions worth writing down ### 1. Legal by construction, not by filtering Every `filter` in a generator is a place where the constraint was expressed too late — it burns draws, skews the distribution and can stall when the acceptance rate is low. Push the constraint into the generator: `Arb.int(1..10)` rather than `Arb.int().filter { it in 1..10 }`; derive the dependent field in `arbitrary { }` rather than generating two and rejecting the mismatches. Reserve `filter` for cheap, rare exclusions. ### 2. Parameterise instead of forking When a test needs a narrower input space, the wrong move is a private copy of the generator. Two right moves: expose the generator as a function with defaults (`fun order(region: Arb<Region> = Gen.region): Arb<Order>`), or compose at the call site (`Gen.order.map { it.copy(status = PAID) }`, `Gen.order.filter { … }` only if rare). Both keep one source of truth for the shape. ### 3. Purity is non-negotiable Generators must draw only from Kotest's random source. A `Random.nextInt()`, a `Instant.now()` or a database lookup inside a generator destroys seed-based replay — the printed seed will not reproduce the failure — and can also fire repeatedly during shrinking, when the property body is re-run with candidate values. Make this an explicit review rule; it is the single most common cause of "the CI failure won't reproduce locally". ### 4. Model the real distribution Uniform is not realistic. If production traffic is 95% well-formed, weight the generator with `Arb.choose(19 to valid, 1 to malformed)` rather than a 50/50 `Arb.choice`, so the iteration budget goes where the logic is deep. Conversely, if a rare shape is where the bugs are, deliberately over-weight it and say so in a comment. Distribution choices are design decisions and deserve a sentence of rationale. ### 5. Invest in shrinking where counterexamples are read For the two or three aggregates that dominate your failure output, it is worth supplying a `Shrinker` or restructuring the composition so component shrinkers apply (favouring `map`/`bind` over hand assembly). For the long tail, coarse counterexamples are fine. This is a cost/benefit call, not a blanket rule. ### 6. Review generators like production code A subtly wrong generator produces a green, meaningless suite — the failure mode nobody notices. Put generators in the same review lane as the code they test: does the range match the domain, does the distribution match reality, is it pure, does it still shrink? ## Trade-offs to acknowledge - **Central library vs. local clarity.** A shared generator is one indirection away from the test that uses it; keep them small, named after the domain type, and grouped so the jump is cheap. - **Constrained vs. adversarial.** Constraining generators to the legal domain is right for domain logic, but for parsers and boundary code you *want* the illegal values. Keep both: a `valid` and a `hostile` generator per boundary type, weighted for the test at hand. - **Reflective shortcuts.** Reflective generation is fine for wire DTOs in codec round-trips; it is wrong for validated domain types, and mixing the two policies without saying which applies where is how illegal values sneak in. ## The signal in an interview This question has no single right answer; it checks whether you treat the generator as part of the specification, whether you have opinions about purity and distribution, and whether you can name the failure modes — drift, rejection sampling, non-reproducible seeds, vacuous green suites.
- A property failure reported in CI cannot be reproduced locally even with the printed seed. What do you check in the generators first?Impurity. Anything the generator reads outside Kotest's random source — kotlin.random.Random, the system clock, an environment variable, a database or file — is not pinned by the seed, so the same seed yields different input. Grep the generator library for those, replace them with arbs bound inside the generator, and make purity a review rule.
- When is a constrained, legal-by-construction generator the wrong default?At the boundaries — parsers, deserializers, validators, HTTP request handling — where illegal input is exactly the input under test. There you want a hostile generator producing empty strings, control characters, huge values and malformed structures, usually weighted alongside the valid one. Keep both generators per boundary type and pick by test intent rather than making one generator serve both jobs.
saying these in an interview costs you the question
- Treating generators as throwaway test scaffolding rather than part of the specification
- Copying a generator and tweaking it instead of parameterising the shared one
- Filtering at the top of a generator instead of constraining its components
- Allowing clocks, Random or I/O inside generators and then wondering why seeds do not replay
- Assuming a uniform distribution is a neutral choice rather than a design decision