skip to content

How do you constrain Kotest's built-in generators — integers, strings and collections — to a specific range, size or character set, and what goes wrong if you leave the defaults in place?

level: seniorimportance: should knowfreq 30%

answer

  1. domain in the generator, not guards in the body
  2. Arb.int(range), Arb.list(gen, sizeRange)
  3. Arb.string(minSize, maxSize, codepoints)
  4. edge cases move to your boundaries
  5. over-constraining = deleting the test

basics

~10 s

Pass the domain to the generator: Arb.int(1..100), Arb.string(minSize, maxSize, codepoints), Arb.list(elementArb, sizeRange). Unconstrained defaults generate values outside your contract — huge numbers, wide unicode, long collections — producing false failures and slow runs.

solid answer

~50 s

Kotest's built-in generators take their domain as parameters rather than expecting you to filter afterwards: - `Arb.int(1..100)`, `Arb.long(0..1_000)`, plus shorthand generators such as `Arb.positiveInt()` and `Arb.nonNegativeInt()`. - `Arb.string(minSize = 1, maxSize = 32, codepoints = Codepoint.alphanumeric())` — the character set matters as much as the length. - `Arb.list(Arb.int(), 1..10)` and the equivalent set/map forms, which bound collection size. Leaving defaults in place has three costs. **False failures**: an unconstrained `Arb.string()` will produce empty strings and unusual unicode, so a property fails on inputs your API never accepts. **Slow tests**: default collection sizes can be large, and if the property calls something expensive you pay for it a thousand times. **Bad edge coverage**: edge cases track the constraint, so a constrained generator tests *your* boundaries. The opposite failure exists too — over-constraining until the property passes removes the very inputs worth testing, so the constraint must reflect a documented contract.

code

kotlin · 6 lines
kotlin
checkAll(
    Arb.string(minSize = 1, maxSize = 32, codepoints = Codepoint.alphanumeric()),
    Arb.list(Arb.int(0..9), 1..10)
) { username, digits ->
    render(username, digits).length shouldBeLessThanOrEqual 64
}

go deeper

for a junior

Know that generators take their domain as parameters — a range for numbers, min/max size for strings, a size range for lists — instead of filtering afterwards.

for a middle

Explain the three consequences of constraining: edge cases move to your boundaries, no wasted discarded draws, and bounded runtime; and give the string character-set example.

for a senior

Bring judgement: constraints must map to documented contracts, over-constraining is equivalent to deleting the test, floating-point specials are a domain decision, and filtering is reserved for structural conditions.

for a principal

Treat generators as the executable specification of the system's input domain, and set a review standard that every constraint is traceable to a contract rather than to a red build.

## Constraints belong in the generator The most common mistake in a first property test is generating a wide domain and then coping with it in the body — filtering, early-returning, or wrapping half the assertion in an `if`. Kotest's built-in generators take their domain as parameters precisely so you do not have to: ```kotlin Arb.int(1..100) // bounded integers Arb.long(0L..1_000_000L) Arb.positiveInt() // shorthand for a common domain Arb.nonNegativeInt() Arb.string(minSize = 1, maxSize = 32, codepoints = Codepoint.alphanumeric()) Arb.list(Arb.int(0..9), 1..10) // element generator + size range ``` This matters for more than tidiness. Three mechanisms follow the constraint. **Edge cases follow the constraint.** An `Arb` mixes declared boundary values into its random stream. When you bound a generator, the boundaries that get injected are *your* boundaries — `1` and `100` for `Arb.int(1..100)`, minimum and maximum size for a bounded list. An unconstrained generator spends its boundary budget on `Int.MIN_VALUE` and the empty string, which may be irrelevant to your contract. **Every iteration stays useful.** Filtering values inside the property (or with a filtering combinator) discards draws: the run does the work of generating and rejecting, and if the acceptance rate is low the effective sample size collapses. A constrained generator produces only usable values. **Runtime is bounded.** Default collection size ranges are generous. If the system under test is expensive — a parser, a serializer, a database round-trip — a thousand iterations over long lists is how a property test becomes the slowest thing in the suite. Bounding size is the cheapest performance fix available. ## Strings: length is only half the problem `Arb.string()` varies both length and character content. The character content is configurable through a codepoint generator — `Codepoint.alphanumeric()`, `Codepoint.ascii()` and similar — and choosing it is a contract decision: - If the field is a username validated to `[A-Za-z0-9]`, generate alphanumeric codepoints. Anything else tests the validator, not the feature. - If the system is supposed to handle arbitrary user text, do **not** narrow the character set — wide unicode is exactly the input class that finds encoding, length-vs-codepoint and normalisation bugs. Many "flaky property" reports are really "the default string generator produced a codepoint the system never promised to accept" or, symmetrically, "the narrow generator hid the multi-byte bug we shipped". ## Numbers: floating point deserves special care Floating-point generators include the classically dangerous values, so a `Double` property can encounter `NaN` and the infinities. `NaN` compares unequal to itself, which breaks naive equality assertions and ordering invariants outright. If those values are out of contract, use a generator restricted to finite numeric doubles or an explicit range — do not litter the property body with guards, and do not "fix" it with a failure tolerance. ## The opposite mistake: over-constraining Narrowing a generator until a property passes is indistinguishable from deleting the test. The discipline is that **every constraint must correspond to a documented contract**: a validated field length, an API-enforced range, a schema limit. If you cannot point at the rule that excludes the input, the input is legal and the code is what needs to change. This is also why constraints are worth reviewing. `Arb.string(minSize = 1)` is a claim that empty input is impossible — if the HTTP layer can deliver an empty body, that claim is wrong and the property has just been told to stop looking at the most likely bug. ## Filter versus constrain Kotest does offer filtering on generators, and it has a legitimate place for conditions that cannot be expressed as a range or size ("a list with no duplicates"). But it has a discard cost: rejected values consume attempts, and a highly selective filter can starve the run. The rule of thumb is: express what you can as a range or size directly on the generator, and reserve filtering for genuinely structural conditions — ideally on a domain already narrowed so the acceptance rate is high. ## A checklist for reviewing a generator 1. Does the numeric range match the validated/documented range? 2. Is the collection size bounded in a way that keeps the property fast? 3. Is the string character set the one the system actually promises to accept — and if the system accepts anything, is the generator wide enough to find encoding bugs? 4. Are floating-point specials in or out of contract, and is that stated in the generator rather than in a guard? 5. Can every constraint be justified by pointing at a rule, or was one added to make a red test green?

  • Why prefer `Arb.int(1..100)` over `Arb.int()` with a filter that keeps 1..100?
    The constrained generator produces only usable values, so every iteration counts, and its injected edge cases become 1 and 100 — the boundaries that matter to your contract. A filter over the full `Int` range rejects effectively all draws, wasting attempts and risking a starved run, while still spending its boundary budget on irrelevant values.
  • When should you deliberately keep an unconstrained `Arb.string()`?
    When the system genuinely promises to accept arbitrary text — a comment body, a search query, a stored note. There the wide character set and the empty string are exactly the inputs that expose encoding, byte-versus-codepoint length and normalisation bugs. Narrowing the generator would hide the class of defect the property exists to find.
  • A property over doubles fails intermittently on an equality assertion. What do you check first?
    Whether the generator can emit `NaN` or an infinity. `NaN` is unequal to itself, so any equality or ordering assertion involving it fails. If those values are out of contract for the system, switch to a finite/ranged double generator; if they are in contract, the code needs to define behaviour for them and the assertion needs to reflect it.

saying these in an interview costs you the question

  • Filtering unwanted values inside the property body instead of constraining the generator
  • Narrowing a generator until the property passes, with no contract that justifies it
  • Forgetting that string generators vary character set as well as length
  • Ignoring collection size defaults and then blaming property testing for slow suites
  • Treating `NaN`/infinity failures as framework flakiness rather than a domain decision

context