skip to content

Arb & Exhaustive Generators

Arb produces random values salted with edge cases while Exhaustive enumerates a full finite domain. Interviewers ask when each applies and which built-ins exist, because generator choice decides whether a property actually explores the input space.

on this pageshow

explore

questions

5

Kotest's `Arb` is usually described as producing random values, yet property runs frequently hit values like 0, `Int.MIN_VALUE`, empty strings and empty lists. What is actually in an Arb's output stream?

level: middleimportance: must knowfreq 42%

answer

  1. Arb = random sampler + declared edge cases
  2. 0, MIN/MAX, "", empty list injected early
  3. uniform random would never hit 0
  4. constraining moves edge cases to your bounds
  5. not a fair statistical sample

basics

~20 s

An Arb is not purely random: it declares a set of edge cases alongside its random sampler, and Kotest interleaves those edge cases into the value stream with a small, configurable probability, so boundary values appear far more often than chance would produce.

solid answer

~50 s

Kotest's `Arb` has two sources of values: a **random sampler** and a declared set of **edge cases** — the boundary values that break code disproportionately often. During a property run the engine injects edge cases into the stream with a small probability rather than leaving them to chance, so `Arb.int()` regularly yields `0`, `Int.MIN_VALUE` and `Int.MAX_VALUE`, `Arb.string()` yields the empty string, and `Arb.list()` yields the empty list — often within the first handful of iterations. This is why property tests find off-by-one and empty-input bugs so quickly: uniform random sampling over the `Int` range would essentially never produce `0`, so the framework does not rely on uniformity. Edge cases flow through derived generators too — mapping or filtering an `Arb` yields an `Arb` that still has an edge-case notion — and they participate in shrinking, which is why a minimised counter-example so often *is* a boundary value. The injection probability is part of the property-test configuration.

code

kotlin · 5 lines
kotlin
// Full range: edge cases include 0 and the Int bounds
checkAll(Arb.int()) { n -> abs(n) shouldBeGreaterThanOrEqual 0 }

// Constrained: edge cases move to 1 and 100
checkAll(Arb.int(1..100)) { n -> bucketOf(n) shouldNotBe null }

go deeper

for a junior

Recall that an Arb mixes deliberate boundary values (0, empty string, empty list, type bounds) into its random output, which is why property tests catch empty-input bugs fast.

for a middle

Explain the two-source model and the injection probability, give concrete edge cases per built-in generator, and note that constraining a generator relocates its edge cases.

for a senior

Add the consequences: no distributional guarantees, edge cases propagating through composed generators, floating-point edge cases like NaN, and constraining rather than filtering to keep iterations useful.

for a principal

Discuss it as domain modelling — generators encode the contract's real input space, and edge-case bias is the mechanism that makes exploratory testing worth its runtime budget.

## Two value sources, one generator In Kotest's property module, `Arb<A>` (short for *arbitrary*) is the generator type for effectively unbounded domains — integers, strings, lists, your own data classes. Its contract has two halves: 1. **A random sampler.** Given a source of randomness, produce a value of type `A`. 2. **Edge cases.** A declared set of values that are disproportionately likely to break code: zero, minimum and maximum bounds, empty collections, empty strings, and similar. A property run does not simply call the sampler N times. The engine decides, per value, whether to emit an edge case or a random sample, using a small injection probability. The practical effect is that boundary values show up early and often — usually within the first few iterations — instead of essentially never. ## Why this design exists Consider `Arb.int()` sampling uniformly over the full 32-bit range. The probability of drawing exactly `0` is about one in four billion; over 1000 iterations you would never see it. Yet `0` is the single most bug-prone integer in most codebases (division, modulo, empty-batch handling, index arithmetic). The same argument applies to the empty string, the empty list, `Int.MIN_VALUE` (whose negation overflows), and the boundaries of any constrained range. So Kotest treats "random" as a *strategy for coverage*, not a statistical guarantee. The generator is really "a curated set of dangerous values, plus randomness to explore everything else". This is the single most important thing to understand about `Arb`, because it explains both why property tests find boundary bugs immediately and why a property run is not a fair statistical sample of the domain — do not use property generators to estimate distributions. ## What the edge cases typically are Each built-in generator declares edge cases appropriate to its type. Broadly: - Integral generators include `0` and the boundaries of their range (for the unbounded form, the type's `MIN_VALUE` and `MAX_VALUE`). - String generators include the empty string, and short strings around the size boundaries. - Collection generators include the empty collection and collections at the boundaries of the configured size range. - Floating-point generators include the classically dangerous values, which is why `Arb.double()` can surface `NaN` and the infinities and why Kotest offers finite-only variants for when those are out of contract. - Nullable generators (see `orNull`) treat `null` as an edge case, so `null` appears early rather than by luck. When you **constrain** a generator — `Arb.int(1..100)`, `Arb.list(Arb.int(), 1..5)` — the edge cases move with the constraint: the boundaries of *your* range become the interesting values. That is a strong argument for expressing the real domain in the generator rather than filtering inside the property body: a constrained generator gets boundary coverage of the domain you actually care about. ## Composition Edge cases survive composition. Mapping an `Arb` produces an `Arb` whose edge cases are the mapped edge cases; combining generators produces combinations that include the constituent edge cases. This is why building a data-class generator from well-chosen field generators yields a generator that automatically tries "empty name", "zero quantity" and "null note" without you enumerating those cases by hand — the composition surface itself is a separate topic, but the edge-case propagation is the reason it pays off. ## Interaction with reproducibility and shrinking Because edge-case injection is probabilistic, two runs of the same property do not necessarily test the same values — which is precisely why property failures print a seed and why replaying is part of the workflow. And because edge cases are the values most likely to break code, minimised counter-examples very frequently *are* edge cases; when a failure report says the shrunk input was `0` or `""`, that is the machinery working as designed. ## Consequences for how you write properties - **Expect the empty case.** If your property will not hold for empty input, that is information: either handle empty input in the code or constrain the generator (`minSize = 1`) because empty is out of contract. - **Do not assume uniformity.** A property that reasons about how often something occurs is misusing the generator. - **Constrain, don't filter, where you can.** A filter that discards most values wastes iterations and can starve the run; a constrained generator keeps every draw useful and relocates the edge cases to your boundaries. - **Beware silent unbounded defaults.** `Arb.string()` and `Arb.list()` come with default size ranges that may be far wider than your system expects; if that produces failures, decide deliberately whether the contract or the generator is wrong. The injection probability itself is part of the property-test configuration surface, so a suite can tune how aggressively edge cases are mixed in — but the default behaviour is the one to internalise: `Arb` is random *plus* a deliberate boundary bias.

  • If `Arb.int(1..100)` is used, are `Int.MIN_VALUE` and `0` still generated?
    No — constraining the generator constrains its whole value space, edge cases included. The interesting values become the boundaries of the constraint, so you would expect `1` and `100` to be favoured. That is exactly why expressing the true domain in the generator is better than filtering: the edge-case machinery then targets the boundaries you care about.
  • Your property over `Arb.double()` fails with a strange comparison result. What is the likely cause?
    Floating-point generators include the classically dangerous values, so a run can produce `NaN` or an infinity. `NaN` compares unequal to itself, which breaks naive equality and ordering assertions. If those values are out of contract for the system under test, use a generator restricted to finite numeric doubles or constrain the range explicitly, rather than special-casing inside the property.
  • Can you rely on `Arb` for a statistical test, for example checking that a sampler is uniform?
    No. Edge cases are injected deliberately, so the value stream is intentionally biased toward boundaries and is not a fair sample of the domain. Property generators are a coverage strategy, not a distribution. A statistical claim needs a purpose-built harness that draws from a known distribution you control.

saying these in an interview costs you the question

  • Describing `Arb` as purely uniform random generation
  • Being surprised that empty strings or empty lists appear and calling them "unrealistic inputs"
  • Assuming edge cases only appear if you enumerate them yourself
  • Using property generators to reason about the distribution of generated values
  • Filtering unwanted values inside the property body instead of constraining the generator

context

open as a page

In Kotest's property module, what does an `Exhaustive` generator such as `Exhaustive.enum<T>()` or `Exhaustive.of(...)` guarantee about the values a test sees, and what happens when the iteration count is larger or smaller than the domain?

level: middleimportance: should knowfreq 34%

basics

~20 s

An Exhaustive holds a finite list of values and the run walks it in order, cycling when there are more iterations than values. Every value is covered only if iterations are at least the domain size; fewer iterations means some values are never tested.

open as a page

How do you constrain Kotest's built-in generators — integers, strings and collections — to a specific range, size or character set, and what goes wrong if you leave the defaults in place?

level: seniorimportance: should knowfreq 30%

basics

~10 s

Pass the domain to the generator: Arb.int(1..100), Arb.string(minSize, maxSize, codepoints), Arb.list(elementArb, sizeRange). Unconstrained defaults generate values outside your contract — huge numbers, wide unicode, long collections — producing false failures and slow runs.

open as a page

What does Kotest's `Arb<A>.orNull()` produce, and how would you use it to exercise a nullable API contract without swamping the run with nulls?

level: middleimportance: nice to knowfreq 22%

basics

~10 s

orNull turns an Arb<A> into an Arb<A?> that emits null with a small probability alongside the underlying values. The nullProbability parameter tunes that rate, so you can make nulls rare, frequent, or effectively absent.

open as a page

How do you decide how tightly to constrain the generators in a property suite — matching the production input distribution versus maximising boundary coverage — and how do you keep those decisions reviewable as a codebase grows?

level: principalimportance: nice to knowfreq 15%

basics

~20 s

Constrain to the documented contract, not to production frequency. Generators should span every input the system promises to accept — including boundaries and absent values — and exclude only what a stated rule forbids. Share them as named, reviewed generators per domain type.

open as a page