Kotest's `Arb` is usually described as producing random values, yet property runs frequently hit values like 0, `Int.MIN_VALUE`, empty strings and empty lists. What is actually in an Arb's output stream?
answer
- Arb = random sampler + declared edge cases
- 0, MIN/MAX, "", empty list injected early
- uniform random would never hit 0
- constraining moves edge cases to your bounds
- not a fair statistical sample
basics
~20 sAn Arb is not purely random: it declares a set of edge cases alongside its random sampler, and Kotest interleaves those edge cases into the value stream with a small, configurable probability, so boundary values appear far more often than chance would produce.
solid answer
~50 sKotest's `Arb` has two sources of values: a **random sampler** and a declared set of **edge cases** — the boundary values that break code disproportionately often. During a property run the engine injects edge cases into the stream with a small probability rather than leaving them to chance, so `Arb.int()` regularly yields `0`, `Int.MIN_VALUE` and `Int.MAX_VALUE`, `Arb.string()` yields the empty string, and `Arb.list()` yields the empty list — often within the first handful of iterations. This is why property tests find off-by-one and empty-input bugs so quickly: uniform random sampling over the `Int` range would essentially never produce `0`, so the framework does not rely on uniformity. Edge cases flow through derived generators too — mapping or filtering an `Arb` yields an `Arb` that still has an edge-case notion — and they participate in shrinking, which is why a minimised counter-example so often *is* a boundary value. The injection probability is part of the property-test configuration.
code
kotlin · 5 lines// Full range: edge cases include 0 and the Int bounds
checkAll(Arb.int()) { n -> abs(n) shouldBeGreaterThanOrEqual 0 }
// Constrained: edge cases move to 1 and 100
checkAll(Arb.int(1..100)) { n -> bucketOf(n) shouldNotBe null }go deeper
Recall that an Arb mixes deliberate boundary values (0, empty string, empty list, type bounds) into its random output, which is why property tests catch empty-input bugs fast.
Explain the two-source model and the injection probability, give concrete edge cases per built-in generator, and note that constraining a generator relocates its edge cases.
Add the consequences: no distributional guarantees, edge cases propagating through composed generators, floating-point edge cases like NaN, and constraining rather than filtering to keep iterations useful.
Discuss it as domain modelling — generators encode the contract's real input space, and edge-case bias is the mechanism that makes exploratory testing worth its runtime budget.
## Two value sources, one generator In Kotest's property module, `Arb<A>` (short for *arbitrary*) is the generator type for effectively unbounded domains — integers, strings, lists, your own data classes. Its contract has two halves: 1. **A random sampler.** Given a source of randomness, produce a value of type `A`. 2. **Edge cases.** A declared set of values that are disproportionately likely to break code: zero, minimum and maximum bounds, empty collections, empty strings, and similar. A property run does not simply call the sampler N times. The engine decides, per value, whether to emit an edge case or a random sample, using a small injection probability. The practical effect is that boundary values show up early and often — usually within the first few iterations — instead of essentially never. ## Why this design exists Consider `Arb.int()` sampling uniformly over the full 32-bit range. The probability of drawing exactly `0` is about one in four billion; over 1000 iterations you would never see it. Yet `0` is the single most bug-prone integer in most codebases (division, modulo, empty-batch handling, index arithmetic). The same argument applies to the empty string, the empty list, `Int.MIN_VALUE` (whose negation overflows), and the boundaries of any constrained range. So Kotest treats "random" as a *strategy for coverage*, not a statistical guarantee. The generator is really "a curated set of dangerous values, plus randomness to explore everything else". This is the single most important thing to understand about `Arb`, because it explains both why property tests find boundary bugs immediately and why a property run is not a fair statistical sample of the domain — do not use property generators to estimate distributions. ## What the edge cases typically are Each built-in generator declares edge cases appropriate to its type. Broadly: - Integral generators include `0` and the boundaries of their range (for the unbounded form, the type's `MIN_VALUE` and `MAX_VALUE`). - String generators include the empty string, and short strings around the size boundaries. - Collection generators include the empty collection and collections at the boundaries of the configured size range. - Floating-point generators include the classically dangerous values, which is why `Arb.double()` can surface `NaN` and the infinities and why Kotest offers finite-only variants for when those are out of contract. - Nullable generators (see `orNull`) treat `null` as an edge case, so `null` appears early rather than by luck. When you **constrain** a generator — `Arb.int(1..100)`, `Arb.list(Arb.int(), 1..5)` — the edge cases move with the constraint: the boundaries of *your* range become the interesting values. That is a strong argument for expressing the real domain in the generator rather than filtering inside the property body: a constrained generator gets boundary coverage of the domain you actually care about. ## Composition Edge cases survive composition. Mapping an `Arb` produces an `Arb` whose edge cases are the mapped edge cases; combining generators produces combinations that include the constituent edge cases. This is why building a data-class generator from well-chosen field generators yields a generator that automatically tries "empty name", "zero quantity" and "null note" without you enumerating those cases by hand — the composition surface itself is a separate topic, but the edge-case propagation is the reason it pays off. ## Interaction with reproducibility and shrinking Because edge-case injection is probabilistic, two runs of the same property do not necessarily test the same values — which is precisely why property failures print a seed and why replaying is part of the workflow. And because edge cases are the values most likely to break code, minimised counter-examples very frequently *are* edge cases; when a failure report says the shrunk input was `0` or `""`, that is the machinery working as designed. ## Consequences for how you write properties - **Expect the empty case.** If your property will not hold for empty input, that is information: either handle empty input in the code or constrain the generator (`minSize = 1`) because empty is out of contract. - **Do not assume uniformity.** A property that reasons about how often something occurs is misusing the generator. - **Constrain, don't filter, where you can.** A filter that discards most values wastes iterations and can starve the run; a constrained generator keeps every draw useful and relocates the edge cases to your boundaries. - **Beware silent unbounded defaults.** `Arb.string()` and `Arb.list()` come with default size ranges that may be far wider than your system expects; if that produces failures, decide deliberately whether the contract or the generator is wrong. The injection probability itself is part of the property-test configuration surface, so a suite can tune how aggressively edge cases are mixed in — but the default behaviour is the one to internalise: `Arb` is random *plus* a deliberate boundary bias.
- If `Arb.int(1..100)` is used, are `Int.MIN_VALUE` and `0` still generated?No — constraining the generator constrains its whole value space, edge cases included. The interesting values become the boundaries of the constraint, so you would expect `1` and `100` to be favoured. That is exactly why expressing the true domain in the generator is better than filtering: the edge-case machinery then targets the boundaries you care about.
- Your property over `Arb.double()` fails with a strange comparison result. What is the likely cause?Floating-point generators include the classically dangerous values, so a run can produce `NaN` or an infinity. `NaN` compares unequal to itself, which breaks naive equality and ordering assertions. If those values are out of contract for the system under test, use a generator restricted to finite numeric doubles or constrain the range explicitly, rather than special-casing inside the property.
- Can you rely on `Arb` for a statistical test, for example checking that a sampler is uniform?No. Edge cases are injected deliberately, so the value stream is intentionally biased toward boundaries and is not a fair sample of the domain. Property generators are a coverage strategy, not a distribution. A statistical claim needs a purpose-built harness that draws from a known distribution you control.
saying these in an interview costs you the question
- Describing `Arb` as purely uniform random generation
- Being surprised that empty strings or empty lists appear and calling them "unrealistic inputs"
- Assuming edge cases only appear if you enumerate them yourself
- Using property generators to reason about the distribution of generated values
- Filtering unwanted values inside the property body instead of constraining the generator