Kotest's Arb type offers map, flatMap and filter. Explain what each does to a generator, and why filter is the one to be careful with.
answer
- map = total transform, shrinking survives
- flatMap = second value depends on first
- filter = rejection sampling → waste, skew, stalls
- construct valid, don't reject invalid
- filter { it > 0 } → Arb.positiveInt()
basics
~20 smap transforms each sampled value (Arb<A> to Arb<B>). flatMap feeds a sampled value into a function that picks the next generator, so it expresses dependent values. filter re-samples until a predicate passes — it wastes samples, can skew the distribution and can stall when matches are rare.
solid answer
~60 s- **`map`** applies a pure transformation to every sample: `Arb.int(0..100).map { it.toString() }`. Cheap, and shrinking still works because it shrinks the underlying int and re-maps. - **`flatMap`** takes the sampled value and returns the *next* generator, so the second value can depend on the first — the only way to say "generate a list, then an index inside it" or "a start date, then an end date after it". - **`filter`** keeps re-sampling until the predicate holds. It is the tempting one and usually the wrong one: every rejected sample is wasted work, the surviving distribution is no longer the one you configured, edge cases can be filtered away entirely, and if the predicate is rarely true the generator spends its budget spinning. Rule of thumb: **construct valid values rather than reject invalid ones.** `Arb.int().filter { it > 0 }` should be `Arb.positiveInt()`; `filter { it.isNotEmpty() }` should be `Arb.string(minSize = 1)`; a date pair should come from `flatMap`, not from generating two dates and filtering out the wrong order.
code
kotlin · 11 lines// map: shape change, no samples lost
val userIds: Arb<UserId> = Arb.long(1..Long.MAX_VALUE).map(::UserId)
// flatMap: the second value depends on the first
val listAndIndex: Arb<Pair<List<Int>, Int>> =
Arb.list(Arb.int(), 1..20).flatMap { xs ->
Arb.int(0..xs.lastIndex).map { i -> xs to i }
}
// filter: only for cheap, rare exclusions
val nonBlank = Arb.string(minSize = 1, maxSize = 10).filter { it.isNotBlank() }go deeper
Know the one-line role of each: map transforms, flatMap makes the next value depend on the previous one, filter drops samples that fail a predicate.
Explain filter's cost model — wasted draws, skewed distribution, possible stall — and give constructive replacements for the common cases.
Add the shrinking angle: map shrinks cleanly through the source, flatMap shrinks worse, and a filter can remove the injected edge cases that were most likely to find the bug.
Treat generators as a cost-bearing pipeline: put constraints as far upstream as possible, keep dependent generation narrow, and review filters in test code the way you review N+1 queries.
## The three combinators In kotest-property an `Arb<A>` is a generator. Three combinators reshape one: ### map — transform each sample ```kotlin val idArb: Arb<UserId> = Arb.long(1..Long.MAX_VALUE).map { UserId(it) } ``` `map` is total: every sample of the source becomes exactly one sample of the result. Nothing is discarded, so the iteration budget is preserved, and shrinking survives — the shrinker shrinks the underlying `Long` and the mapping is re-applied, so you still get a minimal counterexample expressed in your own type. Prefer `map` for wrappers, value classes, formatting and any derivation that cannot fail. One caveat: `map` with a lossy function collapses the space. `Arb.int().map { it % 2 }` generates exactly two values however many iterations you configure — that is a `Exhaustive`-shaped domain expressed as an `Arb`, and it also makes the shrunk counterexample look strange because shrinking happens in the pre-image. ### flatMap — dependent generation ```kotlin val listAndIndex: Arb<Pair<List<Int>, Int>> = Arb.list(Arb.int(), 1..20).flatMap { xs -> Arb.int(0..xs.lastIndex).map { i -> xs to i } } ``` `flatMap` gives you the sampled value and asks you for the *next generator*. That is the mechanism for any correlation: an index inside a list, an end date after a start date, a payload whose size matches a declared length field. This is the answer to "my two fields must agree" — `Arb.bind` samples components independently and cannot express it. The cost is shrinking quality. When the first value shrinks, the second generator changes, so the shrink search over a `flatMap` is less well-behaved than over a `map`. Keep the dependent part small — derive rather than re-generate where you can (`Arb.list(...).map { xs -> xs to xs.size / 2 }` needs no flatMap at all). ### filter — rejection sampling ```kotlin val evens = Arb.int(0..100).filter { it % 2 == 0 } // works, but pays for it ``` `filter` re-draws from the source until the predicate passes. Four things go wrong as the predicate gets stricter: 1. **Wasted work.** Every rejected draw costs a sample and, for expensive generators, real time. 2. **Distribution skew.** You configured a range or a size distribution; filtering silently reshapes it. `Arb.int().filter { it in 1..10 }` does not give you the same distribution as `Arb.int(1..10)`, and it is vastly slower. 3. **Edge cases can vanish.** Kotest's generators inject known-nasty values; a filter that rejects them removes exactly the samples most likely to find a bug — sometimes correctly, often accidentally. 4. **Stalling.** A predicate that is almost never true turns generation into a long spin. "My property test hangs" is very often a too-narrow `filter`. ## The design rule **Construct, do not reject.** Almost every `filter` has a constructive equivalent: | Instead of | Write | |---|---| | `Arb.int().filter { it > 0 }` | `Arb.positiveInt()` | | `Arb.int().filter { it in 1..10 }` | `Arb.int(1..10)` | | `Arb.string().filter { it.isNotEmpty() }` | `Arb.string(minSize = 1)` | | `Arb.int(0..100).filter { it % 2 == 0 }` | `Arb.int(0..50).map { it * 2 }` | | two dates + `filter { a < b }` | `flatMap` the second from the first | A narrow, cheap filter that removes a rare degenerate case is fine — the danger is proportional to the rejection rate, not to the existence of the call. When the condition genuinely belongs to the *property* rather than to the generator ("this law only holds for non-empty inputs"), the honest options are to encode it in the generator, or to state the precondition in the property itself. ## Interview framing Interviewers use this question to check whether you think of generators as data pipelines with a cost model, or as a bag of helper functions. The signal is: `map` for shape, `flatMap` for dependency, `filter` only for cheap rare exclusions — and "prefer constructing valid values" as the one-line rule.
- A colleague's property test suddenly takes minutes instead of seconds after they added a generator constraint. Where do you look first?At a `filter` with a low acceptance rate. Rejection sampling re-draws until the predicate passes, so a predicate that is true for a tiny fraction of the source turns each iteration into many draws. Replace it with a constrained or derived generator — a narrower range, a `map` onto the valid subset, or a `flatMap` when the constraint relates two values.
- You need pairs (start, end) where end is after start. Which combinator, and why not the others?`flatMap`: generate the start, then return a generator for the end that is bounded below by the start. `Arb.bind` samples both independently and cannot correlate them, and generating two independent dates then filtering out the wrong order throws away roughly half the samples and skews the distribution of both.
saying these in an interview costs you the question
- Reaching for filter as the default way to constrain a generator
- Believing filter is free because it "just skips" bad values
- Thinking map breaks shrinking (it does not — the source shrinks and the mapping re-applies)
- Using bind where the second value must depend on the first
- Not noticing that a filter can strip out exactly the injected edge cases you wanted