When is property-based testing the right tool versus example tests, and how do you design properties (laws, round-trips, oracles) that avoid tautologies and flakiness?
answer
- Patterns: round-trip, laws, invariants, oracle
- Anti-pattern: tautology (re-implement logic)
- Keep generators pure → reproducible (seed)
- Bound generation cost to avoid timeouts
- Mix Exhaustive (critical) + Arb; classify to confirm coverage
basics
~20 sUse property tests when there's a general rule the code must always obey, like 'encode then decode returns the original'. Pick rules that don't just re-run the code under test, and keep generators deterministic so tests don't flake.
solid answer
~40 sPBT shines for invariants and algebraic laws; example tests stay best for specific known scenarios and regression pins. Good property patterns: **round-trips** (`decode(encode(x)) == x`), **algebraic laws** (associativity, identity, idempotence), **invariants** (output always sorted/non-negative), and **oracles** (compare against a simpler reference implementation). The classic anti-pattern is the **tautology** — reimplementing the production logic inside the assertion, so both agree and prove nothing. Avoid flakiness by keeping generators pure (only the injected `RandomSource`), logging the seed for reproduction, bounding generation cost, and not depending on wall-clock or external state. For coverage gaps, combine `Exhaustive` for small critical spaces with `Arb` for the rest, and use Kotest's coverage/classification (e.g. labeling generated inputs) to confirm interesting cases actually appear.
code
kotlin · 14 linesimport io.kotest.property.*
import io.kotest.property.arbitrary.*
import io.kotest.matchers.shouldBe
import io.kotest.matchers.collections.shouldContainExactlyInAnyOrder
suspend fun sortProperties(sortFn: (List<Int>) -> List<Int>) {
checkAll(Arb.list(Arb.int(), 0..50)) { xs ->
val out = sortFn(xs)
// invariant: ordered
out shouldBe out.sorted()
// invariant: same multiset (oracle-free permutation check)
out shouldContainExactlyInAnyOrder xs
}
}go deeper
Can name one good property (round-trip) and knows example tests still have a place.
Lists multiple property patterns (laws, invariants, oracle) and recognizes the tautology anti-pattern.
Designs independent properties, controls generation cost, and ensures reproducibility via pure generators and seeds.
Sets org-wide PBT strategy: shared reviewed generator library, coverage classification, CI iteration/seed policy, and pairing with mutation testing to prove properties bite.
## When PBT vs example tests **Reach for PBT when a general law exists:** - Serialization round-trips, parsers, codecs. - Algebraic structures (sorting, set ops, math, monoids). - Invariants over outputs (length preserved, no duplicates, always sorted). **Stay with example tests when:** - The behavior is a fixed, enumerated mapping (specific input → specific output). - You're pinning a known regression. - There's no generalizable rule, only business-specific cases. The two are complementary; mature suites use both. ## Designing good properties ### 1. Round-trip (there-and-back) ```kotlin checkAll(userArb) { user -> Json.decodeFromString<User>(Json.encodeToString(user)) shouldBe user } ``` Proves encode/decode are inverses without restating the format. ### 2. Algebraic laws - **Associativity:** `(a ∘ b) ∘ c == a ∘ (b ∘ c)` - **Identity:** `x ∘ e == x` - **Idempotence:** `f(f(x)) == f(x)` (e.g. normalizing, sorting again) - **Commutativity:** `a + b == b + a` ### 3. Invariants The output always satisfies a structural rule: `sort(xs)` is ordered and a permutation of `xs`. ### 4. Oracle / model-based Compare the optimized implementation against a slow-but-obviously-correct reference (the *oracle*): `fastSum(xs) shouldBe xs.fold(0L, Long::plus)`. ## Anti-patterns to avoid ### Tautology The deadliest mistake — copying the production algorithm into the test: ```kotlin // BAD: re-implements the same logic, so it can't catch its bugs checkAll(Arb.int(), Arb.int()) { a, b -> add(a, b) shouldBe (a + b) // fine if add wraps +, but if add IS a+b reimplemented, it's circular } ``` Prefer an *independent* characterization (an invariant or a different reference) so the test can disagree with the code. ### Flakiness sources - **Impure generators**: reading clocks, files, or shared mutable state → non-reproducible failures. Generate only from the injected `RandomSource`. - **Unbounded cost**: huge collections per iteration make tests slow/timeout; bound sizes (`Arb.list(elem, 0..50)`). - **No seed logging**: makes failures irreproducible; Kotest prints the seed — capture it. ## Confirming coverage of interesting cases Random generation can under-sample the cases you care about. Mitigate by: - Mixing **`Exhaustive`** for small critical subspaces (states/flags) with **`Arb`** for the rest. - Using **classification/labeling** of generated values (Kotest's `collect`/classification utilities) to assert that, say, empty inputs and large inputs both actually occurred. - Asserting *negative* properties too (the function rejects invalid input), not only happy-path laws. ## Operational strategy - Keep **iteration budgets** sane for CI (more in nightly, fewer in PR). - Standardize **seed logging** so any failure is reproducible. - Treat generators as a **shared, reviewed library** (domain Arbs with good shrinkers), since a bad generator silently weakens many tests. - Combine with **mutation testing** if you want evidence the properties truly bite.
- How do you know a property test isn't a tautology?The assertion must use an independent characterization (an invariant or a separate reference oracle), so it can fail when the production code is wrong rather than mirroring its logic.
- How would you confirm that interesting inputs (empty, huge) actually appear?Use Kotest's classification/collect to label generated values and assert each category occurred, or mix Exhaustive for the critical small cases.
- How do you keep property tests reproducible in CI?Keep generators pure (only the injected RandomSource), log the seed on failure, and rerun with PropTestConfig(seed = …) to reproduce exactly.
A property is a law of physics the code must obey; an example test is a single measurement. Don't let the law just restate the apparatus.
saying these in an interview costs you the question
- Writes properties that re-implement production logic (tautology)
- Generators read clocks/files/shared state, causing irreproducible flakes
- Generates unbounded huge inputs, causing timeouts
- Claims PBT replaces all example/regression tests
- No plan to confirm interesting cases are actually sampled