skip to content

A Kotest property test fails on roughly one CI run in twenty. A teammate proposes setting `maxFailure` on Kotest's `PropTestConfig` so the build goes green again. How would you evaluate and handle that?

level: principalimportance: nice to knowfreq 16%

answer

  1. intermittent property failure = a finding
  2. reproduce from seed, read shrunk input
  3. classify: bug / generator / assertion / nondeterminism / statistical
  4. budget only for a rate-shaped contract
  5. graduate counter-examples into example tests

basics

~20 s

Diagnose first: reproduce from the printed seed, look at the shrunk input, and decide whether it is a real corner-case bug, an out-of-contract generator, an over-strong assertion, or a nondeterministic system. A failure budget is only correct for a genuinely statistical specification.

solid answer

~50 s

An intermittently failing property is a *finding*, not noise — the whole point of generated inputs is that a rare input shows up rarely. So before touching `maxFailure` I reproduce: the failure report carries the shrunk counter-example and the seed, so the failure can be pinned into a deterministic run. Then I classify it: 1. **Real bug in a corner of the input space** — fix the code; keep the shrunk input as a regression example test. 2. **Generator wider than the contract** — constrain the `Arb` (ranges, sizes, character set) so the property states its true precondition. 3. **Assertion stronger than the guarantee** — weaken it to the property that actually holds; often expressed better with `checkAll` and clue-carrying matchers than a bare `forAll` boolean. 4. **Nondeterministic system or environment** — fix the determinism; a failure budget just launders flakiness into a green build. Only a genuinely statistical specification earns `minSuccess`/`maxFailure`, and then with the rate documented next to it.

go deeper

for a junior

Know that a property failure prints a shrunk counter-example and a seed, and that the first move is to reproduce it, not to relax the configuration.

for a middle

Walk the classification — real bug, too-wide generator, too-strong assertion, nondeterministic system — and name the fix for each.

for a senior

Emphasise diagnosis and reproduction discipline, promoting counter-examples into deterministic example tests, and why a budget cannot tell causes apart.

for a principal

Frame it as suite-health policy: the review bar for admitting a tolerance, the cost to CI signal, and how the team keeps exploratory properties and deterministic regression tests in their proper roles.

## Why the proposal is tempting and usually wrong `PropTestConfig(maxFailure = n)` tells a Kotest property to tolerate up to *n* failing iterations before failing the test; `minSuccess` sets the floor of iterations that must pass. Reaching for either in response to an intermittently red build converts a signal into silence. And the signal here is unusually valuable: a property test explores a space no example test visits, so a 1-in-20 failure typically means "about 1 in 20 *runs* draws an input from a region where the system is wrong" — a rare-input bug, which is exactly what property testing is for. ## Step one: make it deterministic before deciding anything Every property failure prints the shrunk counter-example and the seed that produced the run. That is enough to turn an intermittent failure into a repeatable one locally, which is the precondition for any real decision. Alongside that, read the *shrunk* input rather than the raw one — the minimised value usually names the bug outright (empty collection, zero, `Int.MIN_VALUE`, a surrogate pair, a duplicate key). If the failure is caused by an assertion inside `checkAll`, the matcher's expected/actual message is in the report; if it is a bare `forAll` boolean, the first improvement is usually to rewrite the property as `checkAll` with assertions, so the next failure explains itself. Improving diagnostics is a legitimate response to intermittency; suppressing them is not. ## Step two: classify the failure **(a) A real bug in a corner of the input space.** Most common outcome. Fix the code, and pin the shrunk value into a fast example-based test so the specific regression is caught deterministically forever after. The property stays as the explorer; the example test stays as the guard. **(b) The generator is wider than the contract.** The property is asserting something the system never promised — an empty string where a non-empty identifier is required, a negative quantity, a 100-element list where the API documents a cap. The fix is to *constrain the generator* (range, size range, character set) so the domain of the property equals the domain of the contract. This is a better fix than filtering inside the property, which silently discards inputs and can starve the run. **(c) The assertion is stronger than the guarantee.** Sometimes the code is right and the property is over-claiming — asserting exact equality where the API guarantees ordering-insensitive equality, or a precise value where a bound is promised. Rewrite the property to the true, weaker claim. A weaker property that always holds is worth far more than a stronger one tolerated 5% of the time. **(d) The system or environment is nondeterministic.** Shared mutable state between iterations, wall-clock or time-zone dependence, iteration order of a hash-based collection, a live dependency. Here the property is innocent. A failure budget would make the build green while leaving a production-relevant nondeterminism in place — and, worse, would also mask category (a) bugs from then on, because the budget cannot distinguish causes. It only counts. **(e) A genuinely statistical specification.** Probabilistic data structures, heuristics, approximation algorithms: the contract *is* a rate. This is the one case where `minSuccess`/`maxFailure` are the correct encoding, because zero tolerance would assert something the system never claimed. ## If a tolerance really is right Make it legible. Set `minSuccess` and `maxFailure` coherently against the iteration count so the numbers read as the guaranteed rate (e.g. 1000 iterations, `minSuccess = 990`, `maxFailure = 10` for a 99% claim), name the test after the guarantee, and put the justification beside the config. Be explicit in review that the property no longer fails deterministically on a regression — that is the price, and it should be paid knowingly. ## The organisational angle The deeper issue is what the team does with flaky signals in general. Two habits keep property tests trustworthy: - **A tolerance requires a stated statistical rationale in review.** "To make CI green" is not one. Without this rule, `maxFailure` becomes the local anaesthetic that spreads through the suite. - **Every diagnosed counter-example graduates into a deterministic example test.** This keeps the property suite exploratory and cheap while the regression protection lives in fast, always-reproducible tests. The answer to the teammate, then, is not "no" but "not yet": reproduce from the seed, read the shrunk input, and let the classification decide. Four of the five categories have a better fix than a budget, and the fifth deserves the budget for a reason you can write down.

  • What do you do with the shrunk counter-example once the bug is fixed?
    Promote it into a fast, deterministic example-based test alongside the property. The property keeps exploring the space for the next unknown input, while the example test guarantees this specific regression fails the build every time and immediately, with no dependence on generator luck or the seed.
  • How does rewriting a `forAll` property as `checkAll` help with an intermittent failure?
    `forAll` only tells you the predicate returned `false` for some input; `checkAll` runs assertions, so the matcher's expected/actual message travels into the failure report. When failures are rare, you may only get one chance to read the output from a CI log, so maximising what that output says is often the highest-value change you can make before deciding anything else.
  • When is constraining the generator the wrong fix?
    When the constraint removes the very inputs the system is supposed to handle. Narrowing an `Arb` until the property passes is indistinguishable from deleting the test's value: you must be able to point at the documented contract that excludes those inputs. If no such contract exists, the input is legal and the code, not the generator, is wrong.

saying these in an interview costs you the question

  • Treating a rare property failure as flakiness rather than as a rare-input finding
  • Setting `maxFailure` without any statistical rationale, purely to green the build
  • Narrowing the generator until the property passes, without a contract that justifies the narrowing
  • Believing a failure budget can distinguish an environmental flake from a genuine corner-case bug
  • Deleting the property test instead of weakening the assertion to the claim that actually holds

context