skip to content

What do the `minSuccess` and `maxFailure` settings on Kotest's `PropTestConfig` do, and when is it legitimate to move them off their defaults?

level: seniorimportance: should knowfreq 22%

answer

  1. default: 1 failure = fail, then shrink
  2. maxFailure = failure budget
  3. minSuccess = successes required
  4. choose both relative to iterations
  5. only honest for statistical specs

basics

~20 s

They set a failure tolerance. By default every iteration must pass — one failure fails the property. Raising maxFailure lets that many iterations fail before the test fails; minSuccess requires at least that many passing iterations overall. Use only for genuinely statistical properties.

solid answer

~60 s

By default a Kotest property is all-or-nothing: the first failing iteration aborts the run, triggers shrinking and fails the test. `PropTestConfig` lets you relax that. - **`maxFailure`** is the number of failing iterations the run will tolerate. Exceed it and the property fails, reporting the failure that broke the budget. - **`minSuccess`** is the number of iterations that must succeed for the property to pass; too few successes at the end is a failure even if the failure budget was never blown. You pass the config as the first argument: `checkAll(PropTestConfig(maxFailure = 10, minSuccess = 990), Arb.int()) { … }`. The honest use case is a genuinely probabilistic property — a heuristic, a sampling algorithm, a bloom-filter false-positive rate — where "holds for almost all inputs" is the real specification. Everything else (a generator producing inputs outside the contract, a flaky dependency, a real bug in a corner of the space) should be fixed at the source, because a tolerance turns a reproducible failure into an intermittently green build.

code

kotlin · 6 lines
kotlin
checkAll(
    PropTestConfig(iterations = 1000, minSuccess = 990, maxFailure = 10),
    Arb.string(minSize = 1, maxSize = 64)
) { s ->
    bloomFilter.mightContain(s) shouldBe true
}

go deeper

for a junior

Know that a Kotest property fails on the first failing iteration by default, and that PropTestConfig exists to change run settings.

for a middle

State what each setting means, that they are chosen relative to the iteration count, and how the config is passed to checkAll/forAll.

for a senior

Focus on judgement: name the genuinely statistical cases where a tolerance is correct, and the four disguised problems (bad generator, corner-case bug, nondeterministic SUT, over-strong assertion) where it is not.

for a principal

Treat it as a policy question — what the team's bar is for admitting a tolerance into the suite, how the statistical rationale is documented and reviewed, and the long-term cost to CI signal when properties stop failing deterministically.

## The default is zero tolerance A Kotest property test is, by default, a universally quantified claim: *for every value the generators produce, this holds*. Concretely that means the first failing iteration ends the run — the engine stops, shrinks the failing input to a minimal counter-example, and fails the test with that input, the underlying assertion error and the seed. `PropTestConfig` is the object that lets you change how the run is executed, and two of its fields specifically change the pass/fail arithmetic. ## `maxFailure` — a failure budget `maxFailure` is the number of failing iterations the run will absorb before declaring the property broken. With the default of zero, a single failure is fatal — which is why shrinking kicks in immediately. Raise it and the engine keeps iterating past failures, only failing the test once the budget is exceeded. The important consequence is diagnostic: with a tolerance in place, the run is no longer "stop at the first anomaly and minimise it". You have told the framework that some failures are expected, so a report arrives only when the *rate* of failure is too high, and the counter-example you get is the one that happened to break the budget — not necessarily the most interesting one. ## `minSuccess` — a success floor `minSuccess` approaches the same question from the other side: how many iterations must actually succeed for the property to count as satisfied. It is checked against the successes accumulated over the run, so a property that mostly discards inputs or mostly fails will be rejected even if no single mechanism blew a budget. The two settings are normally chosen together, relative to the iteration count: with 1000 iterations, `minSuccess = 990` and `maxFailure = 10` express "at least 99% of inputs must satisfy this". Setting them incoherently (a `minSuccess` above the iteration count, for instance) is a way to make a property that can never pass. ```kotlin checkAll( PropTestConfig(iterations = 1000, minSuccess = 990, maxFailure = 10), Arb.string(minSize = 1, maxSize = 64) ) { s -> bloomFilter.mightContain(s) shouldBe true } ``` ## When relaxing is honest There is a small, real class of properties where the specification itself is statistical: - **Probabilistic data structures** — a bloom filter's false-positive rate, a sketch's error bound. - **Heuristics and approximations** — a ranker, a fuzzy matcher, a compression estimator that is allowed to be wrong occasionally. - **Sampling or randomised algorithms** where the guarantee is "with high probability". For these, zero tolerance is not merely inconvenient, it is *wrong*: the test would be asserting a property the system never claimed. Encoding the rate in `minSuccess`/`maxFailure` makes the real contract explicit and reviewable, and is far better than deleting the test. ## When relaxing is a smell Every other reason to reach for these knobs is really a different problem wearing a disguise: 1. **The generator produces inputs outside the contract.** A property over `Arb.string()` that breaks on empty strings is telling you the domain is wrong, not that the code is 99% right. Constrain the generator instead. 2. **A real bug lives in a corner of the input space.** A tolerance hides exactly the inputs property testing exists to find — the rare ones. 3. **The system under test is nondeterministic** (time, ordering, shared state, an external dependency). Then the property is not the flaky part; the environment is, and a failure budget converts a design problem into intermittent CI noise. 4. **The assertion is too strong.** Often the right move is to assert the weaker property that is actually true, which is more informative than the strong property tolerated 1% of the time. ## Review guidance If a diff sets `maxFailure`, the diff should also say *why* in a comment or the test name — with the rate the system actually guarantees. A tolerance without a stated statistical rationale should be challenged: it is a build-greening device, and its effect on the suite is corrosive because the property no longer fails deterministically on a real regression. Prefer, in order: fix the code, constrain the generator, weaken the assertion to the true property, and only then encode a documented failure rate. Finally, note that these settings tolerate failures — they do not *retry* them. There is no re-execution of a failing input, and no attempt to distinguish a flaky failure from a deterministic one. The engine simply counts.

  • How does raising `maxFailure` change the counter-example you get in the failure report?
    With the default of zero, the run aborts on the first failure and shrinks it, so you get a minimal counter-example. With a budget, the run continues past failures and only reports once the budget is exceeded, so the input you are shown is the one that happened to break the budget rather than the most illuminating one. You trade minimality for a rate check.
  • A property fails only for empty strings. Would you set `maxFailure = 1`?
    No — that is a domain problem, not a statistical one. Either the code should handle empty strings, in which case fix the code, or empty input is out of contract, in which case constrain the generator (a minimum size) so the property states its real precondition. A tolerance would hide the very edge case property testing is designed to surface.
  • Do `minSuccess` and `maxFailure` make a property test retry flaky iterations?
    No. They only change the arithmetic of pass and fail — the engine counts successes and failures and compares them against the thresholds. No input is ever re-executed, and a flaky failure is indistinguishable from a deterministic one to the framework. Flakiness must be addressed in the system under test or its environment.

saying these in an interview costs you the question

  • Using `maxFailure` to quiet a flaky test rather than diagnosing the flakiness
  • Believing these settings retry failing iterations
  • Thinking the default is some non-zero tolerance rather than all-or-nothing
  • Setting `minSuccess` higher than the iteration count and expecting it to pass
  • Assuming shrinking still yields a minimal counter-example once a failure budget is in place

context