skip to content

Property tests in your CI pipeline occasionally fail on inputs nobody has seen before, and the team is debating pinning fixed seeds so the suite becomes deterministic. How would you decide, and what policy would you put in place?

level: principalimportance: nice to knowfreq 20%

answer

  1. pinned seed = property degraded to unnamed examples
  2. new-input failure ≠ flakiness
  3. two artefacts: seed (temporary) + counterexample (durable)
  4. non-reproducing → hunt impurity, not noise
  5. tune iterations and shrink budget, don't freeze inputs

basics

~20 s

Do not pin seeds permanently — that trades away the exploration that justifies property testing. Keep properties random, capture every failure's seed and shrunk counterexample, promote counterexamples to example-based regression tests, and treat a genuinely flaky property as a bug in the property or the code.

solid answer

~50 s

Pinning `PropTestConfig(seed = …)` across the suite makes CI deterministic and makes the tests permanently blind: they will run the same handful of inputs forever, which is a slower, more complicated way of writing example-based tests. The policy I would set: 1. **Properties stay random.** New inputs every run is the feature, not the flakiness. 2. **Every failure produces two artefacts** — the printed seed (for immediate reproduction) and the shrunk counterexample (durable). 3. **Counterexamples become example-based tests** in the same fix commit; the seed is never committed. 4. **Non-reproducing failures are treated as impurity**, not as noise: a clock, `Random`, shared state or real concurrency somewhere in the generator, the body or the code under test. 5. **Cost is tuned, not silenced** — iteration counts sized per suite, shrink budget bounded where bodies are expensive. The only legitimate uses of a fixed seed are temporary local reproduction and, rarely, a deliberately deterministic benchmark-style test that should not be a property in the first place.

code

kotlin · 4 lines
kotlin
// TEMPORARY, local only — never committed
checkAll(PropTestConfig(seed = 8891233L), Arb.string(), Arb.int()) { s, i ->
   parse(render(s, i)) shouldBe (s to i)
}

go deeper

for a junior

Know that seeds are for reproducing a failure locally and that the regression should end up as a concrete example test.

for a middle

Explain why a permanently pinned seed removes the exploration that justifies property testing, and describe the reproduce-fix-promote-unpin workflow.

for a senior

Add the diagnosis angle: non-reproducing failures point at impurity, and cost should be tuned via iterations and shrink budget rather than by freezing inputs.

for a principal

Own the policy and the conversation with stakeholders: what a red property build means, which artefacts CI must capture, where deterministic inputs legitimately belong, and how the trade-off is reviewed over time.

## The tension Property tests generate new inputs on every run. That is precisely what makes them find bugs your example tests missed, and precisely what makes a red CI job unrepeatable. Teams under delivery pressure resolve the tension the fast way: pin a seed, get a deterministic suite, move on. It works, and it removes most of the value. ## Why permanent pinning is the wrong default A pinned seed converts a property into a fixed set of examples — chosen at random, once, by nobody. You keep the machinery cost (generators, shrinking, a slower run) and lose the benefit (new inputs). Worse, the fixed set is undocumented: nobody knows which cases the suite actually covers, so nobody can reason about the gap. A handful of hand-picked example tests would be faster, clearer and more honest. There is a second-order harm: pinning teaches the team that unrepeatable failures are a property-testing quirk to be configured away, rather than a signal. Most non-reproducing failures are impurity somewhere in the stack, and impurity in generators or in the code under test is worth finding. ## The policy **1. Properties stay random in CI.** Accept that a property can go red on an input that has never appeared before. That is the system working. **2. Capture two artefacts on every failure.** The seed gives immediate local reproduction; the shrunk counterexample is the durable description of the bug. CI logs must retain both — if your reporting truncates test output, fix that before you debate seeds. **3. Reproduce with the seed, temporarily.** `checkAll(PropTestConfig(seed = …), …)` locally, fix, verify, remove. The pinned seed never reaches `master`. **4. Promote the counterexample.** In the same commit as the fix, add an ordinary example-based test asserting the concrete minimal input. It is immune to generator edits, Kotest upgrades and seed arithmetic, and it names the regression for future readers. **5. Treat non-reproduction as a defect hunt.** If the seed does not reproduce the failure, something outside Kotest's random source is in play: `kotlin.random.Random` or a clock in a generator or body, shared mutable fixtures, database or environment state, or real concurrency in the code under test. Each of those is worth fixing on its own merits — and each of them also undermines every other test you own. **6. Tune cost deliberately.** Iteration counts sized to the value of the suite; shrink budget bounded (`ShrinkingMode.Bounded`) or disabled for the few properties with expensive bodies. Do not let "property tests are slow" become the unexamined reason for pinning. **7. Keep generators pure and reviewed.** Purity is what makes seeds meaningful; it is the foundation the whole policy rests on. ## Legitimate exceptions - **Local reproduction.** Obviously fine; it is what the seed is for. - **A deliberately fixed corpus.** If a test genuinely needs a fixed set of inputs — a golden-file comparison, a performance baseline — write it as an example-based or data-driven test rather than a property with a frozen seed. Same determinism, far clearer intent. - **A quarantined property under investigation.** Bounding or disabling shrinking, or reducing iterations, while a slow or noisy property is being diagnosed is acceptable as a temporary measure with an owner and an end date. Pinning the seed is not, because it hides the very variability you are diagnosing. ## Cultural framing The deeper argument to make with the team is about what a red build means. A property failing on a new input is not flakiness — it is a newly discovered defect, and the correct response is the same as for any other failing test: reproduce, understand, fix, add a regression test. Genuine flakiness — a property that passes and fails on the *same* input — is a different animal, and it is always a bug in the test or in the code, never something a seed should be papering over. A useful metric for the discussion: how many production defects were first caught by properties on inputs no human wrote? If the answer is more than zero, pinning would have cost you those. ## What a strong answer contains The explicit trade (determinism versus exploration), the distinction between a new-input failure and true flakiness, the durable-artefact rule (counterexample as example test, seed never committed), the impurity checklist for non-reproducing failures, and a short list of exceptions where deterministic inputs really are the right design — expressed as example-based tests rather than frozen properties.

  • How do you distinguish a property that is genuinely flaky from one that keeps finding new bugs?
    Check whether it fails on the same input twice. Re-run with the reported seed: if it fails again, the input is a real defect and the property is doing its job. If the same input passes on a re-run, the flakiness lives in the test or the code — impure generators, shared state between tests, timing or concurrency — and that is a bug to fix rather than a seed to freeze.
  • A stakeholder argues that unpredictable CI failures block releases and any determinism is worth it. What do you offer instead of pinned seeds?
    Fast triage rather than frozen inputs: guaranteed capture of the seed and shrunk counterexample in CI logs, a documented reproduce-and-promote workflow, iteration counts and shrink budgets tuned so failures surface quickly, and properties confined to pure logic so failures are cheap to diagnose. If a specific pipeline stage truly cannot tolerate exploration, run the property suite as a separate job rather than degrading it into fixed examples.

saying these in an interview costs you the question

  • Committing fixed seeds as the standard way to stabilise a property suite
  • Calling every unrepeatable property failure flakiness without checking the seed
  • Keeping a pinned seed after the bug is fixed instead of promoting the counterexample
  • Losing the shrunk counterexample because CI truncates test output
  • Reducing iterations to near-zero so properties never fail, and calling the suite stable

context