How would you decide which parts of a regression pack should become property-based tests?
answer
- Different oracles, decided behaviour by behaviour
- A rule exists, or only a table does
- Cheap runs make exploration affordable
- Gates want determinism; nightly wants breadth
- Pin every finding as an example
basics
~20 sConvert behaviour that is expressible as a general rule with an independent, cheap oracle - codecs, normalisers, calculators, queue and ordering logic. Keep example tests where the expected value is a business decision, where a case is a communication artefact, or where each run is expensive.
solid answer
~50 sDecide per behaviour, never by mandate. Properties pay where the input space is large and structured, where a rule genuinely independent of the implementation exists, and where a single case runs in milliseconds - serialisation, normalisation, ordering, pricing arithmetic, idempotent handlers. Example tests still win in four situations: the expected value is a decision rather than a derivable rule, so the only property would restate the table; the test is a communication artefact a non-engineer reads; each run costs real time or real infrastructure, so thousands of runs are unaffordable; and a release gate needs deterministic, fast failures rather than a never-seen input at the worst moment. My usual shape is a broad generated-input job outside the gate, with every finding shrunk and pinned as an example inside it. Claims that one style finds more defects per hour than the other are contested, so I decide on the shape of the behaviour, not on a study.
go deeper
Know that the two styles answer different questions - one checks named cases, the other checks a rule over many generated inputs - and that neither replaces the other wholesale.
Be ready to name concrete behaviours that suit each style: a codec or normaliser for a property, a negotiated rounding table or an acceptance scenario for examples, and say why.
Show the operational side: where a generated-input job runs relative to the release gate, how findings become pinned deterministic tests, and why an over-strong property costs more trust than it earns.
Own the standard and its economics - review for provenance, price the triage load, refuse blanket mandates, and steer by locally measured escaped defects rather than by contested industry effectiveness claims.
### The decision is per behaviour, not per team The failure mode I most want to avoid is a mandate - *every module gets a property* - which produces tautologies in the modules where no rule exists and a maintenance burden nobody defends. Property-based tests and example tests supply different oracles, and the question for each behaviour is which oracle is available and what it costs. ### Where properties earn their place Four signals, and I want at least three of them: - **The input space is large and structured**, so no realistic set of hand-written cases covers it. Anything parsing, encoding, merging, ordering, deduplicating or normalising qualifies. - **A rule exists that is independent of the implementation** - an inverse to round-trip against, an invariant from the requirement, a naive reference implementation, or a metamorphic relation. Without independence the property is a tautology. - **Defects are input-shaped**: boundaries, duplicates, equal keys, empty collections, unusual character classes, arithmetic edges. These are exactly what a human writing examples fails to imagine. - **A single case is cheap**, milliseconds and in-process, so thousands of runs are affordable. In the playlist service, the share-string codec, the queue builder with its cap and pinning rules, and the batch-update conflict resolution all score four out of four. The conflict resolution is where it paid: at a 1,200-request-per-minute peak, two updates to one playlist within a window are routine, and the defect was an **ordering assumption** no example test author had thought to write down. ### Where example tests still win **The expected value is a decision, not a rule.** Rounding tables, tier boundaries negotiated with finance, the exact wording of a notification, a regulated figure. Any *property* over these would restate the table, and a test that restates the table is a copy of the specification with the same defects. Write the examples, one per documented case. **The test is a communication artefact.** An acceptance case a product owner or an auditor reads is a scenario with concrete values. Its readability is the point; replacing it with a universally quantified rule loses the audience that the artefact exists for. **Each run is expensive.** End-to-end journeys, tests touching real infrastructure, anything measured in seconds. Property-based testing trades many cheap runs for coverage of the space; at forty seconds a run that trade collapses. Push the property down to the pure core and leave the expensive layer with a handful of examples. **The failure must be deterministic and legible in a gate.** A release gate that can fail on an input nobody has seen creates a triage decision at the moment a team is least able to make it, and the pressure to disable it is enormous. Determinism is a property of the artefact, not of its quality. **Regression pinning.** A specific past incident deserves a specific named test, permanently, independent of whether a generator would rediscover it. ### The shape I default to Properties run in a generated-input job outside the release gate - on merge to the mainline and nightly, with a larger budget nightly. Findings are shrunk, triaged and pinned as example tests inside the gate. The gate stays deterministic and fast; the exploration keeps running with a bigger budget and a longer time horizon. Where a property is cheap, mature and has not produced a false alarm in months, it graduates into the gate with a small fixed run count. ### The organisational costs to price in A good property is harder to write than a good example, and there are two ways to write a bad one: the **tautology** that can never fail, and the **over-strong** clause that fails on legitimate input. The second is more damaging, because a suite whose failures are not trusted is worse than no suite. So review properties for **provenance** - each clause traceable to a requirement, with a plausible defect it would catch - and count the triage cost of generated failures as part of the cost of ownership. Team capability matters too: if two people can read a property and disagree about what it claims, it is not ready to gate anything. ### What I do not decide on I would not justify the split with a claimed defect-detection rate. The comparative evidence between testing styles is genuinely contested and heavily context-dependent, and industry-wide numbers of that kind do not survive contact with a specific codebase. What I would do is measure locally: over two or three quarters, record which style caught each escaped defect and what each style cost in triage time. That is a number about this system, and it is the only kind worth steering by.
- A team lead proposes a property for every public function as a coverage policy. What is your response?I would reject it as a mandate. Where no implementation-independent rule exists, the only property anyone can write restates the code, passes forever and costs maintenance. I would replace the rule with a standard: any behaviour with a large structured input space and an independent oracle gets a property, and every property is reviewed for provenance - each clause traceable to a requirement, with a named defect it would catch.
- How do you protect a release gate from generated-input failures without losing their value?Run the broad generated-input job outside the gate, on merge and nightly with a larger budget, and require every finding to be shrunk, triaged and pinned as a deterministic example test inside the gate. The gate then fails only on known, readable cases. A mature, cheap property with no recent false alarms can graduate into the gate with a small fixed run count.
- Which behaviours would you deliberately leave as example tests even though a property is technically writable?Anything whose expected value is a negotiated decision - rounding tables, tier boundaries, regulated figures - because the property would restate the table and inherit its defects. Also acceptance cases read by non-engineers, where concreteness is the point, and expensive end-to-end journeys where thousands of runs are unaffordable. In the last case I push the property down to the pure core instead.
- How would you measure whether the split you chose was right?Locally and over quarters, not from published comparisons. For each defect that escaped to production, record which style should have caught it and why it did not; alongside that, track triage hours spent on generated-input failures and how many turned out to be over-strong properties. Those two series tell you whether to move a behaviour between styles. Cross-industry effectiveness claims are contested and do not transfer.
saying these in an interview costs you the question
- Mandates a property for every module regardless of behaviour
- Cites a specific defect-detection multiplier as settled fact
- Puts a broad generated-input job directly in the release gate
- Deletes example tests once properties exist
- Writes properties over negotiated tables and rounding rules
- Ignores the triage cost of failures on never-seen inputs