skip to content

Property-Based Testing

Stating a rule that must hold for every input and letting a generator hunt for a counterexample. Interviewers probe whether you can express behaviour as an invariant rather than an example.

on this pageshow

questions

5

What is a round-trip property in property-based testing, and what bug can it miss?

level: juniorimportance: must knowfreq 63%

answer

  1. Do it, then undo it
  2. The input is its own expected value
  3. Both halves can be wrong together
  4. Order-blind equality hides the defect
  5. Golden samples check the format itself

basics

~20 s

A round-trip property asserts that transforming a value and then reversing the transformation returns the original value, for every generated input. It misses paired defects: when both directions are wrong in mirror-image ways, the trip still returns the input unchanged.

solid answer

~50 s

A round-trip property pairs a transformation with its inverse - write then read, encode then decode, serialise then parse - and asserts that the result equals the original for thousands of generated inputs. It is the cheapest property to state because the input is its own expected value, so you need no independently computed result. Its blind spot is symmetry: if the writer omits a field and the reader supplies the same wrong default, the trip is clean and both halves stay broken, so a passing round trip says nothing about conformance to a published format. It is also only as strong as the equality it compares with - a comparison that ignores element order will pass happily while ordering is destroyed. Strengthen it with an invariant on the intermediate form plus a couple of fixed, hand-written samples.

code

pseudocode · 6 lines
pseudocode
property "encoding a playlist and parsing it back returns the same playlist":
    for each generated playlist p:
        s = encode(p)
        parsed = parse(s)
        assert parsed.tracks == p.tracks   # sequence equality: order counts
        assert parsed.name  == p.name

go deeper

for a junior

Be ready to state the rule in one line - transform, reverse, compare with the original - and to name two things it fits, such as a serialiser and its parser. Say out loud that the input serves as the expected value.

for a middle

An interviewer expects you to explain why the comparison used inside the property is part of the property, and to give a concrete case where order-insensitive equality hides a real defect while the test stays green.

for a senior

Show that you know self-consistency is not conformance: describe the symmetric defect where writer and reader share the same deviation, and say what you add - fixed samples from the specification, invariants on the encoded artefact - to catch it.

for a principal

Own the judgement of when a round trip is worth the effort at all: it pays where an inverse genuinely exists and the pair is independently authored, and it is theatre where one half is generated from the other.

### Why a property needs an oracle, and why the round trip is the cheapest one An **oracle** is whatever decides that observed behaviour is wrong. In an example-based test you supply the oracle by hand: you write the input and you write the expected output next to it. That does not scale to generated inputs, because nobody can hand-write the expected output for ten thousand values nobody has seen yet. A property replaces the table of expected values with a **rule that must hold for every input**, and the rule is the oracle. The round trip is the first property most teams reach for because it needs no independent expected value at all. If `decode` is the inverse of `encode`, then for every value `v`: ``` decode(encode(v)) == v ``` The input is its own oracle. That is why round trips fit codecs, serialisers, compressors, query-string builders and parsers, forward and backward schema migrations, and encryption and decryption. ### A worked failure A music-streaming playlist service turns a playlist into a share string that another client can paste back. The property is stated once and run against generated playlists: random track counts, unicode titles, repeated tracks, empty playlists. On a run of 3,700 generated playlists the property fails. The failing input has 214 tracks; the shrinker reduces it to two tracks that were added within the same second. The share encoder had collected tracks into a structure with no defined iteration order, so two tracks with equal added-at timestamps came back in whichever order the structure felt like. That is an **ordering assumption**: the code assumed a stable order that the data structure never promised. No example test in the pack had two tracks sharing a timestamp, so it had been invisible for months. ### The two ways a round trip lies to you **Symmetric defects.** The round trip checks the pair against *itself*, not against a specification. If the encoder writes the track duration in a field the published format reserves for something else, and the reader reads the same field, the trip passes forever - and the mobile client written from the format document cannot read a single share string. Self-consistency is not conformance. The fix is not more generated inputs; it is a small set of **golden samples**: fixed strings taken from the specification, parsed and compared against the value they are documented to mean. **Weak equality.** The property is only as sharp as the comparison inside it. If playlist equality is defined over the *set* of track ids, the ordering defect above passes every run. Whenever you write a round-trip property, ask explicitly what the equality compares, and whether every field the users can observe takes part in it. ### Tautology and direction A round trip asserts nothing if the two halves are not independent. If the reader is generated from the writer's own lookup table, you are asserting that a table agrees with itself. Independent authorship, or a golden sample, is what makes the assertion real. Direction also matters. `decode(encode(v)) == v` is usually true by design. The reverse, `encode(decode(s)) == s`, is a different and much stronger claim, and it is often false for correct code: two different strings can decode to the same value once whitespace, field order or optional defaults are normalised away. The honest form is to round-trip through the canonical shape: ``` decode(encode(decode(s))) == decode(s) ``` ### What to pair it with A round trip is a **partial oracle** - it constrains behaviour without fully specifying it. Pair it with: - an **invariant** on the intermediate artefact (the encoded string is non-empty; it carries exactly one separator per track boundary); - a couple of fixed samples from the format document, so conformance is checked as well as self-consistency; - generation that actually reaches the interesting shapes - duplicated ids, equal timestamps, an empty collection - since a generator that only produces two-to-five distinct tracks would never have found the ordering defect at all. Finally, a passing property is evidence, not proof. It says the rule held for the inputs that were tried this run. Different runs try different inputs, which is exactly why the run that finally produced two equal timestamps was the run that told you something new.

  • Why is encode(decode(s)) == s often false even when both halves are correct?
    Because many different encoded strings can mean the same value once optional fields, whitespace or field order are normalised. Re-encoding produces the canonical form, which need not be the string you started from. The honest version compares canonical forms: decode(encode(decode(s))) equals decode(s). Asserting the raw string round trip forces the encoder to preserve incidental formatting, which is rarely a real requirement.
  • How do you stop a round-trip property from being a tautology?
    Make the two halves independent. If the reader is derived from the writer's own table, or one calls the other, the property asserts only that the code agrees with itself. Add fixed samples drawn from the format specification, and check that the encoded artefact satisfies invariants stated in the document - length, separators, required fields - rather than only that parsing undoes encoding.
  • A round-trip property passes for a year, then fails on one generated input. Is the property or the code at fault?
    Either is possible and triage decides. The generator simply reached a shape it had not produced before, so the passing year was never proof. Shrink to the smallest failing input and read it: if it is a legitimate value a user can create, the code is wrong; if the generator manufactured something the domain forbids, the generator is too loose and the property was asserting more than the requirement promises.

Translating a sentence into another language and back proves the two translators agree with each other - not that either of them speaks the language the dictionary describes.

saying these in an interview costs you the question

  • Claims a passing round trip proves the published format is honoured
  • Compares results with equality that ignores element order
  • Derives the reverse function from the forward function's own table
  • Asserts encode(decode(s)) equals the original string unconditionally
  • Says a passing property proves no counterexample exists
  • Applies round trips to behaviour that has no inverse at all

context

open as a page

How do you state a property for a function that has no inverse to round-trip against?

level: middleimportance: must knowfreq 56%

basics

~10 s

Assert a rule the output must satisfy rather than a specific value: an invariant or postcondition, idempotence, agreement with a simpler reference implementation, or a metamorphic relation between two runs on related inputs.

open as a page

A property test fails on a generated input in nightly CI. How do you turn it into a durable regression test?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Shrink the failure to the smallest input that still fails, decide whether the code or the property is wrong, fix it, and commit the minimal case as a named example test so the defect is caught deterministically instead of waiting for the generator to rediscover it.

open as a page

Why does a naive random generator explore almost none of a structured input space?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Because independently randomised fields almost never assemble into the interesting shapes: most values are rejected at the first validation check, or are small and unrelated, so the states that carry the defects - collisions, duplicates, boundaries - are never reached.

open as a page

How would you decide which parts of a regression pack should become property-based tests?

level: principalimportance: should knowfreq 40%

basics

~20 s

Convert behaviour that is expressible as a general rule with an independent, cheap oracle - codecs, normalisers, calculators, queue and ordering logic. Keep example tests where the expected value is a business decision, where a case is a communication artefact, or where each run is expensive.

open as a page