skip to content

A property test fails on a generated input in nightly CI. How do you turn it into a durable regression test?

level: seniorimportance: should knowfreq 44%

answer

  1. The failing input is mostly noise
  2. Simplify while it still fails
  3. Decide: code, property, or generator
  4. Commit the small case, not the seed
  5. Keep the hunter running after the fix

basics

~20 s

Shrink the failure to the smallest input that still fails, decide whether the code or the property is wrong, fix it, and commit the minimal case as a named example test so the defect is caught deterministically instead of waiting for the generator to rediscover it.

solid answer

~50 s

Start with the **minimal counterexample**: the shrinker repeatedly simplifies the failing input and keeps any simplification that still fails, so a 214-track playlist becomes two tracks added in the same second. That minimal case is the diagnosis - it names the condition and strips the confounds. Then triage: a genuine defect, a property that asserted more than the requirement promises, or an input the domain can never contain. After fixing, **pin the minimal case as an explicit example test** named for the rule it protects, and confirm it fails before the fix and passes after. Keep the property running as well; the example guards the known case, the property keeps hunting for the next one. Treat a recorded seed as a diagnostic aid for reproducing while you work, not as the permanent artefact - it is tied to the generator definitions and library version and rots when either changes.

code

pseudocode · 6 lines
pseudocode
# minimal counterexample from the nightly run, shrunk from 214 tracks to 2
test "tracks added in the same second keep their insertion order":
    p = playlist(tracks = [ track("t-4471", addedAt = "09:14:22"),
                            track("t-9083", addedAt = "09:14:22") ])
    q = buildQueue(p)
    assert idsOf(q) == ["t-4471", "t-9083"]   # red before the fix, green after

go deeper

for a junior

Know that the failing input gets simplified to a minimal counterexample, and that the small case should be copied into an ordinary test so the defect stays caught after the fix.

for a middle

Be ready to explain what shrinking is doing - simplify, re-check, keep what still fails - and why a pinned example plus the surviving property is better than either artefact alone.

for a senior

Show the triage judgement: code wrong, property over-strong, or input outside the domain, with a reason for the choice; and separate a property finding new inputs from genuine flakiness on unchanged input.

for a principal

Own the policy: where generated-input jobs run relative to the release gate, how findings are converted into the pinned corpus, and how the team is stopped from quarantining a detector because its failures are inconvenient.

### Why the raw counterexample is nearly useless When a generated input fails, the input it failed on is usually enormous and mostly irrelevant. A nightly run against the playlist service produced a failing playlist of 214 tracks, 9 of them pinned, with unicode titles and three duplicated ids. Somewhere in that is one relevant fact and 213 distractions. **Shrinking** is the search that finds the relevant fact. Conceptually it is simple: propose a simpler version of the failing input - fewer elements, smaller numbers, shorter strings - re-run the property, and if it still fails, adopt the simpler version and repeat; if it passes, try a different simplification. What comes out is a **minimal counterexample**: an input where removing anything makes the failure disappear. Here the shrink landed on two tracks with the same added-at timestamp. That is the whole diagnosis. The queue builder sorted by added-at and relied on the sort leaving equal elements in their original order - an **ordering assumption** that the underlying sort never promised. Reading *two tracks, equal timestamps, wrong order out* takes a second; reading the 214-track original takes an afternoon. ### Triage before you fix A failing property has three possible causes, and choosing the wrong one produces bad tests: 1. **The code is wrong.** The minimal input is something a real user can create, and the behaviour violates the requirement. Fix the code. 2. **The property is wrong.** It asserted something the requirement never promised - a total ordering where ties are permitted, a canonical form the specification leaves free. Weaken the clause to what is actually required. 3. **The input is impossible.** The generator produced a value the domain forbids. Tighten the generator - and be honest about whether the domain really forbids it, because *impossible* is the most over-used of the three. Write down which one it was. That sentence is what a reviewer needs to judge the change. ### Pinning the case The generated failure is not repeatable by construction: the next run generates different inputs. So convert the minimal counterexample into an ordinary example test, checked into the same suite: - **name it for the rule**, not for the incident number - the name is the documentation a future reader gets; - **assert the specific behaviour**, not merely absence of an exception; - **verify it red before the fix and green after.** A pinned case that never failed is a case you have no evidence about; - keep it small - the value of the pinned case is that it reads as a statement of the rule. And keep the property. The two artefacts do different jobs: the example test guarantees this exact regression is caught on every run, in a fraction of a second, deterministically, and it fails in the gate in a way anyone can read; the property keeps searching the space for the next member of the failing class. Deleting the property because it found something is a common and self-defeating instinct. ### Seeds are a diagnostic, not an artefact Most libraries print a seed with a failure so the same generated sequence can be replayed. That is genuinely useful for the twenty minutes you spend reproducing and stepping through. It is a poor permanent regression test, because a seed only reproduces the same inputs while the generator definitions, the library version and everything nondeterministic inside the system under test stay identical. Change the generator - which the fix may well require - and the seed replays something else entirely. Record it in the failure report and the incident notes; commit the shrunk value. ### Is a nightly property failure flaky? This distinction matters and interviewers probe it. **Flaky** normally means nondeterministic on unchanged code and unchanged input - a race, a clock, a shared fixture. A property failing on a *different input each night* is not that: the code is deterministic for each input, and the run keeps finding new members of a real failing class. Calling it flaky, and quarantining it, discards a working defect detector. The right response is to shrink one of them, look for the shared shape - here, equal timestamps - fix the class, and pin one representative. The opposite case exists too: a property that fails on the *same* generated input inconsistently is genuinely flaky, and the cause is nearly always shared state between iterations or a dependence on wall-clock time or environment. Fix that before trusting anything the property reports. ### The corpus you accumulate Over a year, the pinned counterexamples become the most valuable part of the suite: every one of them is a real defect that a human would not have thought to write down, expressed in the smallest form that exhibits it. Keep them together, keep them named for their rules, and let the property keep adding to the collection.

  • Why not simply commit the recorded seed as the regression test?
    Because a seed reproduces a sequence, not a value. It replays the same inputs only while the generator definitions, the library version and any nondeterminism inside the system under test are unchanged - and fixing the defect often changes the generator. A seeded test can then pass while testing something else entirely, which is worse than no test. Record the seed for reproduction during diagnosis; commit the shrunk input.
  • A property has failed on four different generated inputs in four nightly runs. Is that flakiness?
    No, in the usual sense. Flaky means nondeterministic on unchanged code and unchanged input; here each input is new and each failure is deterministic for that input. The generator is finding successive members of one real failing class. Shrink two of them, look for the shared shape, fix that class, and pin one representative. Quarantining the property would discard a working detector.
  • How do you decide whether a failure means the property was over-strong rather than the code being wrong?
    Go back to the requirement, not the code. Read the minimal counterexample and ask whether the requirement genuinely promises the asserted behaviour for that input. If the specification permits ties, several valid orderings, or a non-canonical form, the property claimed more than was promised and the clause should be weakened. If a user can create the input and the observed result contradicts a documented rule, the code is wrong.

saying these in an interview costs you the question

  • Commits the seed instead of the shrunk input
  • Deletes the property once the defect is fixed
  • Quarantines a property that finds a new input each run
  • Pins the original huge counterexample unshrunk
  • Adds the pinned test without seeing it fail first
  • Blames the generator for every unexpected failing input

context