skip to content

A drafted batch of service-level cases asserts the exact responses captured when the drafts were made. What can that suite not detect?

level: middleimportance: should knowfreq 47%

answer

  1. Ask where the expectation came from
  2. Observed output is not intended output
  3. A defect at capture becomes expected
  4. Intended change reddens a wide band
  5. Re-recording repins, it does not repair

basics

~20 s

Expectations copied from observed output encode whatever the system does today, defects included. Such a suite detects change rather than wrongness: it stays green over a defect that was present at capture, and reddens on every intended change.

solid answer

~40 s

An assertion is a claim about **intended** behaviour; a captured value is only a record of **current** behaviour, and drafting at volume swaps the two silently because running the system is the cheapest way to obtain a plausible expectation. The suite that results cannot fail for the reason you want: if the behaviour was already wrong when the drafts were taken, the wrongness is now the expected value and is protected by a passing case. It does fail for the reason you do not want - every deliberate change reddens a wide, unrelated-looking band at once, because a whole-payload comparison observes far more than any case meant to check. The cheapest repair, re-capturing the expectations, re-pins the suite to the new current behaviour. The suite ends up tracking the system instead of judging it.

code

pseudocode · 11 lines
pseudocode
# pinned: the expectation is whatever ran on the day the draft was made
expected = capture(quote_for(order))
assert quote_for(order) == expected            # passes even if it was wrong then

# anchored: the expectation is derived from the rule the requirement states
subtotal = sum(line.price * line.qty for line in order.lines)
expected_total = subtotal - discount_for(order.tier, subtotal)

result = quote_for(order)
assert result.total == expected_total          # fails when system and rule disagree
assert result.currency == order.currency      # nothing else is claimed here

go deeper

for a junior

Recall the difference between what a system did and what it was supposed to do, and be ready to say where the expected value in a case you wrote actually came from.

for a middle

Explain the mechanics: an expectation harvested from a running system freezes current behaviour, so a defect present at capture becomes the expected value and every intended change reddens a wide band of cases at once.

for a senior

Show the production judgement. You can recognise a suite that tracks rather than judges from its failure pattern alone, and you treat bulk re-recording as a decision that destroys evidence rather than as routine repair.

for a principal

Own the standard for where expectations may come from. Decide when captured values are permitted at all, and make derivation from the stated requirement the default that drafted work has to meet before it counts.

## Two claims that look identical on the page An assertion is a **claim about intended behaviour**: this input should produce this outcome, because the stated rule says so. A captured value is a **record of current behaviour**: this input did produce this outcome, on the day somebody looked. Written into a case file the two are indistinguishable - both end up as an expected value sitting beside an actual one. The entire difference is **where the expected value came from**, and that provenance is exactly what the file does not record. Bulk drafting swaps them silently, because the cheapest way to produce a plausible expectation at volume is to run the system and write down what it said. The suite that results does not judge the system. It **tracks** it. ## Where the pinning gets in - A draft is produced by exercising the running system and recording its response, then comparing the whole payload on every later run. - A draft is produced from the implementation rather than from the requirement, so it restates what the code computes instead of what the rule demands. - A first run fails and the expectation is edited until it matches the actual value - a repair that always terminates and always pins. - A stored snapshot is refreshed as a routine step whenever the suite goes red, so the expectation is permanently one release behind reality and is never a claim about anything. ## What each kind of expectation actually does | | Expectation captured from the system | Expectation derived from the rule | | --- | --- | --- | | Fails when | any observable output changes | the system disagrees with the rule | | Stays green when | the behaviour was already wrong at capture | never, while that disagreement stands | | Cost of an intended change | a wide band of cases reddens at once | only the cases the rule touches redden | | What a green run means | nothing moved since capture | the rule still holds | Two consequences follow, and a pinned suite exhibits both at the same time. **It cannot see an existing defect.** If a total was miscalculated on the day the drafts were made, the wrong total is now the expected value, and the case passes on it every night. Worse, the case actively defends the defect: correcting the calculation turns the suite red, and that red reads like a regression to anyone who was not there. **It fails for the wrong reason.** Any intended change to observable behaviour trips every pinned case that happens to observe it, which is usually a wide and unrelated-looking band, because a whole-payload comparison observes far more than any single case meant to check. The failures carry no information about which of them was legitimate. ## The pattern you can see from outside Three symptoms, roughly in the order they appear: 1. **Wide, synchronous failure bands.** One small deliberate change reddens dozens of cases across unrelated areas, and the differences all involve fields nobody in the change had thought about. 2. **Repair by re-capture.** The standard fix becomes regenerating expectations in bulk. It takes minutes and it always works, which is precisely why it becomes the habit. 3. **A suite that has never found anything.** Over months the red runs correlate with the team's own changes rather than with defects, and no failure has ever been the first sign of a bug. The second symptom is the destructive one. Bulk re-capture is not a repair; it re-pins the suite to behaviour nobody verified, and it discards the single moment at which the failing set was evidence. The failing set after a deliberate change is a map of what that change really touched, including the places it should not have touched. Regenerating expectations burns the map before anyone reads it. ## Re-anchoring an expectation The repair is to give the expected value a source that is not the system under test. - Compute it inside the case from the rule the requirement states, so the case fails when the system and the rule disagree. - Assert the **invariant** rather than the literal: the total equals the sum of the lines less the stated discount; the identifier returned when a record is created is the one returned when it is read back; the count after two additions is two greater than before. - Assert the **transition** rather than the snapshot - what changed as a result of the action, not the whole shape of the response around it. - Check only the fields this case is about and leave the rest unchecked, so an unrelated field changing cannot redden a case that never meant to claim anything about it. ## Where a captured value is still legitimate There are honest uses. Where the value **is** the specification and its source of truth sits outside the system - a fixed protocol code, a published rate table, wording supplied verbatim by whoever owns the words - writing it into the case is authoring, not echoing. The test is one question about provenance: **could this expected value have been written before the system existed?** If yes, it is a claim. If it could only have been obtained by running the system, it is a recording, and a suite of recordings certifies whatever it was shown first.

  • Eighty pinned cases have just reddened after a deliberate pricing change. What is the wrong move?
    Re-capturing all eighty expectations from the new behaviour. It clears the board in minutes and restores exactly the condition that made the failures uninformative, because the suite is pinned again, now to behaviour nobody has checked either. That failing set is a map of what the change actually touched, and it is the only chance to ask which of the eighty should have moved.
  • Is there anywhere that asserting a captured value is the right choice?
    Yes, where the value itself is the requirement and its source of truth sits outside the system - a fixed protocol code, a published rate table, wording supplied verbatim by whoever owns it. The distinction is provenance: an expected value authored from the specification is a claim, while one harvested from the running system is an echo of it.

A contract and a receipt look similar on paper. The contract says what was supposed to happen; the receipt only says what did.

saying these in an interview costs you the question

  • Says a passing case proves the behaviour is correct
  • Captures expected values from the running system and calls it a baseline
  • Repairs a wide failure band by re-recording every expectation
  • Cannot say where an expected value originally came from
  • Treats every red case after a deliberate change as noise