skip to content

Why does a failure record need both the actual-versus-expected pair and the log of steps that led there?

level: juniorimportance: should knowfreq 52%

answer

  1. One says what, the other says how
  2. A discrepancy is not an explanation
  3. The log carries outcomes and durations
  4. Earlier steps may have quietly not worked
  5. Together they remove the need to re-run

basics

~20 s

The value pair says what is wrong — the discrepancy itself. The step log says how the run got there — which interactions ran, what each returned, and how long they took. Neither answers the other's question.

solid answer

~50 s

They answer different questions, and a failure needs both. The **actual-versus-expected pair** localises the discrepancy: which value is wrong and by how much — a missing field, a count that is one short, a total off by a factor of a hundred. What it cannot say is why that value is what it is. The **interaction log** — the ordered record of the calls and screen actions the case performed, each with its response and its duration — supplies the history: whether the earlier steps actually did what the case assumed, whether a setup call quietly returned an error the case ignored, whether something took far longer than usual. With only the pair you have the symptom; with only the log you have a sequence but no statement of what broke. Together they turn a red case into a diagnosis nobody has to re-run for.

code

yaml · 10 lines
yaml
case: checkout-applies-percentage-discount
failure:
  observation: order total in minor units
  expected: 4275
  actual: 4750
steps:
  - { action: create cart,          ok: true,  ms: 41 }
  - { action: add item,             ok: true,  ms: 63 }
  - { action: apply discount code,  ok: false, ms: 812, note: stand-in returned an error the case did not check }
  - { action: read order total,     ok: true,  ms: 55 }

go deeper

for a junior

Be ready to name the two halves of a useful failure record: the value that was expected against the value that arrived, and the ordered steps the case ran before it. Say what each one lets you rule out.

for a middle

Explain what makes a step log usable — one entry per meaningful interaction, an outcome and an elapsed time on each, scoped so parallel cases do not interleave. Explain why a large comparison should print the diverging path rather than both objects.

for a senior

Show judgement about signal: which interactions earn a line, how you keep the record short enough that somebody reads it, and how you recognise the common case where an earlier step failed silently and the final observation merely reported the consequence.

for a principal

Own the contract rather than the instance. Decide what every failure record in the estate must carry before a case is allowed in, and be ready to defend the logging that produces it against a team that wants faster runs.

## Two records, two different questions A failed case leaves behind two very different kinds of evidence, and engineers routinely argue for one as though it made the other redundant. It does not. The **actual-versus-expected pair** is a statement about a single observation: this value was supposed to be that, and it was this instead. The **interaction log** is a statement about history: these steps ran, in this order, each returning this and taking this long. One localises the symptom. The other supplies the story. Diagnosis is the act of joining them, and a report carrying only one of them forces the reader to invent the other — which is exactly the moment somebody decides to re-run the case and watch. ## What the value pair can and cannot say The pair is precise where prose is vague. A total that is 475 minor units too high, a collection with four entries where three were expected, a field that arrived absent, a value shaped differently from the contract — these are all *exact*, and exactness is what turns a red case into a hypothesis. It also tells you how badly wrong things are, which is often the fastest route to a cause: being off by one is a boundary bug, being off by a factor of a hundred is a units bug, and being off by an entire record is a data bug. What the pair cannot say is anything about how the value came to be. It is a photograph of the finish line. If the third of five setup steps quietly returned an error and the case carried on regardless, the pair reports the consequence and says nothing about the cause. Worse, it can actively mislead: a discrepancy that looks like a calculation bug is very often a setup that did not do what the case assumed. ## What the interaction log adds | Question a reader asks | The value pair | The step log | |---|---|---| | What exactly is wrong? | Yes — the precise discrepancy | No — it records outcomes, not correctness | | Where in the case did it go wrong? | Only at the final observation | Yes — the step whose outcome broke | | Did the setup actually work? | No | Yes — every step carries its own outcome | | Was something unusually slow? | No | Yes — durations per step | | Did the case reach the behaviour at all? | No | Yes — the last step recorded says so | | Which side of the pair should be believed? | Neither — that is the reader's judgement | Neither | The log's value is entirely in its fields. A stream of lines with no outcome and no duration is prose; a line per meaningful interaction with a success flag and an elapsed time is data. Three properties make the difference: - **One entry per meaningful interaction**, not per internal call. The reader wants the case's story, not the application's. - **An outcome on every entry**, so a step that failed and was ignored stands out without being read for. - **An elapsed time on every entry**, because "this normally takes 40 milliseconds and took 812" is a diagnosis in one line. - **Scoped to the case**, so a parallel run does not interleave two stories into one unreadable file. ## The pairing in practice A worked shape, and it is the commonest one in real suites: 1. The report opens with `expected 4275, actual 4750`. That is a discrepancy of 475 — suspiciously like a discount that was not applied. 2. The log shows four steps. Three succeeded in tens of milliseconds. The third, applying the discount, recorded `ok: false` after 812 milliseconds. 3. The case never checked that third step's outcome, so it continued and asserted on a total that was correct *for a cart with no discount*. Neither record alone gets you there. The pair alone leaves you reading the discount calculation, which is fine. The log alone tells you a step failed but not whether it mattered to the outcome the case was checking. Together they take about fifteen seconds. ## Where each one is usually weak The value pair is weakest on large objects. Printing two big structures in full makes the reader perform the comparison the harness has already performed; the useful output is the **path where they diverged**, the two values at that path, and a bounded summary of everything else. A diff that scrolls off the screen is only marginally better than none. The log is weakest on volume. The temptation is to record everything, on the theory that more is safer. More is not safer: an undifferentiated stream at a single level, with no case boundaries, buries the four lines that mattered. If a reader has to search the log rather than read it, the log has failed at the one job the value pair could not do for it.

  • What turns a step log into noise rather than evidence?
    Volume without structure. A log that records every internal call at one level, with no case boundaries, no outcome per step and no durations, forces the reader to reconstruct the story by hand. Keep one line per meaningful interaction, mark whether it succeeded, stamp the elapsed time, and scope it to the case so a parallel run does not interleave two stories.
  • How do you decide what the value pair should show when the compared object is large?
    Show the difference, not both objects. Printing two large structures in full makes the reader redo the comparison the harness already did. Print the path at which they diverged, the two values at that path, and a bounded summary of the rest. A diff that scrolls off the screen is only marginally better than none.

A thermometer reading tells you the patient has a fever; the chart of the last six hours tells you what has been happening to them. Treating either one as the whole picture produces the wrong diagnosis.

saying these in an interview costs you the question

  • Thinks the bare failure text alone is enough to diagnose
  • Logs every internal call at one level with no case boundaries
  • Prints both whole objects instead of where they diverged
  • Treats the step log as optional because the comparison says enough
  • Records steps without outcomes or durations