skip to content

Why does a stored manual test case carry an expected result on every numbered step rather than one expected result for the whole case?

level: middleimportance: must knowfreq 62%

answer

  1. instruction versus check
  2. somewhere to point at
  3. row seven, not just the case
  4. "works correctly" asserts nothing
  5. cost is staleness per assertion

basics

~20 s

Per-row expectations make each step independently checkable, so a divergence can be attached to the action where it happened instead of to the case as a whole. A single case-level expectation pushes that detail into free text nobody can search.

solid answer

~50 s

An action alone is an instruction; it becomes a check only when something states what should be observed afterwards. Putting that statement on each row makes the row a complete small test — do this, see that — and gives anything later a specific place to point. With one expectation for the whole case, a divergence at row seven and a divergence at row two look identical from outside, and the executor has to explain the difference in a free-text comment that no report can read. Per-row expectations also make the case reviewable: someone can check the logic row by row before it is ever run. The cost is real — more to write, more to keep current — which is why an expectation should be added where a failure would mean something distinct, not on every row by reflex.

go deeper

for a junior

Be ready to say that each numbered row states what should be seen after its own action, and that a row without such a statement is an instruction rather than a check.

for a middle

Explain the mechanics: a per-row expectation gives a divergence somewhere specific to attach to, and one case-level expectation forces the same information into unsearchable free text.

for a senior

Demonstrate the tradeoff you have lived with. More assertions mean more maintenance, so show how you pick the rows whose divergence carries distinct information rather than asserting everywhere.

for a principal

Own the argument that this field is worth defending in review. Without it, downstream outcome modelling and reporting can never say where behaviour diverged, because the definition named no place for it.

## The expectation is what makes a row a test An action on its own is an instruction: *submit the form*. It becomes a **test** only when something states what should be observed afterwards. In a stored manual case, each numbered row carries that statement for itself, so the row is a complete small check: do this, see that. That is not a formatting preference. It decides how much a later reader can learn. When every row carries its own expectation, a disagreement attaches to the row where it happened, and the case's history says *where* behaviour diverged rather than merely *that* it did. When the expectation lives only at the bottom, that information was never captured in a structured place, and no amount of downstream tooling can recover it. ## What a single case-level expectation costs - **Location is lost.** Every divergence reads as "this case did not do what it should", whether it broke at the first action or the last. - **Free text takes over.** The executor writes a paragraph explaining where things went wrong. That paragraph is unsearchable, inconsistent between people, and often missing when the same case diverges again months later. - **Partial progress has nowhere to live.** "Rows one to six behaved, row seven did not" is genuinely useful information, and a definition with no per-row expectations gives it nothing to hang on. - **Review gets harder.** A reviewer cannot check the logic of a case whose only assertion is a closing sentence; they have to simulate the whole flow in their head to find out what the author assumed. ## What makes an expectation checkable 1. **Observable.** It names something a person can actually see or read on screen, in a record, or in a message — not an internal state nobody can inspect from where they are standing. 2. **Singular.** One row, one claim. "The order appears in history and the confirmation message arrives" is two checks that will diverge separately and should sit in two rows. 3. **Independent of the executor's judgment.** "Works correctly", "behaves as expected", "looks right" are not expectations. Two competent people can read the same screen and reach opposite conclusions, which means the row proves nothing. 4. **Written in the domain where possible.** "The basket total shows the discounted price" survives a redesign. "The badge in the top right turns green" is invalidated by a change that broke nothing. ## Rows that genuinely have nothing to expect Not every action produces an observation worth naming. A row that only navigates somewhere has two honest treatments: fold it into the following row's action, or write the smallest true observation, such as which page is now shown. What to avoid is filling the column with "n/a" over and over, because that trains executors to stop reading the column at all — and then the rows that do carry a real expectation get skimmed with the rest. ## The cost, stated honestly Per-row expectations are more work to write and more work to maintain. A behaviour change touches the row that asserts it, and a case with fifteen assertions has fifteen chances to go stale and start producing arguments about whether the case or the product is wrong. Teams that feel this pain usually have the **granularity** wrong rather than the shape wrong: they are asserting on rows whose failure would not mean anything different from the failure of the row above. A case is not improved by asserting more. It is improved by asserting the things whose divergence carries distinct information. ## Where the definition stops The definition supplies the expectations. What an execution then does with them — which outcome values the product offers, how a row-level outcome relates to the case-level one — is a separate concern decided by the product and the team's conventions. The point about the definition is narrower and comes first: > If the rows carry no expectations, nothing downstream can say where the behaviour diverged, because the definition never named a place where it could. That is why this field is the one part of the case anatomy that is worth defending in review. A vague title is annoying; a missing precondition is recoverable by a knowledgeable executor. An empty expectation column turns a repository of tests into a repository of instructions.

  • A row's expected result reads "the page behaves as expected". What do you tell the author?
    That the row currently proves nothing, because two competent executors can read the same screen and disagree. Replace it with something observable and singular — a value shown, a message displayed, a record visible — stated in domain terms rather than in terms of the current layout, so it survives a redesign that broke no behaviour.
  • Should every row have an expectation, even a purely navigational one?
    No. A row that only moves the executor somewhere is better folded into the following action, or given the smallest true observation, such as which page is now shown. Filling the column with placeholder text everywhere trains people to stop reading it, which then costs you the rows that carry a real assertion.

Per-step expected results are the shown working of a test case: with only a final answer on the page, a wrong result tells the marker nothing about where the reasoning broke, but with each line written out they can circle the exact one.

saying these in an interview costs you the question

  • Writes one closing expectation for a fifteen-row case
  • Uses "works correctly" as an expected result
  • Packs two independent assertions into one row
  • Fills the expectation column with placeholder text
  • Thinks per-row expectations are only a formatting nicety