skip to content

When a manual test case records a result on each step, how is the case-level result derived from those step results?

level: middleimportance: should knowfreq 56%

answer

  1. not a score and not a vote
  2. worst outcome wins
  3. what happens to steps after the failure
  4. the summary can be set without the steps
  5. evidence belongs where it is local

basics

~20 s

The case result is derived by precedence, not by counting: any failed step makes the case failed, otherwise any blocked step makes it blocked, and only an all-passed set makes it passed. Unexecuted steps never count as passes.

solid answer

~50 s

Most repositories roll a case up by precedence rather than arithmetic. A single failed step decides the case, whatever the other steps say; if nothing failed but something was blocked, the case is blocked; the case only passes when every step passed. There is no averaging and no majority vote, because a failure is evidence and evidence is not outvoted. The consequence people forget is that execution usually stops at the first failing step, so the steps after it stay unexecuted - they are unknown on this build, not implicitly fine. Two things to watch for in practice: many tools let a tester set the case-level result directly without touching the steps, which is how a case marked passed ends up sitting over a failed step; and the roll-up loses locality, so the comment and evidence belong on the step that failed rather than on the case.

go deeper

for a junior

Be able to state the precedence rule out loud: one failed step makes the whole case failed, and a case passes only when every step passed. No averaging is involved.

for a middle

Explain why unexecuted steps after a failure are unknown rather than passing, and why setting the item result directly can leave it inconsistent with the step results underneath it.

for a senior

Show how you handle a record whose levels disagree, and describe your rule for when a tester continues past a failing step instead of stopping. Expect to justify re-running from the failure rather than the failing step alone.

for a principal

Own the tradeoff between per-step recording and a single item outcome: the detail costs execution time on every run and pays back only if the team records evidence where it is local and reads the summary as lossy.

A manual case that records an outcome per step has two levels of truth: what happened at each step, and one summary outcome for the item. The summary is derived, and knowing the derivation rule is what lets you read a case correctly - and spot when the two levels disagree. ## Precedence, not arithmetic The rule is a priority order over the step outcomes, applied to the whole case: | If any step is | The case rolls up to | Reasoning | |---|---|---| | Failed | Failed | A failure is an observation, and it cannot be outvoted by passes | | Blocked (and none failed) | Blocked | Something prevented part of the run; the case is not verified | | Passed, all of them | Passed | Every stated expectation was checked and met | | Untested (and nothing else recorded) | Untested | The case has not been exercised in this cycle | There is deliberately **no averaging**. Four passed steps and one failed step is not eighty percent passed; it is a failed case with four checks that happened to hold. Percentages belong to counting items across a cycle, never to deriving one item's outcome. ## Why the steps after a failure stay unexecuted When a tester hits a failing step, they usually stop, because the product is now in a state the remaining steps were not written for. That leaves a tail of steps with no result. The correct reading of that tail is **unknown on this build**: - they are not passes, because nothing was observed - they are not failures, because nothing misbehaved in front of anyone - they are not blocked in the environmental sense, because nothing external stood in the way This matters after the fix lands. A reader who assumes the tail was fine will re-run only the failing step; a reader who understands the tail is unverified re-runs the case from the failure onward. It is also why continuing past a failure is sometimes worth it: if the later steps check behaviour genuinely independent of the failed check, running them turns unknowns into evidence for the price of a few minutes. ## When the two levels disagree The most common inconsistency is a case-level result that does not match its steps - a case reading passed while one of its steps reads failed. That is almost never a roll-up bug. It usually means: 1. **The case result was set directly.** Many tools let a tester record an item outcome without opening the steps, and setting the summary does not propagate downward. 2. **Steps were edited after the run.** The step list changed under a stored result, so the recorded step outcomes no longer line up with the steps as they now read. 3. **A later attempt updated the item but not each step**, leaving stale step results underneath a fresh summary. Whatever the cause, treat the disagreement as a signal that the record is not trustworthy and go back to the evidence rather than to either status. ## What the roll-up loses, and how to compensate The summary answers whether the case is good; the steps answer where it went wrong. Rolling up throws the second away, so the recording discipline has to put it back: - **Attach evidence at the step**, not at the case, when the tool supports it. A screenshot filed against a fifteen-step case makes the reader hunt; the same file against the failing step is self-locating. - **Put the actual-versus-expected pair on the failing step**, because the expectation being contradicted is that step's expectation, not the case's. - **Say in the comment where execution stopped**, so the unexecuted tail is visibly unknown rather than ambiguous. - **Do not restate the whole case in the case-level comment.** The step results already carry the sequence; the summary comment is for what the steps cannot say, such as the build, the data or the account used. ## Practical rules that hold across products 1. Read the case status as a **worst-outcome** summary, never as a score. 2. Treat unexecuted trailing steps as unverified, and re-run from the failure rather than only the failing step. 3. If the item status and the step results disagree, believe the evidence, not either status word. 4. Record at the level where the information lives - step-level detail on the step, run-level context on the case. The general point is that a per-step case is a small structured record and the item outcome is a lossy view of it. The view is what everyone reads; the structure is what makes the view recoverable when someone needs to know what actually happened.

  • When is it worth continuing past a failing step rather than stopping the case there?
    When the later steps check behaviour independent of the failed one and the product is still in a usable state. You convert unknowns into evidence cheaply and may find a second, unrelated defect in the same pass. Stop instead when the failure leaves the system in a state the remaining steps were never written for, since anything you observe afterwards is unreliable.
  • A case reads passed while one of its steps reads failed. What is your first assumption?
    That the case-level result was set directly without touching the steps, since setting a summary does not propagate downward in most tools. The other likely causes are steps edited after the run, or a later attempt updating the item but not each step. In all three, trust the attached evidence over either status word.

saying these in an interview costs you the question

  • Averaging step outcomes into a case percentage
  • Reading unexecuted trailing steps as passes
  • Assuming the case status always matches its steps
  • Attaching all evidence at case level by habit
  • Re-running only the failing step after a fix