skip to content

Which statuses beyond pass and fail should an automated suite record for a case, and what does each let a reader decide?

level: middleimportance: must knowfreq 62%

answer

  1. Not every non-pass is a failure
  2. Did it run, or did it not?
  3. Each status enables a different next action
  4. Blocked, skipped, known-issue, flaked
  5. A reason code beside every status

basics

~20 s

Record at least skipped, blocked, known-issue and flaked alongside pass and fail. Each answers a different reader question: was the case deliberately not run, prevented from running, failing for an accepted reason, or unreliable and not evidence of anything?

solid answer

~40 s

A two-value model forces every non-pass into `FAILED`, so a reader cannot tell a genuine regression from a case the machinery never let start. Four extra statuses carry most of the value: **skipped** (the suite chose not to run it), **blocked** (it could not run — an unmet precondition, a setup step that never completed, an unreachable target), **known-issue** (it ran and failed for an accepted, recorded reason) and **flaked** (it produced both outcomes for the same code, so it is evidence of nothing). The bar for adding a status is decision value: if two statuses always lead a reader to the same next action, collapse them into one. Each non-pass also needs a reason code beside it, because `BLOCKED` with no stated cause is only slightly more useful than `FAILED`.

code

pseudocode · 17 lines
pseudocode
STATUS = { PASSED, FAILED, BLOCKED, SKIPPED, KNOWN_ISSUE, FLAKED }

record_attempt(case_id, attempt_index, outcome, reason_code, detail):
    require reason_code is present when outcome != PASSED
    attempts[case_id].append({
        index:  attempt_index,
        outcome: outcome,
        reason:  reason_code,       # "target-unreachable", "filtered-out", "accepted-3417"
        detail:  detail             # one human sentence
    })

derive_case_status(case_id):
    outcomes = distinct(attempts[case_id].outcome)
    if outcomes == { PASSED }              -> PASSED
    if PASSED in outcomes and FAILED in outcomes -> FLAKED    # never PASSED
    if outcomes == { BLOCKED }             -> BLOCKED
    otherwise                              -> worst_of(outcomes)

go deeper

for a junior

Be ready to say that a suite records more than pass and fail, and that a case which never ran is not the same as one that ran and failed. Know the words skipped and blocked and roughly what each claims.

for a middle

Explain what each status asserts and what a reader does next because of it. Expect to be asked why a reason code has to sit beside the status, and what goes wrong when everything that is not a pass is recorded as a failure.

for a senior

Show that you have watched a vocabulary decay in a real suite: blocked used to hide regressions, an unreliable case recorded as a pass, statuses nothing branches on. Say how you would derive a case's status from its attempts rather than writing it directly.

for a principal

Own the vocabulary as a shared contract. Argue for the smallest set that supports distinct actions, insist each value is defined in writing, and explain how you stop two suites in the same organisation from giving one status opposite meanings.

## Why pass and fail is not a vocabulary A result model with two values forces an implicit claim on every row: anything that is not a pass is a failure of the product. That claim is false often enough to be expensive. A case that could not start because the target it exercises was unreachable did not test anything. A case filtered out of this run tested nothing either. A case failing for a reason already recorded and accepted tells the reader nothing new. And a case that produced two different outcomes for the same code is not evidence in either direction. Collapsing all four into `FAILED` costs the reader the one thing a results document exists to give them: **a next action they can take without opening the run**. The reader's question is almost never "how many failed". It is "is there something here I have to do, and what is it". ## The statuses that earn a place The bar for adding a status is **decision value**: two statuses that always lead a reader to the same next action should be one status. | Status | What it asserts | What the reader does next | |---|---|---| | `PASSED` | The case ran and its checks held | Nothing | | `FAILED` | The case ran and a check did not hold | Investigate the change to the product | | `SKIPPED` | The suite chose not to run it | Confirm the choice is still intended | | `BLOCKED` | It could not run — an unmet precondition, a setup step that never completed, an unreachable target | Fix the surroundings, not the product | | `KNOWN_ISSUE` | It ran and failed for an accepted, recorded reason | Nothing new; the reason already has an owner | | `FLAKED` | It produced both outcomes for the same code | Distrust it; it is not evidence | The two most valuable are the two most often missing. `BLOCKED` separates **"the product is wrong"** from **"the machinery around the case is wrong"**, and those two findings go to different people on different days. `SKIPPED` separates a deliberate omission from an accidental one — but only if the model also records *why* it was skipped. ## A status without a reason is barely better than a failure Every non-pass needs a machine-readable **reason code** beside it, plus one human sentence: - `SKIPPED` — which filter or unmet precondition excluded it - `BLOCKED` — what was unavailable, and at which stage the case stopped - `KNOWN_ISSUE` — the reference to the accepted entry, so a reader can check it still applies - `FLAKED` — how many attempts ran, and what each of them produced The reason code is what lets a reader group two hundred non-passes into three causes instead of reading them one at a time. It is also what stops `BLOCKED` from quietly becoming the bucket a real regression hides in: if every blocked row must name what was unavailable, a row that cannot name anything stands out. ## Deriving a case's status from its attempts Where the model records attempts, the case-level status is **derived**, never written directly: 1. Every attempt carries its own outcome and its own evidence. 2. The case-level status is computed from the list of attempts. 3. A case whose attempts disagree is `FLAKED`, not `PASSED` — the disagreement *is* the finding. Overwriting the first attempt's outcome with a later one is the most common way a result model lies to its readers. It turns a suite that is quietly unreliable into a suite that looks green, and it does so without anyone deciding to hide anything. ## Keeping the vocabulary small and shared Three failure modes show up in real suites: - **Vocabulary inflation** — a dozen statuses, half of which nothing branches on. Delete any status no consumer distinguishes; it is costing every reader a decision and paying nobody. - **Vocabulary drift** — two suites in the same organisation use `BLOCKED` for opposite things, so a combined view is meaningless. Write the definitions down once, beside the model, and treat a change in a value's meaning as a change to a contract. - **Status as policy** — the model records the fact. Who owns an accepted failure and when it expires is a separate policy question and does not belong inside the status vocabulary. A good check on a proposed vocabulary: take one row out of the results document with the run stripped away, hand it to someone who has not seen the suite, and ask what they would do next. If the status plus its reason code answers that, the vocabulary is doing its job. If the honest answer is "I would have to open the run", the vocabulary is decoration.

  • A case stopped because the target it exercises was unreachable. Why record that as blocked rather than failed?
    Because the two send a reader to different people. A failure claims the product changed and someone should look at the change; blocked says the surroundings the case needed were absent, so the fix is in the environment or the setup. Recording it as a failure also inflates the failing set with rows nobody can act on, which trains readers to skim the list instead of reading it.
  • How do you stop a status vocabulary from growing into a dozen values nobody agrees on?
    Add a status only when a reader would take a genuinely different next action because of it, and delete any status no consumer branches on. Write each value's definition down beside the model, together with the action it implies. Push everything else into the reason code, which can stay open-ended without forcing every reader to learn another value.
  • A case failed on its first attempt and passed on a second within the same run. What must the result model preserve?
    Both attempts, each with its own outcome and its own evidence, and a case-level status derived from the pair rather than overwritten by the last one. Discarding the first attempt turns a case that disagreed with itself into a clean pass, destroying the only interesting signal. Whether attempting again is permitted at all is a separate decision from how the run records what happened.

saying these in an interview costs you the question

  • Records every non-pass as a failure
  • Uses blocked as a bucket that hides real regressions
  • Writes a status with no reason code beside it
  • Overwrites a first attempt's outcome with a later one
  • Invents statuses no reader ever branches on
  • Says skipped and blocked mean the same thing