A generated specification report renders scenarios that never executed as documented behaviour. How do you fix it?
answer
- More than two outcomes per scenario
- Only passed and failed are evidence
- Say which revision and which selection
- Counts on the front page
- Refuse to publish an overstatement
basics
~20 sRender every scenario with its run status and make not-executed visually distinct, stamp the revision, environment and any tag selection on the header, publish executed-versus-total counts, and fail the publishing step when too much of the suite did not run.
solid answer
~50 sThe report is trusted because readers assume its pages describe a real run, so any scenario rendered as ordinary prose without a pass or fail behind it is an unsupported claim. Start by enumerating the states a run produces: passed, failed, steps not implemented, skipped after an earlier failure, filtered out by a tag selection, or newer than the results file. Only passed and failed are evidence; everything else must render as a distinct not-executed state, never as neutral text. Then stamp provenance on the artefact — revision, finish time, environment, and the selection the run applied — so a reader can see it describes a slice. Put executed-versus-total counts on the front page, and fail the publish step when the not-run share crosses a small threshold: a missing document for a day beats an overstated one for a month. Never carry a previous run's green status forward as current.
code
pseudocode · 22 linesrun = load_last_run()
all_s = load_scenarios_from_repository()
executed = [s for s in all_s if run.status_of(s) in ["PASSED", "FAILED"]]
not_run = [s for s in all_s if s not in executed]
report.header = {
"revision": run.revision,
"finished_at": run.finished_at,
"environment": run.environment,
"selection": run.tag_expression,
"executed": len(executed),
"total": len(all_s),
}
for s in not_run:
report.page(s).status = "NOT EXECUTED" # never rendered as plain prose
if len(not_run) * 100 / len(all_s) > 2:
fail_publish("report would describe scenarios that did not run")
publish(report)go deeper
Know that a scenario can end a run in more states than pass or fail — not implemented, skipped, or simply not selected — and that a report should show which happened.
Explain how a selection expression shrinks a run, and why a report generated from a subset must say so in its header rather than presenting itself as a picture of the whole system.
Show that you would rather publish nothing than publish an overstatement: a publish gate on the not-run share, provenance on the header, and a hard rule against carrying an older run's status forward.
Own the trust contract. Decide what the published artefact is allowed to claim, who is accountable when it overstates the system, and how that shows up in an audit or a partner conversation.
### The failure A generated specification report is trusted because readers assume its pages describe what the system did on a real run. That assumption breaks the moment the report renders a scenario the run never executed with the same visual weight as one that passed. The reader sees prose describing behaviour; nothing verified that behaviour; and the document is now confidently asserting something no evidence supports. This is worse than a stale hand-written page, because the artefact carries an implicit claim of freshness. Scenarios end a run in more states than pass and fail: - **Passed** — every step executed and the assertions held. - **Failed** — a step raised or an assertion did not hold. - **Not implemented** — the scenario text exists but one or more steps have no automation behind them, so the runner reports them as undefined or pending. - **Skipped** — an earlier step failed, or a precondition hook aborted the scenario. - **Filtered out** — the run selected a subset by tag, and this scenario was never a candidate. - **Absent** — the scenario is in the repository, but the results file the generator read predates it. Only the first two are evidence. Every other state must be visible in the published document. ### The fixes, in order of leverage **1. Render status per scenario, and make "not executed" loud.** The generator should have no default that lets an unknown status render as ordinary prose. Treat any scenario without a pass-or-fail result as a first-class, visually distinct state — not-run — never as neutral body text. A scenario with undefined steps in particular is frequently mistaken for a passing one, because the text reads perfectly. **2. Stamp provenance on the artefact.** The header should carry the revision, the finish time, the environment, and the selection expression the run used. A reader who can see "generated from a run filtered to the smoke tag" knows immediately that the document describes a slice, not the system. A report with no provenance cannot be audited and cannot be trusted a week later. **3. Publish the counts, not just the pages.** Total scenarios in the repository, executed, passed, failed, not run. If those five numbers are on the front page, the misreading is much harder to sustain. **4. Gate the publish step.** If more than a small share of scenarios did not execute, the publishing step should fail rather than emit a misleading artefact. A document that is missing for a day is a smaller problem than a document that overstates the system for a month. **5. Never carry a previous run's status forward.** The tempting optimisation — "this scenario was not in today's filtered run, so show yesterday's green" — is exactly how a report ends up asserting behaviour on a revision that never demonstrated it. If it must be shown, label it with the revision it last passed on and mark it as not current. ### A worked example The four-person team on the parcel-tracking gateway runs a fast subset on every commit, filtered by tag, and the full suite nightly. Their published report was generated from the *commit* run. Of 214 scenarios in the repository, 118 were candidates under that filter; the remaining 96 were rendered as ordinary pages with no status marking. Among those 96 was the scenario for serving a status lookup after a cache expiry — which, on the nightly run, had been failing for eleven days because of a stale-cache read. The published specification of record was describing behaviour the system had not exhibited for a week and a half, and the document gave the reader no way to know it. The fix they shipped: generate the published artefact only from a full run; stamp the revision and the filter on the header; and fail the publish step when the not-run share exceeds a small threshold. ### The boundary worth stating Note what this is not. Counting how often the suite is red over time, or how long the pipeline takes, is a different conversation about suite and pipeline health. This is narrower and about the *document*: whether a reader can distinguish what the last run demonstrated from what it merely described. Getting that distinction right is what makes the artefact usable as a specification of record rather than as decoration. ### How to answer well Enumerate the statuses; say that anything other than passed or failed is not evidence and must render distinctly; insist on provenance in the header; and offer the publish gate as the mechanism that stops a misleading artefact from shipping at all. Mentioning that you would rather publish nothing than publish an overstatement is a strong senior signal.
- A scenario passed on last night's full run but was outside today's tag selection. What should today's report show?Not-executed for this revision, optionally annotated with the revision it last passed on. Carrying yesterday's green forward as current status is how a report ends up asserting behaviour on a build that never demonstrated it. The annotation is fine; the false claim of currency is not.
- Should the publishing step fail the pipeline when too many scenarios did not execute?Yes, with a threshold the team sets deliberately and reviews. The trade is a temporarily missing document against a published artefact that overstates the system, and the second is worse because readers cannot tell it is wrong. Failing loudly also stops a broken selection expression from quietly shrinking the report.
- Why is a scenario with unimplemented steps particularly dangerous in a published report?Because the prose reads perfectly. Nothing about the sentence signals that no automation stands behind it, so a reader takes it as verified behaviour when it is a wish. Render it as a distinct state and, if the team is disciplined, treat unimplemented steps as failing rather than neutral.
saying these in an interview costs you the question
- Renders unexecuted scenarios the same as passing ones
- Publishes a report with no revision or environment stamp
- Carries a previous run's green status forward as current
- Generates the published artefact from a tag-filtered subset silently
- Treats unimplemented steps as neutral rather than not verified
- Prefers publishing something misleading over publishing nothing