skip to content

What does a single line of Cucumber's Messages NDJSON stream represent, and what consumes that stream?

level: middleimportance: should knowfreq 37%

answer

  1. one line, one thing
  2. newline-delimited, not one big array
  3. each envelope wraps a single message
  4. a pickle per Examples row
  5. formatters all subscribe to it

basics

~20 s

Each line is one JSON envelope holding exactly one message: a parsed feature file, a compiled runnable scenario, a step result, an attachment. Cucumber's built-in HTML, JSON and JUnit formatters all consume that same stream rather than writing independently.

solid answer

~40 s

Cucumber emits **Cucumber Messages** as NDJSON, one JSON object per line, each an *envelope* wrapping exactly one message. Early lines describe the run: `meta` (environment), `source` (the raw `.feature` text) and `gherkinDocument` (its parsed tree). Then comes one `pickle` per runnable scenario, which is where each `Examples` row becomes its own item. Execution streams `testCaseStarted`, `testStepFinished` (carrying a `status` such as `PASSED`, `FAILED` or `UNDEFINED`, plus a duration), `testCaseFinished` and finally `testRunFinished`; `attachment` messages carry whatever the run captured. The architectural point is that the built-in `html`, `json` and `junit` formatters all subscribe to this one stream, so no report can contain anything the messages do not. Choosing the `message` formatter simply writes the raw stream to a file.

code

json · 3 lines
json
{"pickle":{"id":"7f1c","uri":"features/membership_renewal.feature","name":"Renewing a lapsed climbing-gym membership","language":"en","steps":[{"id":"a1","text":"the member lapsed 43 days ago"}]}}
{"testCaseStarted":{"id":"c9","testCaseId":"t3","attempt":0,"timestamp":{"seconds":1757116800,"nanos":0}}}
{"testStepFinished":{"testCaseStartedId":"c9","testStepId":"s1","testStepResult":{"status":"PASSED","duration":{"seconds":0,"nanos":41000000}},"timestamp":{"seconds":1757116800,"nanos":41000000}}}

go deeper

for a junior

Recall that Cucumber can write a machine-readable record of a run as newline-delimited JSON, and that the HTML report you open is generated from it rather than typed by the runner.

for a middle

Be ready to name the main message types in order and explain that a pickle is one runnable scenario, so an Examples row counts as its own item. Explain why formatters are subscribers to one stream.

for a senior

An interviewer expects you to treat the NDJSON file as the archived artefact of record: it answers questions later that no rendered report can, and it survives a killed run as a readable prefix.

for a principal

Own the position that reporting is a consumer problem, not an instrumentation problem. Standardising on one archived stream keeps teams from bolting per-team hooks into shared step definitions.

Cucumber's `message` output is the run's primary record. Everything a report can show is derived from it, so understanding its shape is the difference between configuring reporting and actually controlling it. ## What NDJSON means here **NDJSON** is newline-delimited JSON: one complete JSON object per line, with no enclosing array and no commas between lines. Each line is an **envelope** carrying exactly one message, and the envelope's single key names the message type. That shape is deliberate, and it has three consequences you feel in practice: - A consumer can read the file **line by line while the run is still going**, which is how live formatters render progress. - A run that dies halfway still leaves a **valid, readable prefix** on disk. A file that stops before `testRunFinished` tells you the run was killed rather than that it failed. - Merging two runs is a file concatenation problem, not a JSON surgery problem. ## The message types, in the order they arrive | Message | What that line carries | |---|---| | `meta` | who ran it: the Cucumber implementation, the runtime, the OS, CI details | | `source` | the raw text of one `.feature` file, with its `uri` | | `gherkinDocument` | that file parsed into a tree of `Feature`, `Rule`, `Background`, `Scenario`, `Examples` | | `pickle` | one **compiled, runnable scenario**: a flat list of steps with placeholders already substituted | | `testCase` | the execution plan for a pickle, its steps mapped to step definitions and hooks | | `testCaseStarted` / `testCaseFinished` | one attempt at running a test case; a retry raises `attempt` | | `testStepStarted` / `testStepFinished` | one step's execution; `testStepResult` holds `status` and `duration` | | `attachment` | anything the run attached, with its media type | | `testRunFinished` | the run ended, and whether it was successful | The `pickle` is the line that surprises people. A `Scenario Outline` with an `Examples` table is **not** one runnable item: Gherkin compiles it into one pickle per `Examples` row, with each `<placeholder>` already replaced. So a directory of 26 feature files holding 287 `Scenario:` and `Scenario Outline:` keywords can easily emit 613 pickles, and every count in every report is a pickle count, not a keyword count. When a stakeholder asks "how many scenarios do we have?", the report and the feature files answer differently, and the pickle is why. `testStepResult.status` comes from a fixed set: `PASSED`, `FAILED`, `SKIPPED`, `PENDING`, `UNDEFINED`, `AMBIGUOUS`, `UNKNOWN`. That fixed vocabulary is exactly why a report can distinguish "no step definition matched this line" from "the step ran and threw" without guessing from a stack trace. ## Why the formatters are consumers, not independent writers Cucumber runs an internal event bus. A plugin registers as a subscriber, receives events as the run produces them, and writes whatever it wants. The built-in formatters are ordinary subscribers to that bus, and the `message` formatter is the one that writes the events out verbatim. Three things follow: 1. **The formats are projections of one truth.** The JUnit XML formatter is a lossy projection: it flattens each scenario into a `testcase` element and drops the Gherkin structure, per-step status and attachments. The JSON formatter keeps more but is a legacy shape. Neither can show something the stream never carried. 2. **The built-in HTML report embeds the messages it was rendered from.** The page is a small viewer around the data, which is why it is self-contained and why it can show you per-step results without a server. 3. **Adding a report never means instrumenting your tests again.** You add a consumer, not a hook in every scenario. ## Where this bites in practice A climbing-gym membership product runs its 26-file feature directory nightly. The team wants to know which scenarios got slower over the quarter. There is no formatter for that, and there does not need to be: `testStepFinished` already carries a duration for every step of every pickle in every run. Archiving the NDJSON file per run gives you the entire history as a dataset; archiving only the rendered HTML gives you 90 pages nobody joins together. The same reasoning applies in reverse. If a report is missing a fact, check whether the stream carries it before writing code. Attachment media types, retry attempts, hook execution and the original feature-file text are all in there. ## Reading it yourself Consuming the stream is a loop over lines, a `json.loads`, and a check for the envelope key you care about. There is no library requirement and no schema compiler needed for simple aggregation. The cost of consuming it is small; the cost of *not archiving* it is that a question asked next month cannot be answered about last month's run.

  • Where in the stream would you find the original plain text of a feature file?
    In the `source` message, which carries the raw file contents and its `uri`. The parsed structure arrives separately as `gherkinDocument`, and the runnable form arrives later as pickles. That separation is why a report can display the author's original Gherkin even though the runner executed compiled pickles.
  • Why does one Scenario Outline with nine Examples rows produce nine pickles?
    A pickle is a compiled, runnable scenario with placeholders already substituted, so each `Examples` row compiles into its own pickle with its own values. The runner never executes an outline; it executes pickles. Report counts, filtering and per-item results therefore all work at row granularity.
  • Why is the file written incrementally rather than assembled at the end?
    Messages are emitted as events happen, so a live formatter can render progress and a crashed run still leaves everything up to the crash on disk. A truncated file missing `testRunFinished` is itself diagnostic: it says the process died rather than that scenarios failed.

The stream is a flight recorder and each report is one way of playing the tape back; picking a different formatter changes the playback, never what was recorded.

saying these in an interview costs you the question

  • Calls the messages file one big JSON array of results
  • Thinks the HTML formatter parses the JSON formatter's file
  • Believes one pickle exists per outline, not per Examples row
  • Assumes nothing is written until the run finishes
  • Confuses the local message stream with a published report link