Several teams read one Cucumber suite's report - how do you decide between a bundled formatter and building on the Messages stream?
answer
- start from who reads it
- cheap to build, expensive to own
- filtering is not rendering
- history is the honest reason
- archive the stream regardless
basics
~20 sStart from the audience and the decision the report drives. Bundled formatters and tag-filtered runs cover most multi-team needs for free. Build on the message stream only when history, merging or per-team slicing survives those cheaper options.
solid answer
~50 sOrder the options by cost, not by appeal. The built-in `html` formatter is free, self-contained and per-run, but it has no history, no cross-run merge and no per-team slice. Cheaper than code: run the suite tag-filtered so each team gets its own report file, or adopt a reporting product that already does trends. Building your own consumer is justified when a need survives those: **merging several runs into one view**, **trend history the product does not model**, or a **per-audience slice** nobody sells. The cost is real and ongoing: you own a consumer of an evolving message schema, a build step and somewhere to host the output. Before any of it, ask what decision the report drives. A nightly run nobody reads is not fixed by a better renderer, and building one is an expensive way to avoid that conversation.
code
python · 12 linesimport json
from collections import Counter
status = Counter()
with open("target/cucumber-messages.ndjson", encoding="utf-8") as stream:
for line in stream:
envelope = json.loads(line)
finished = envelope.get("testStepFinished")
if finished:
status[finished["testStepResult"]["status"]] += 1
print(status)go deeper
Recall that the bundled formatters cover an ordinary run and that custom reporting is a choice with a cost, not a default step. Know the message stream is what any custom report would read.
Be able to compare the options concretely: what a per-run HTML page cannot do, what a tag-filtered run solves, and why merging or trending across runs is the case that needs a consumer.
Show that you would diagnose readership before rendering. Name the ongoing costs of a consumer you own, and insist the raw stream is archived so decisions stay reversible.
Own the standard across teams: one archived artefact, one agreed formatter set, per-audience slices before bespoke code, and a review date with a readership signal that can retire the thing you built.
This is a build-versus-buy decision with an unusually cheap build, which is exactly what makes it dangerous. Consuming Cucumber's message stream is a loop over lines of NDJSON, so the prototype takes an afternoon and the maintenance takes years. ## What each option actually gives you | Option | Audience | History across runs | Ongoing cost | Control | |---|---|---|---|---| | built-in `html` | anyone with the file | none, it is one run | zero | full, it is a local file | | JUnit XML into CI | the CI test tab | whatever CI keeps | zero | full, but structure is lost | | tag-filtered runs, one report each | one per team | none | one CI job per slice | full | | a third-party reporting product | everyone, one site | yes, that is its selling point | an upgrade treadmill | its model, not yours | | hosted publish | anyone with the link | limited | zero | least, data leaves you | | your own message consumer | exactly who you design for | whatever you build | a service you own forever | total | Read that table top to bottom before writing code. Most "we need a custom report" requirements are actually "team B does not want to scroll past team A's 180 scenarios", and a tag-filtered run answers that for the price of a CI job. ## The questions that decide it 1. **Who opens it, and what do they decide?** If nobody can name a decision, no report format fixes anything. 2. **Does the need span runs?** Bundled formatters model exactly one run. Trends, flakiness rates and week-over-week comparisons are inherently multi-run and are the honest reason to build. 3. **Does it span suites?** Merging a Cucumber-JVM suite and a cucumber-js suite into one view is straightforward on the message stream and impossible with two HTML pages. 4. **Is the slice per-audience?** A product owner wanting only the scenarios tagged for their area is a filtering requirement, not a rendering one. 5. **Would an existing product do it?** Adopting one costs an upgrade treadmill; building costs a service. Both are ongoing; only one has someone else fixing it. ## What owning a consumer really costs - **Schema drift.** The message format evolves. Your consumer must ignore unknown envelope keys and tolerate new message types instead of failing on them, and someone must own it when it stops matching. - **A build step and a home.** A report nobody can reach is not a report. You now run a job that generates it and a place that serves it, with the access control that implies. - **Correctness nobody checks.** A bundled formatter that miscounts gets reported by thousands of users. Yours gets reported by nobody, and a wrong report is worse than no report because people act on it. - **A bus factor of one.** Reporting tools written by an enthusiastic engineer routinely outlive their author's interest in them. ## The failure this usually hides A climbing-gym membership product has a 26-file feature directory and a nightly run at 02:15 that nobody reads any more. The proposal on the table is a custom dashboard built from the message stream. The diagnosis is usually somewhere else: the run reports 41 failures of which 3 are real, the report arrives eight hours after the change that caused it, and no team believes a red result means their code is broken. A dashboard makes ignored data prettier. Cutting the failing-for-environmental-reasons scenarios, moving the run onto the change that caused it and giving each team its own tag-filtered slice fixes readership. Then, if a trend question remains, build for it. ## A staged path 1. Always emit and archive the raw message stream, from day one, whatever else you do. It costs one formatter entry and it is the option value for everything below. 2. Split by audience with tag-filtered runs before writing any code. 3. If trends are the real need, evaluate an existing reporting product against that need specifically. 4. Only then write a consumer, and keep it small: aggregate the archived streams into a dataset and render that, rather than replacing the per-run report that already works. 5. Set a review date and a readership signal. A custom report that nobody opens for a quarter should be deleted, not maintained. The principal-level position is not "never build". It is that the message stream makes building so cheap that the discipline has to come from the audience question rather than from the effort.
- What is the cheapest thing to try before building anything, when two teams want different views?Run the suite twice with different tag filters and give each team its own report file. It costs a CI job, uses only bundled formatters, and answers the real complaint, which is usually scrolling past another team's scenarios rather than a missing visualisation.
- How do you keep a custom consumer from breaking when Cucumber evolves?Treat the message format as an external contract: read defensively, ignore envelope keys and message types you do not recognise, and never assume a field you have not checked for. Archive the raw NDJSON so a report can be rebuilt after a fix, and keep a bundled formatter running alongside as the fallback everyone can still open.
- What would make you retire a custom report you had already built?No readership. Instrument it, and if the people it was built for have not opened it in a quarter, delete it rather than carry it. The same test applies to the run behind it: an unread nightly report and an unread nightly run are usually the same problem.
saying these in an interview costs you the question
- Builds a custom report before naming who reads it
- Treats a prettier report as the fix for an ignored run
- Consumes the legacy JSON output instead of the message stream
- Archives only rendered HTML, so nothing can be rebuilt
- Assumes the message schema will never change
- Ignores that a wrong custom report is worse than none