skip to content

Format Limits and Dialects

What this format cannot say, and what each writer does about it: no official schema, no step tree, no attachment channel, and two incompatible answers to a control character in a stack trace.

on this pageshow

explore

questions

5

JUnit 5's `LegacyXmlReportGeneratingListener` writes a `@hostname` attribute on `<testsuite>` that the Maven Surefire report schema never declares. Why does that not break the tools reading the file, and what else differs between the two writers?

level: middleimportance: must knowfreq 52%

answer

  1. the file validates against nothing
  2. one writer adds what the schema omits
  3. hostname, written unconditionally
  4. readers look names up, they do not validate

basics

~20 s

Nothing validates these files. JUnit 5's legacy writer emits hostname unconditionally, falling back to the literal <unknown host>; the Surefire schema declares no such attribute, so the file fails that schema and readers, which never validate, consume it anyway.

solid answer

~40 s

The two writers produce different attribute sets under the same element name, and nothing in the chain notices. `LegacyXmlReportGeneratingListener` writes `@name`, the four counts, `@time`, `@hostname` and `@timestamp` — `@hostname` unconditionally, falling back to the literal `<unknown host>` when the machine name cannot be resolved. Maven Surefire writes `xmlns:xsi` and `xsi:noNamespaceSchemaLocation` pointing at `surefire-test-report.xsd`, a `@version` carrying the schema version, `@name`, the counts, `@flakes` unconditionally, and `@group`, `@time` and `@timestamp` only when it has them. The schema declares no `@hostname`, so JUnit's legacy XML does not validate against it — the cleanest live example of dialect drift in this format. It survives because nobody validates: readers look attributes up by name. Allure 2's `JunitXmlPlugin` goes further and actually reads `@hostname` into the suite information it builds.

code

xml · 7 lines
xml
<?xml version="1.0" encoding="UTF-8"?>
<testsuite name="JUnit Jupiter" tests="2" skipped="0" failures="0" errors="0"
           time="0.211" hostname="&lt;unknown host&gt;"
           timestamp="2026-03-06T12:56:57">
  <testcase name="addsItem()" classname="com.example.CartTest" time="0.014"/>
  <testcase name="clearsCart()" classname="com.example.CartTest" time="0.006"/>
</testsuite>

go deeper

for a junior

Be able to say that two tools writing the same kind of report do not put the same attributes on the suite element, and that no validation step catches the difference.

for a middle

Name the concrete drift: JUnit 5's legacy listener writes hostname, the Surefire schema declares none, and the readers survive it because they look attributes up by name instead of validating.

for a senior

Show the judgement about absence. Explain why a missing attribute cannot be read as a missing measurement here, and why turning on schema validation would fail builds for a difference nothing downstream cares about.

for a principal

Own the position on strictness: whether your organisation validates report files at all, what it does with a producer it cannot change, and how a difference like this gets recorded so it is not rediscovered every year.

## Two writers, one nominal format Both Maven Surefire and JUnit 5's `LegacyXmlReportGeneratingListener` write files people call *JUnit XML*, and both root them at `<testsuite>`. That is roughly where the agreement stops. Each writer emits the attribute set its own authors needed, in its own order, and neither checks itself against the other. Because the format was never standardised, there is no arbiter to say which one is right — only a schema that one of them ships and the other has never heard of. The concrete divergence people meet first is `@hostname`. JUnit 5's listener writes it on every suite, **unconditionally**, and when the machine name cannot be resolved it writes the literal placeholder `<unknown host>` rather than omitting the attribute. `surefire-test-report.xsd` declares no attribute of that name on `<testsuite>` at all. Put a JUnit 5 legacy report through a validator armed with that schema and it fails — not because anything is wrong with the report, but because the schema describes a different writer. ## What each writer actually puts on the suite element | attribute | Maven Surefire | JUnit 5's legacy listener | |---|---|---| | `xmlns:xsi`, `xsi:noNamespaceSchemaLocation` | written, pointing at its own schema | not written | | `@version` | written, carrying the schema version | not written | | `@name` | written | written | | `@tests` `@errors` `@skipped` `@failures` | written | written | | `@time` | written when known | written | | `@timestamp` | written when known | written | | `@group` | written when known | not written | | `@flakes` | written **unconditionally**, `0` when there were none | not written | | `@hostname` | not declared, not written | written **unconditionally**, fallback `<unknown host>` | Two details in that table trip people repeatedly. The first is that `@flakes` is written *outside* any condition, so its absence never means "reruns were off" — a Surefire suite that never repeated anything still carries `flakes="0"`. The second is the mirror image: `@hostname` from the JUnit writer is never absent either, so an empty-looking value is a placeholder, not a missing field. **Absent, empty and zero are three different states here, and each writer picks a different one.** The attribute *order* differs too. Surefire opens with its namespace and schema pointers and puts the counts after the name; JUnit's listener writes name, then the counts, then time, hostname and timestamp. Order is meaningless to an XML parser, but it matters to anything that compares report files as text. ## Why the mismatch is invisible Three things have to be true for a drift like this to survive in production, and all three are: 1. **Nothing validates.** `xsi:noNamespaceSchemaLocation` is a hint to a validating parser about where a schema lives for a namespace-less document. It is not an instruction, and no step in a normal pipeline turns validation on. 2. **Readers look names up.** A consumer walks the parsed document asking for the attributes it cares about. An attribute it never asks for is invisible; an attribute it asks for and does not find is simply null. 3. **The dialect acquired consumers.** Allure 2's `JunitXmlPlugin` reads `hostname` off `<testsuite>` and stores it on the suite information it constructs. An attribute the published schema never declared now has a first-class reader, which makes it part of the real contract regardless of what the schema says. ## Where the drift does bite - **Anyone who switches validation on.** A build step that validates reports against `surefire-test-report.xsd` will fail on every JUnit 5 legacy report, for a difference no downstream tool cares about. That gate has a large false-positive rate by construction. - **Strict converters.** A translator with a fixed, closed list of attributes either drops what it does not know or refuses the document. Dropping is worse, because it is quiet. - **Cross-writer comparison.** Diffing or fingerprinting report files from two writers finds differences on every line even when the outcomes are identical. - **Inference from absence.** Concluding "no `@flakes`, so no reruns happened" is wrong for Surefire's output, where the attribute is always there, and meaningless for JUnit's, where it never is. ## How to tell which writer produced a file Open the suite element and look at three things: a `xsi:noNamespaceSchemaLocation` plus `@version` says Surefire; a `@hostname` says JUnit 5's legacy listener; and `unique-id:` lines in the free-text output element confirm the latter. That check takes seconds and it settles most arguments about whose bug it is before anybody opens a reader's source code.

  • Which value does JUnit 5's legacy writer put in `@hostname` when the host name cannot be resolved?
    The literal string `<unknown host>`, escaped into the attribute value. The attribute is written unconditionally, so it is always present — a consumer reading it has to treat that placeholder as "no host known", not as a machine name it can look up.
  • Which attributes does Maven Surefire put on `<testsuite>` that JUnit 5's legacy writer never emits?
    The namespace declaration and `xsi:noNamespaceSchemaLocation` pointing at its schema, a `@version` carrying the schema version, an optional `@group`, and `@flakes`. JUnit's listener writes none of those; its whole suite attribute set is `@name`, the four counts, `@time`, `@hostname` and `@timestamp`.

saying these in an interview costs you the question

  • Assumes something in the chain validates the file
  • Thinks every JUnit-XML writer emits the same attributes
  • Reads a missing attribute as a missing measurement
  • Treats the Surefire schema as the format's standard
open as a page

The Maven Surefire report schema `surefire-test-report.xsd` declares `<testsuite>` as its only root element, yet many JUnit-XML result files use a `<testsuites>` root holding several suites. What is going on, and what must a reader of these files do about it?

level: middleimportance: must knowfreq 60%

basics

~20 s

No single official schema governs this format. The only published one, surefire-test-report.xsd, has a single root, testsuite, and declares no testsuites wrapper; other writers emit that wrapper anyway, so readers dispatch on whichever root they find.

open as a page

The legacy JUnit-XML shape has no field for a JUnit 5 test's unique id or its full display name. Where does `LegacyXmlReportGeneratingListener` put them, and what does `OpenTestReportGeneratingListener` do instead?

level: middleimportance: should knowfreq 42%

basics

~10 s

The legacy XML has no field for either, so JUnit 5's listener writes unique-id: and display-name: lines into the free-text system-out element. Open Test Reporting gives them real elements instead: uniqueId, legacyReportingName and type.

open as a page

Your build's JUnit-XML result files are read by several tools and none of them validates against `surefire-test-report.xsd`. How do you decide which dialect to emit and what to treat as the contract?

level: principalimportance: should knowfreq 34%

basics

~20 s

The contract is what your consumers actually read, not the one published schema. Emit the shape every reader handles and pin it with a checked-in sample. Add a second format only when the legacy one cannot carry what you need.

open as a page

A stack trace contains a control character that XML 1.0 forbids. Maven Surefire and JUnit 5's `LegacyXmlReportGeneratingListener` answer that differently — what does each do, and what does the difference cost a tool that reads both?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

The two writers disagree. Surefire preserves the character by escaping it twice, so a reader that knows the convention can recover it; JUnit 5 substitutes the replacement character U+FFFD and splits a section-closing sequence across two CDATA sections.

open as a page