Your build's JUnit-XML result files are read by several tools and none of them validates against `surefire-test-report.xsd`. How do you decide which dialect to emit and what to treat as the contract?
answer
- the schema is a vendor's, not a standard
- readers define the real contract
- pin a sample, gate on nothing
- add an artefact, never invent an attribute
basics
~20 sThe contract is what your consumers actually read, not the one published schema. Emit the shape every reader handles and pin it with a checked-in sample. Add a second format only when the legacy one cannot carry what you need.
solid answer
~50 sStart by rejecting the tempting answer: the published schema is not the contract. `surefire-test-report.xsd` describes one vendor's writer, and a writer you probably do not control — JUnit 5's legacy listener — already violates it by emitting `@hostname`. Adopting it as a gate fails builds for a difference nothing downstream cares about. The real contract is the intersection of what your consumers demonstrably look up, and the only durable way to state it is a **checked-in sample file plus a list of the elements and attributes each reader needs**, diffed when a producer is upgraded. Emit the shape that intersection supports, prefer the singular `<testsuite>` root that both major writers here produce, and handle both roots on the reading side. When you need something the shape has no channel for — real identity, node type — emit a second artefact such as Open Test Reporting alongside rather than inventing attributes.
go deeper
Know that the result file your build writes is read by other tools, and that changing what it contains can quietly break them because nothing checks the file for you.
Explain why the published schema describes one writer rather than the format, and why a build gate based on it fails for reasons no downstream consumer cares about.
Show how you would establish the real contract: enumerate what each consumer looks up, commit a sample report per producer, and diff it when a producer is upgraded rather than trusting a schema.
Own the strategy end to end — what your organisation gates on, how a private extension is refused, when a second reporting artefact is worth running in parallel, and what would have to be true to retire the old one.
## Three things people mean by "the contract" When a team argues about what its result files must look like, three different answers are usually in the room at once, and they are not compatible. 1. **The published schema.** `surefire-test-report.xsd`, `version="3.0.2"`, is the only formally published description in this family. It is precise, it is checkable, and it is one vendor's account of one vendor's writer. 2. **The union of the dialects.** Everything any producer emits: both roots, every attribute anyone writes, every escaping convention. Permissive, uncheckable, and useless as a target. 3. **What the readers actually look up.** The set of elements and attributes the specific consumers in your chain ask for by name. Unwritten by default, but the only one whose violation actually breaks something. Only the third is a contract in the sense that matters — breaking it produces a visible failure, and honouring it produces working reports. The other two are, respectively, too narrow and too wide. ## Why the published schema is the wrong gate It is worth being concrete about why the obvious move fails, because it is proposed on most teams eventually. - **It describes a writer, not the format.** It was written to document Surefire's output, not to standardise an ecosystem, and there is no body that could have standardised one. - **A writer you rely on already violates it.** JUnit 5's `LegacyXmlReportGeneratingListener` writes `@hostname` unconditionally, and the schema declares no such attribute. Validate its output and every build fails, permanently, over an attribute no consumer objects to. - **The failure it produces is uninformative.** A validation error tells you the document differs from one vendor's model. It does not tell you that any reader will struggle, which is the question you actually had. - **It cannot see the real risks.** The divergences that hurt — control characters represented two ways, a plural root some readers flatten and others reject — are either invisible to a schema or outside its reach entirely. None of that means validation is worthless as a **diagnostic**. Running it once, deliberately, to enumerate exactly how a producer's output differs from that schema is a fine way to discover your dialect. Wiring it as a gate that fails builds is what does not survive contact with reality. ## What to pin instead The workable substitute is small and boring: - **A checked-in sample.** One real report file per producer you run, committed beside the code, regenerated and diffed when a producer is upgraded. It captures the root element, the attribute set, the attribute order and the escaping convention in one artefact that a human can read. - **A named list of what each consumer needs.** Written down once: this dashboard reads the counts and the suite name; that ingester needs `@classname`; this one reads `@hostname`. The list is short, it changes rarely, and it converts "will this upgrade break anything?" from a discussion into a check. - **A read-side that accepts both roots.** Anything you write yourself should dispatch on the root element name and handle both the singular and the plural form. It costs a few lines and removes an entire class of silent, empty reports. - **An explicit position on absence.** Decide, and write down, that a missing attribute means "this writer does not emit it" and never "the measurement was zero" — because in this format some attributes are written unconditionally with a zero value and others are omitted entirely. ## When a second format earns its keep Sooner or later something you need has no channel in the legacy shape at all — a stable identity, a node type, anything the file simply cannot say. There are only two moves, and one of them is a trap. The trap is **inventing an attribute**. It is easy, it feels harmless, and it produces exactly the drift this whole problem is made of: a file that is fine for your reader and subtly wrong for everyone else's, with nothing recording why. The move that works is emitting a **second artefact** beside the first. JUnit 5's Open Test Reporting output is the worked example: the legacy XML keeps its consumers, and anything that wants real structure reads the new document. You accept two artefacts per run and a migration you may never finish, and in exchange nothing that reads the old file has to change on your schedule. The judgement to make explicit before committing: 1. Which consumers can read the second format today, and which will never be upgraded. 2. Whether the value you need is worth an artefact nobody currently reads. 3. What has to become true before the legacy artefact can stop being produced — and whether anyone will notice if that condition is never met. ## The position worth defending State the contract as *what our readers require*, evidence it with a committed sample, gate on nothing, and treat the published schema as documentation of one producer among several. It is less satisfying than a schema check, and it is the only formulation that stays true when the next producer joins the chain.
- Is there any use for validating against `surefire-test-report.xsd` at all?Yes, as a one-off diagnostic. Running it deliberately enumerates exactly how a producer's output differs from that schema, which is a fast way to discover your dialect and document it. What does not work is wiring it as a gate, since a writer you rely on violates it permanently and harmlessly.
- A consumer needs one field the format cannot carry. Why not just add an attribute for it?Because it produces exactly the drift the format already suffers from: a file that suits your reader and is subtly wrong for every other, with nothing recording that a private extension exists. Emitting a second artefact alongside keeps the old file honest and leaves its consumers untouched.
saying these in an interview costs you the question
- Treats the vendor schema as the ecosystem's standard
- Gates the build on validating report files
- Adds private attributes to a shared interchange file
- Assumes all producers emit the same dialect
- Never records which shape consumers actually require