A stack trace contains a control character that XML 1.0 forbids. Maven Surefire and JUnit 5's `LegacyXmlReportGeneratingListener` answer that differently — what does each do, and what does the difference cost a tool that reads both?
answer
- the format says nothing about it
- one writer keeps it, one replaces it
- double escaping versus a replacement character
- same failure, two different strings
basics
~20 sThe two writers disagree. Surefire preserves the character by escaping it twice, so a reader that knows the convention can recover it; JUnit 5 substitutes the replacement character U+FFFD and splits a section-closing sequence across two CDATA sections.
solid answer
~50 sXML 1.0 forbids most control characters outright — they cannot appear in a document even as an escape — and a stack trace is arbitrary program output, so a writer must decide what to do. The format says nothing, so each writer invented an answer. Maven Surefire keeps the value: it detects the illegal character and writes it doubly escaped, an `&#N;` form inside an attribute and an `&#N;` form in text, on the reasoning that a recoverable encoding beats a lost byte. JUnit 5's legacy listener does the opposite: it substitutes the replacement character U+FFFD for anything not permitted, and where the text would close a CDATA section early it splits the `]]>` across two sections instead. The consequence for a consumer is that the same failure, captured by two writers, yields two different strings.
go deeper
Know that some characters a program prints simply cannot appear in an XML document, and that the tool writing the report has to do something about them before the trace reaches the file.
Explain the two answers concretely: one writer encodes the character twice so it survives, the other substitutes a replacement character and splits a section that would otherwise close early.
Show where the divergence causes real damage — pattern rules that stop matching, fingerprints that split a single failure in two — and how you would normalise trace text on ingest instead of chasing symptoms.
Own the position that trace text from different producers is not comparable evidence, and decide where a value that must survive byte-for-byte is kept instead of inside a format that forbids the bytes.
## The hole in the format XML 1.0 does not merely require that certain characters be escaped; it forbids most control characters from appearing in a document **at all**, escaped or not. A stack trace, meanwhile, is whatever a program produced — and programs emit terminal control sequences, framing bytes, and the occasional stray byte from a corrupted stream. So every writer of this format hits the same wall: it has been handed text it is required to record and forbidden to reproduce. The interchange format offers no guidance whatsoever. There is no rule about lossy substitution, no declared sentinel, no place to record that a substitution happened. Each writer therefore invented its own answer, and the two most common writers invented **incompatible** ones. That is the whole of this problem, and it is a good example of what "the format is silent here" costs in practice: silence does not produce an absence, it produces divergence. ## Surefire's answer: keep the value, escape it twice Maven Surefire scans the text it is about to write for characters illegal in XML 1.0. When it finds one, it does not drop it and does not replace it. It emits a token built from the character's numeric code, deliberately escaped a second time so that the parser cannot turn it back into the forbidden character on the way in: - inside an attribute value, an `&#N;` form; - inside element text, an `&#N;` form, whose ampersand is itself escaped again on the way out. The trade is explicit in the design: the document stays legal, and **the original value is still in the file** for a consumer that knows the convention. The cost is that the value is no longer plain text. Anything that reads the trace naively sees the escape token, not the character, and a human reading the report sees it too. ## JUnit 5's answer: replace the character, split the section JUnit 5's legacy listener takes the opposite branch. It walks the text and, for every code point not permitted in XML, appends the Unicode replacement character **U+FFFD** in its place. The forbidden byte does not survive in any form. It has a second rule for a related problem. Test output is written inside CDATA sections, and a CDATA section ends at the first `]]>` — so text that itself contains that sequence would terminate the section early and corrupt the document. Rather than abandon CDATA, the writer splits the run between the `]]` and the `>`, emitting two consecutive sections whose contents concatenate back to the original text when parsed. Both rules are applied to attribute values and to output text alike, so the substitution is uniform across the document. ## What a consumer actually sees | | Maven Surefire | JUnit 5's legacy listener | |---|---|---| | forbidden control character | preserved as a doubly escaped numeric token | replaced by U+FFFD | | original value recoverable | yes, by a reader that knows the convention | no, the byte is gone | | every forbidden character | keeps its own distinct encoding | collapses to the same replacement | | a `]]>` inside output | handled by the escaping rules | section is split in two, text is unchanged | | the fact that anything happened | inferable from the token | inferable only from the U+FFFD glyph | The asymmetry that matters is **reversibility**. Surefire's answer is lossless and needs decoding; JUnit's is lossy and needs nothing. A pipeline that ingests both and treats trace text as opaque will show one of them as escape noise and the other as a scattering of replacement glyphs, and neither will match a trace captured any other way. ## Where this actually bites 1. **Pattern matching over trace text.** Any consumer that buckets or classifies failures by matching a pattern against the message or the stack trace is matching against a string one writer altered and the other encoded. A rule that fires on one CI job can silently stop firing when the same suite is run by the other writer. 2. **Fingerprinting and deduplication.** Hashing trace text to group repeated failures produces two different hashes for one failure captured twice, so the grouping silently splits. 3. **Diffing across writers.** Comparing reports from two producers finds differences that are entirely an artefact of the writers' escaping choices. 4. **Reading the report as a human.** A stack trace peppered with escape tokens or replacement glyphs is the first symptom, and it is easy to misdiagnose as an encoding problem in the test itself. ## Living with both There is no fix at the format level, because there is no format to fix. What works is narrower: know which writer produced a given file before you interpret its trace text, normalise on ingest if you are going to match patterns against it, and never treat trace text from two producers as directly comparable. If a value genuinely must survive a round trip byte-for-byte, the trace field of an interchange file is the wrong place to keep it — attach it as its own artefact instead of asking a format that forbids the bytes to carry them.
- Which of the two behaviours can a consumer reverse, and which cannot?Double escaping is reversible: the original code point is still encoded in the text, so a consumer that knows the convention can decode it. Substitution is lossy — every forbidden character collapses onto the same replacement glyph, so nothing in the file distinguishes them and the original value is gone for good.
- Given a result file, how would you tell which of the two writers produced it?Look at the suite element. Surefire's output carries a schema-location attribute and a `@version`, plus `@flakes`; JUnit 5's legacy output carries `@hostname` and `unique-id:` lines inside its captured-output element. Once you know the writer, its escaping behaviour follows without inspecting the trace at all.
saying these in an interview costs you the question
- Assumes both writers produce identical text for one trace
- Thinks a replacement character can be decoded back
- Believes the format specifies how illegal characters are handled
- Expects a schema check to catch the difference