skip to content

In REST Assured, what does XmlPath.CompatibilityMode.HTML change about parsing a body?

level: middleimportance: should knowfreq 38%

answer

  1. two constants only: XML and HTML
  2. HTML mode wraps tagsoup, not the DTD flags
  3. the three booleans are XML-mode only
  4. features and properties apply to both
  5. same GPath, different parser underneath

basics

~20 s

XmlPath.CompatibilityMode swaps the parser, not the path language. Its XML constant builds an XmlSlurper from XmlConfig's validating, namespaceAware and allowDocTypeDeclaration flags; its HTML constant wraps the same slurper around tagsoup, which repairs ill-formed markup instead of rejecting it.

solid answer

~50 s

`XmlPath.CompatibilityMode` is an enum with exactly two constants, `XML` and `HTML`, and it picks which parser sits under the same GPath expressions. In `XML` mode REST Assured builds `new XmlSlurper(validating, namespaceAware, allowDocTypeDeclaration)` from `XmlConfig`, so a body that is not well-formed fails. In `HTML` mode it builds `new XmlSlurper(new org.ccil.cowan.tagsoup.Parser())` — the same slurper wrapped around **tagsoup**, which supplies implied elements and closes unclosed tags instead of rejecting the document. The practical consequence is that those three `XmlConfig` booleans reach the constructor only in `XML` mode; `feature(...)`, `property(...)` and declared namespaces are applied afterwards and so apply to both. You select the mode with `new XmlPath(CompatibilityMode.HTML, body)`, and it is exactly the mode `htmlPath()` presets. Tagsoup supplies the implied elements an HTML document may omit and closes tags the markup left open, so a page an XML parser would reject still yields a tree.

code

java · 13 lines
java
import io.restassured.path.xml.XmlPath;

import static io.restassured.path.xml.XmlPath.CompatibilityMode.HTML;

String page = "<html><head><title>Spray diary 2026</title></head>"
            + "<body><ul id='blocks'><li>Northfield<li>Longacre</ul></body></html>";

// tagsoup closes the two <li> elements that the markup left open
XmlPath report = new XmlPath(HTML, page);

String title = report.getString("html.head.title");   // Spray diary 2026
int blocks = report.getInt("html.body.ul.li.size()"); // 2
String first = report.getString("html.body.ul.li[0]");// Northfield

go deeper

for a junior

Know that CompatibilityMode has exactly two constants and that HTML mode exists so you can path over a rendered page. Recall that the GPath expressions themselves do not change.

for a middle

Explain the mechanism: XML mode builds the slurper from the three XmlConfig booleans, HTML mode builds it over tagsoup, and only features, properties and declared namespaces survive both.

for a senior

Argue about when HTML assertions earn their keep at all, and insist that a malformed XML response be reported as a defect rather than parsed leniently to keep a suite green.

for a principal

Set the boundary: how much page-structure assertion belongs in an API suite before it becomes a brittle UI test in disguise, and where that coverage should live instead.

## One class, two parsers `XmlPath` parses two rather different things, and `XmlPath.CompatibilityMode` — a nested enum with exactly two constants, `XML` and `HTML` — decides which. The choice is made once, when the path object is built, and it changes only the parser underneath. The expression language above it does not move an inch: `html.head.title` in HTML mode is the same GPath dialect as `sprayDiary.block[0].@name` in XML mode. Mechanically, REST Assured builds the parser like this: - **`XML`** — `new XmlSlurper(validating, namespaceAware, allowDocTypeDeclaration)`, with all three booleans taken from `XmlConfig` (or `XmlPathConfig` when you build an `XmlPath` yourself). - **`HTML`** — `new XmlSlurper(new org.ccil.cowan.tagsoup.Parser())`, the same Groovy slurper wrapped around **tagsoup**, a SAX parser written for HTML as it is found in the wild rather than HTML as a spec would like it. After construction, and identically in both modes, REST Assured applies every entry from `XmlConfig.features()` via `setFeature`, every entry from `properties()` via `setProperty`, and then declares the namespaces from `declaredNamespaces()` on the parsed result. ## What tagsoup buys you An XML parser has one answer to malformed input: reject it. Tagsoup never rejects; it repairs. That is the whole point of the mode. - Unclosed tags are closed for you, so `<ul><li>Northfield<li>Longacre</ul>` yields two `li` elements. - Implied structure is supplied, so paths still start at `html.body...` even when the markup omitted a tag. - Void elements such as `<br>` and `<img>` do not have to be self-closed — the HTML sample in REST Assured's own documentation contains a bare `<br>`, which an XML parse would refuse outright. - Attribute and element naming is normalised the way an HTML parser normalises it, rather than being treated as case-sensitive XML. That is why a server-rendered page — an orchard spray-diary summary report, a login form, an error page — is readable at all. REST Assured leans on this internally: its own form-login filter and its CSRF token finder parse the login page in HTML mode before they can locate the form fields. ## Which settings still apply in which mode | `XmlConfig` setting | `XML` mode | `HTML` mode | |---|---|---| | `validating(boolean)` | applied via the slurper constructor | not applied | | `namespaceAware(boolean)` | applied via the slurper constructor | not applied | | `allowDocTypeDeclaration(boolean)` | applied via the slurper constructor | not applied | | `feature(uri, enabled)` / `features(Map)` | applied after construction | applied after construction | | `property(name, value)` / `properties(Map)` | applied after construction | applied after construction | | `declareNamespace(prefix, uri)` | declared on the parsed result | declared on the parsed result | The three booleans carry that restriction in their own Javadoc — each one says it is only applicable when the compatibility mode is `XML` — and the code matches: the tagsoup branch calls the single-argument `XmlSlurper` constructor, which those flags never reach. ## Selecting the mode 1. Standalone: `new XmlPath(CompatibilityMode.HTML, body)`. The mode is the first argument, and every source overload has a mode-carrying twin — `String`, `InputStream`, `InputSource`, `File`, `Reader` and `URI`. 2. From a response: `htmlPath()` is defined as `xmlPath(CompatibilityMode.HTML)`, so it is the same object with the mode preset rather than a separate implementation. 3. Statically importing `io.restassured.path.xml.XmlPath.CompatibilityMode.HTML` keeps the call site short, which is how the documentation writes it. ## Judgment calls - Use HTML mode for markup that is genuinely HTML: rendered reports, form pages, error pages you want to assert something cheap about. - Do **not** reach for HTML mode as a lenient-XML escape hatch. If an API's XML response is not well-formed, that is a defect in the response, and quietly parsing it with tagsoup buries the finding instead of reporting it. - Remember that leniency cuts both ways: tagsoup produces *a* tree from almost anything, so a badly wrong page yields a wrong-but-parsable structure and your path silently matches nothing rather than failing at parse time. - If you need a DOCTYPE or namespace behaviour switched, check which mode you are in first — flipping `allowDocTypeDeclaration` and then wondering why nothing changed is nearly always an HTML-mode call. - Keep HTML assertions shallow. Page structure churns far faster than an API contract, so a deep `html.body.div[2].div[0].table.tr[3]` path is a maintenance liability; anchor on an id or a class with a `find { }` instead.

  • You set XmlConfig.allowDocTypeDeclaration(true) and the HTML page still parses the same way. Why?
    Because `allowDocTypeDeclaration`, `validating` and `namespaceAware` reach the parser only through the three-argument `XmlSlurper` constructor, which REST Assured uses in `XML` mode. The `HTML` branch constructs the slurper from a tagsoup parser instead, so those flags are never consulted; only `feature(...)`, `property(...)` and declared namespaces still apply.
  • Does switching to HTML mode change the path expressions you write?
    No. Both modes produce a Groovy `XmlSlurper` tree, so the same GPath applies: dots for child steps, `@` for attributes, `size()`, `find { }` and `**`. Only the shape of the tree differs, because tagsoup inserts the implied `html`, `head` and `body` elements that HTML documents are allowed to omit.

saying these in an interview costs you the question

  • Thinks HTML mode is a different path language, not a different parser
  • Expects validating or allowDocTypeDeclaration to take effect in HTML mode
  • Uses HTML mode to swallow malformed XML from an API
  • Believes tagsoup can fail on ill-formed markup
  • Assumes CompatibilityMode has a lenient or auto constant