In REST Assured, what does XmlPath.CompatibilityMode.HTML change about parsing a body?
answer
- two constants only: XML and HTML
- HTML mode wraps tagsoup, not the DTD flags
- the three booleans are XML-mode only
- features and properties apply to both
- same GPath, different parser underneath
basics
~20 sXmlPath.CompatibilityMode swaps the parser, not the path language. Its XML constant builds an XmlSlurper from XmlConfig's validating, namespaceAware and allowDocTypeDeclaration flags; its HTML constant wraps the same slurper around tagsoup, which repairs ill-formed markup instead of rejecting it.
solid answer
~50 s`XmlPath.CompatibilityMode` is an enum with exactly two constants, `XML` and `HTML`, and it picks which parser sits under the same GPath expressions. In `XML` mode REST Assured builds `new XmlSlurper(validating, namespaceAware, allowDocTypeDeclaration)` from `XmlConfig`, so a body that is not well-formed fails. In `HTML` mode it builds `new XmlSlurper(new org.ccil.cowan.tagsoup.Parser())` — the same slurper wrapped around **tagsoup**, which supplies implied elements and closes unclosed tags instead of rejecting the document. The practical consequence is that those three `XmlConfig` booleans reach the constructor only in `XML` mode; `feature(...)`, `property(...)` and declared namespaces are applied afterwards and so apply to both. You select the mode with `new XmlPath(CompatibilityMode.HTML, body)`, and it is exactly the mode `htmlPath()` presets. Tagsoup supplies the implied elements an HTML document may omit and closes tags the markup left open, so a page an XML parser would reject still yields a tree.
code
java · 13 linesimport io.restassured.path.xml.XmlPath;
import static io.restassured.path.xml.XmlPath.CompatibilityMode.HTML;
String page = "<html><head><title>Spray diary 2026</title></head>"
+ "<body><ul id='blocks'><li>Northfield<li>Longacre</ul></body></html>";
// tagsoup closes the two <li> elements that the markup left open
XmlPath report = new XmlPath(HTML, page);
String title = report.getString("html.head.title"); // Spray diary 2026
int blocks = report.getInt("html.body.ul.li.size()"); // 2
String first = report.getString("html.body.ul.li[0]");// Northfieldgo deeper
Know that CompatibilityMode has exactly two constants and that HTML mode exists so you can path over a rendered page. Recall that the GPath expressions themselves do not change.
Explain the mechanism: XML mode builds the slurper from the three XmlConfig booleans, HTML mode builds it over tagsoup, and only features, properties and declared namespaces survive both.
Argue about when HTML assertions earn their keep at all, and insist that a malformed XML response be reported as a defect rather than parsed leniently to keep a suite green.
Set the boundary: how much page-structure assertion belongs in an API suite before it becomes a brittle UI test in disguise, and where that coverage should live instead.
## One class, two parsers `XmlPath` parses two rather different things, and `XmlPath.CompatibilityMode` — a nested enum with exactly two constants, `XML` and `HTML` — decides which. The choice is made once, when the path object is built, and it changes only the parser underneath. The expression language above it does not move an inch: `html.head.title` in HTML mode is the same GPath dialect as `sprayDiary.block[0].@name` in XML mode. Mechanically, REST Assured builds the parser like this: - **`XML`** — `new XmlSlurper(validating, namespaceAware, allowDocTypeDeclaration)`, with all three booleans taken from `XmlConfig` (or `XmlPathConfig` when you build an `XmlPath` yourself). - **`HTML`** — `new XmlSlurper(new org.ccil.cowan.tagsoup.Parser())`, the same Groovy slurper wrapped around **tagsoup**, a SAX parser written for HTML as it is found in the wild rather than HTML as a spec would like it. After construction, and identically in both modes, REST Assured applies every entry from `XmlConfig.features()` via `setFeature`, every entry from `properties()` via `setProperty`, and then declares the namespaces from `declaredNamespaces()` on the parsed result. ## What tagsoup buys you An XML parser has one answer to malformed input: reject it. Tagsoup never rejects; it repairs. That is the whole point of the mode. - Unclosed tags are closed for you, so `<ul><li>Northfield<li>Longacre</ul>` yields two `li` elements. - Implied structure is supplied, so paths still start at `html.body...` even when the markup omitted a tag. - Void elements such as `<br>` and `<img>` do not have to be self-closed — the HTML sample in REST Assured's own documentation contains a bare `<br>`, which an XML parse would refuse outright. - Attribute and element naming is normalised the way an HTML parser normalises it, rather than being treated as case-sensitive XML. That is why a server-rendered page — an orchard spray-diary summary report, a login form, an error page — is readable at all. REST Assured leans on this internally: its own form-login filter and its CSRF token finder parse the login page in HTML mode before they can locate the form fields. ## Which settings still apply in which mode | `XmlConfig` setting | `XML` mode | `HTML` mode | |---|---|---| | `validating(boolean)` | applied via the slurper constructor | not applied | | `namespaceAware(boolean)` | applied via the slurper constructor | not applied | | `allowDocTypeDeclaration(boolean)` | applied via the slurper constructor | not applied | | `feature(uri, enabled)` / `features(Map)` | applied after construction | applied after construction | | `property(name, value)` / `properties(Map)` | applied after construction | applied after construction | | `declareNamespace(prefix, uri)` | declared on the parsed result | declared on the parsed result | The three booleans carry that restriction in their own Javadoc — each one says it is only applicable when the compatibility mode is `XML` — and the code matches: the tagsoup branch calls the single-argument `XmlSlurper` constructor, which those flags never reach. ## Selecting the mode 1. Standalone: `new XmlPath(CompatibilityMode.HTML, body)`. The mode is the first argument, and every source overload has a mode-carrying twin — `String`, `InputStream`, `InputSource`, `File`, `Reader` and `URI`. 2. From a response: `htmlPath()` is defined as `xmlPath(CompatibilityMode.HTML)`, so it is the same object with the mode preset rather than a separate implementation. 3. Statically importing `io.restassured.path.xml.XmlPath.CompatibilityMode.HTML` keeps the call site short, which is how the documentation writes it. ## Judgment calls - Use HTML mode for markup that is genuinely HTML: rendered reports, form pages, error pages you want to assert something cheap about. - Do **not** reach for HTML mode as a lenient-XML escape hatch. If an API's XML response is not well-formed, that is a defect in the response, and quietly parsing it with tagsoup buries the finding instead of reporting it. - Remember that leniency cuts both ways: tagsoup produces *a* tree from almost anything, so a badly wrong page yields a wrong-but-parsable structure and your path silently matches nothing rather than failing at parse time. - If you need a DOCTYPE or namespace behaviour switched, check which mode you are in first — flipping `allowDocTypeDeclaration` and then wondering why nothing changed is nearly always an HTML-mode call. - Keep HTML assertions shallow. Page structure churns far faster than an API contract, so a deep `html.body.div[2].div[0].table.tr[3]` path is a maintenance liability; anchor on an id or a class with a `find { }` instead.
- You set XmlConfig.allowDocTypeDeclaration(true) and the HTML page still parses the same way. Why?Because `allowDocTypeDeclaration`, `validating` and `namespaceAware` reach the parser only through the three-argument `XmlSlurper` constructor, which REST Assured uses in `XML` mode. The `HTML` branch constructs the slurper from a tagsoup parser instead, so those flags are never consulted; only `feature(...)`, `property(...)` and declared namespaces still apply.
- Does switching to HTML mode change the path expressions you write?No. Both modes produce a Groovy `XmlSlurper` tree, so the same GPath applies: dots for child steps, `@` for attributes, `size()`, `find { }` and `**`. Only the shape of the tree differs, because tagsoup inserts the implied `html`, `head` and `body` elements that HTML documents are allowed to omit.
saying these in an interview costs you the question
- Thinks HTML mode is a different path language, not a different parser
- Expects validating or allowDocTypeDeclaration to take effect in HTML mode
- Uses HTML mode to swallow malformed XML from an API
- Believes tagsoup can fail on ill-formed markup
- Assumes CompatibilityMode has a lenient or auto constant