In a Karate feature file, match //teacher[@department='science']/subject == ['math', 'physics'] passes. What does that same step produce when the science teacher has only one subject element, and why does comparing it to ['math'] then fail?
answer
- Count the nodes, not the path
- Zero, one and many differ
- One node is unwrapped
- A list only when it repeats
- data types don't match
basics
~20 sAn XPath selecting several nodes comes back as a list, so you compare it to a JSON array. A single node is unwrapped, so with one subject the step yields the string 'math' and ['math'] fails with data types don't match.
solid answer
~50 sKarate turns the XPath result into an ordinary value before the comparison runs, and the rule is **arity-based**: zero nodes is not present, exactly one node is unwrapped to a single value, and two or more become a `List`. That is what lets `match //teacher[@department='science']/subject == ['math', 'physics']` read so naturally - a node-set is just a JSON array. The trap is that the same step against a teacher with one subject yields the plain string `'math'`, not `['math']`, so the array form fails with `data types don't match`. The arity of the *data* decides the shape, not the shape of the XPath. When the count can be one, assert with `contains` on a wrapped value, assert the count separately with `count(...)`, or convert the document to JSON first, where a repeating element is a list only when it actually repeats.
code
gherkin · 13 lines* def response =
"""
<teachers>
<teacher department="science"><subject>math</subject><subject>physics</subject></teacher>
<teacher department="arts"><subject>english</subject></teacher>
</teachers>
"""
# two nodes -> a list
* match //teacher[@department='science']/subject == ['math', 'physics']
# one node -> a scalar, NOT ['english']
* match //teacher[@department='arts']/subject == 'english'
# arity stated out loud
* match response count(//teacher[@department='arts']/subject) == 1go deeper
Learn the headline: several nodes come back as an array, one node comes back on its own. Check how many the document really has before writing brackets on the right side.
Explain the arity rule end to end - zero is not present, one is unwrapped, many become a list - and name data types don't match as the failure you get when the wrapping is wrong.
Treat any XML assertion whose arity can vary by environment as fragile. Pin the count with count(...) or a positional predicate so a one-child payload does not fail a step nobody edited.
Decide as a standard whether XML suites assert through XPath or convert to JSON at the boundary. Mixed conventions mean every reviewer re-derives the wrapping rules on every diff.
## A node-set becomes a Karate value before the match runs An XPath on the left of a `match` is evaluated against the DOM and then **converted**. The conversion is small enough to hold in your head, and it is the same in both Karate lines: | Nodes selected | What the match sees | |---|---| | 0 | not present, so `== '#notpresent'` passes | | 1, element with no child elements | that element's text, as a **string** | | 1, element with child elements | a fresh XML document you compare to an XML literal | | 1, attribute node | the attribute's value, as a **string** | | 2 or more | a **list**, each item converted by the rules above | So the array in `match //teacher[@department='science']/subject == ['math', 'physics']` is not a special XML feature. Two `<subject>` elements were selected, each has no child elements, so each became its text, and the two texts became a list. Comparing a list to a JSON array is the ordinary `match` you would write over JSON. ```gherkin * def response = """ <teachers> <teacher department="science"><subject>math</subject><subject>physics</subject></teacher> <teacher department="arts"><subject>english</subject></teacher> </teachers> """ * match //teacher[@department='science']/subject == ['math', 'physics'] # one node, so no list is built * match //teacher[@department='arts']/subject == 'english' ``` ## Why the single-node case bites The conversion is driven by **how many nodes the data actually has**, not by how the XPath looks. `//teacher[@department='arts']/subject` is written exactly like the two-node path, but the arts teacher has one subject, so the result is the string `'english'`. Assert it as `['english']` and the match fails with `data types don't match` - a string on the left, an array on the right, and the engine refuses to coerce between the two. That makes the assertion **data-dependent in a way the step text hides**. A suite written against a fixture where every parent has two children passes; the first environment where one parent has a single child fails on a step nobody changed. It is the XML twin of the JSON habit of assuming a repeated key is always an array. The same rule bites in the other direction. `match //teacher/subject == 'math'` looks like it asserts the first subject; against the document above it selects three nodes, builds a list, and fails. ## The signals that a step is arity-fragile When reviewing an XML suite, three things mark an assertion that will break on a different payload: - a **bracketed right side** on a path with no positional predicate - the step is asserting that the data repeats, silently; - a path ending in a **child element name** rather than a `[n]` or an attribute - `/teacher/subject` can be one or many, `/teacher/@id` cannot; - an expected array whose length is **one** - almost always a fixture captured from a payload that happened to have a single child. ## Three ways to write it so arity does not matter 1. **Assert the count separately.** `match response count(//teacher[@department='science']/subject) == 2` states the arity out loud, and then a positional path such as `//teacher[@department='science']/subject[1]` always selects exactly one node. 2. **Use `contains` when membership is the point.** Under `contains` a non-list expected value is wrapped in a single-element list before the walk, so the shapes line up in the common cases and the assertion reads as "physics is among the subjects". 3. **Convert the document to JSON first.** `* json data = response` gives you a map where a repeating element is a list; you are then in JsonPath territory with the same one-versus-many caveat, but with JSON operators and a JSON diff on failure. ## What the failure looks like The message is `data types don't match`, with the two types named. It is worth recognising on sight, because the instinct is to suspect the XPath - and the XPath is fine. The path selected exactly what you asked for; only the wrapping differs. When you see it on an XML assertion, print the left side first: ```gherkin * def subjects = //teacher[@department='arts']/subject * print subjects ``` If `print` shows a bare string rather than a bracketed list, the path matched one node, and the assertion - not the data - needs adjusting. ## The wider point This is the price of the design that makes XML pleasant here: a document is flattened into the same value model as JSON so one `match` engine serves both. The flattening has to make a choice about a one-element node-set, and Karate chose to unwrap it, because reading a single element is the far more common intent. Knowing that choice turns a puzzling failure into a one-line fix, and it is the detail an interviewer uses to separate someone who has run an XML suite from someone who has only read the syntax.
- What does an XPath that selects one element with children give you?A fresh XML document holding that element, not its text. So `match response //teacher[@department='arts'] == <teacher department="arts"><subject>english</subject></teacher>` is the right shape. Text is only produced for an element with **no** child elements, which is why a leaf gives `'english'` but its parent gives a chunk.
- How do you assert on just the second of several selected nodes?Put the position in the XPath: `match foo /records/record[2] == 'b'`. XPath positions are **1-based**, unlike a Karate array index. A positional predicate also guarantees a single-node result, which sidesteps the list-versus-scalar problem entirely.
- Does the same one-versus-many rule apply after converting the XML to JSON?Yes, in spirit. `json data = response` builds a map in which a child element name maps to a list only when that element actually repeats under its parent. A single occurrence is a plain value. The engine that builds the map detects repeats as it walks, so the shape still follows the data.
saying these in an interview costs you the question
- Assumes an XPath always returns an array
- Says Karate coerces 'math' to ['math'] automatically
- Blames the XPath when the failure is the wrapping
- Thinks the node count is fixed by the path text
- Reads XPath positional predicates as 0-based