skip to content

In a Karate feature file, match //teacher[@department='science']/subject == ['math', 'physics'] passes. What does that same step produce when the science teacher has only one subject element, and why does comparing it to ['math'] then fail?

level: middleimportance: must knowfreq 46%

answer

  1. Count the nodes, not the path
  2. Zero, one and many differ
  3. One node is unwrapped
  4. A list only when it repeats
  5. data types don't match

basics

~20 s

An XPath selecting several nodes comes back as a list, so you compare it to a JSON array. A single node is unwrapped, so with one subject the step yields the string 'math' and ['math'] fails with data types don't match.

solid answer

~50 s

Karate turns the XPath result into an ordinary value before the comparison runs, and the rule is **arity-based**: zero nodes is not present, exactly one node is unwrapped to a single value, and two or more become a `List`. That is what lets `match //teacher[@department='science']/subject == ['math', 'physics']` read so naturally - a node-set is just a JSON array. The trap is that the same step against a teacher with one subject yields the plain string `'math'`, not `['math']`, so the array form fails with `data types don't match`. The arity of the *data* decides the shape, not the shape of the XPath. When the count can be one, assert with `contains` on a wrapped value, assert the count separately with `count(...)`, or convert the document to JSON first, where a repeating element is a list only when it actually repeats.

code

gherkin · 13 lines
gherkin
* def response =
  """
  <teachers>
    <teacher department="science"><subject>math</subject><subject>physics</subject></teacher>
    <teacher department="arts"><subject>english</subject></teacher>
  </teachers>
  """
# two nodes -> a list
* match //teacher[@department='science']/subject == ['math', 'physics']
# one node -> a scalar, NOT ['english']
* match //teacher[@department='arts']/subject == 'english'
# arity stated out loud
* match response count(//teacher[@department='arts']/subject) == 1

go deeper

for a junior

Learn the headline: several nodes come back as an array, one node comes back on its own. Check how many the document really has before writing brackets on the right side.

for a middle

Explain the arity rule end to end - zero is not present, one is unwrapped, many become a list - and name data types don't match as the failure you get when the wrapping is wrong.

for a senior

Treat any XML assertion whose arity can vary by environment as fragile. Pin the count with count(...) or a positional predicate so a one-child payload does not fail a step nobody edited.

for a principal

Decide as a standard whether XML suites assert through XPath or convert to JSON at the boundary. Mixed conventions mean every reviewer re-derives the wrapping rules on every diff.

## A node-set becomes a Karate value before the match runs An XPath on the left of a `match` is evaluated against the DOM and then **converted**. The conversion is small enough to hold in your head, and it is the same in both Karate lines: | Nodes selected | What the match sees | |---|---| | 0 | not present, so `== '#notpresent'` passes | | 1, element with no child elements | that element's text, as a **string** | | 1, element with child elements | a fresh XML document you compare to an XML literal | | 1, attribute node | the attribute's value, as a **string** | | 2 or more | a **list**, each item converted by the rules above | So the array in `match //teacher[@department='science']/subject == ['math', 'physics']` is not a special XML feature. Two `<subject>` elements were selected, each has no child elements, so each became its text, and the two texts became a list. Comparing a list to a JSON array is the ordinary `match` you would write over JSON. ```gherkin * def response = """ <teachers> <teacher department="science"><subject>math</subject><subject>physics</subject></teacher> <teacher department="arts"><subject>english</subject></teacher> </teachers> """ * match //teacher[@department='science']/subject == ['math', 'physics'] # one node, so no list is built * match //teacher[@department='arts']/subject == 'english' ``` ## Why the single-node case bites The conversion is driven by **how many nodes the data actually has**, not by how the XPath looks. `//teacher[@department='arts']/subject` is written exactly like the two-node path, but the arts teacher has one subject, so the result is the string `'english'`. Assert it as `['english']` and the match fails with `data types don't match` - a string on the left, an array on the right, and the engine refuses to coerce between the two. That makes the assertion **data-dependent in a way the step text hides**. A suite written against a fixture where every parent has two children passes; the first environment where one parent has a single child fails on a step nobody changed. It is the XML twin of the JSON habit of assuming a repeated key is always an array. The same rule bites in the other direction. `match //teacher/subject == 'math'` looks like it asserts the first subject; against the document above it selects three nodes, builds a list, and fails. ## The signals that a step is arity-fragile When reviewing an XML suite, three things mark an assertion that will break on a different payload: - a **bracketed right side** on a path with no positional predicate - the step is asserting that the data repeats, silently; - a path ending in a **child element name** rather than a `[n]` or an attribute - `/teacher/subject` can be one or many, `/teacher/@id` cannot; - an expected array whose length is **one** - almost always a fixture captured from a payload that happened to have a single child. ## Three ways to write it so arity does not matter 1. **Assert the count separately.** `match response count(//teacher[@department='science']/subject) == 2` states the arity out loud, and then a positional path such as `//teacher[@department='science']/subject[1]` always selects exactly one node. 2. **Use `contains` when membership is the point.** Under `contains` a non-list expected value is wrapped in a single-element list before the walk, so the shapes line up in the common cases and the assertion reads as "physics is among the subjects". 3. **Convert the document to JSON first.** `* json data = response` gives you a map where a repeating element is a list; you are then in JsonPath territory with the same one-versus-many caveat, but with JSON operators and a JSON diff on failure. ## What the failure looks like The message is `data types don't match`, with the two types named. It is worth recognising on sight, because the instinct is to suspect the XPath - and the XPath is fine. The path selected exactly what you asked for; only the wrapping differs. When you see it on an XML assertion, print the left side first: ```gherkin * def subjects = //teacher[@department='arts']/subject * print subjects ``` If `print` shows a bare string rather than a bracketed list, the path matched one node, and the assertion - not the data - needs adjusting. ## The wider point This is the price of the design that makes XML pleasant here: a document is flattened into the same value model as JSON so one `match` engine serves both. The flattening has to make a choice about a one-element node-set, and Karate chose to unwrap it, because reading a single element is the far more common intent. Knowing that choice turns a puzzling failure into a one-line fix, and it is the detail an interviewer uses to separate someone who has run an XML suite from someone who has only read the syntax.

  • What does an XPath that selects one element with children give you?
    A fresh XML document holding that element, not its text. So `match response //teacher[@department='arts'] == <teacher department="arts"><subject>english</subject></teacher>` is the right shape. Text is only produced for an element with **no** child elements, which is why a leaf gives `'english'` but its parent gives a chunk.
  • How do you assert on just the second of several selected nodes?
    Put the position in the XPath: `match foo /records/record[2] == 'b'`. XPath positions are **1-based**, unlike a Karate array index. A positional predicate also guarantees a single-node result, which sidesteps the list-versus-scalar problem entirely.
  • Does the same one-versus-many rule apply after converting the XML to JSON?
    Yes, in spirit. `json data = response` builds a map in which a child element name maps to a list only when that element actually repeats under its parent. A single occurrence is a plain value. The engine that builds the map detects repeats as it walks, so the shape still follows the data.

saying these in an interview costs you the question

  • Assumes an XPath always returns an array
  • Says Karate coerces 'math' to ['math'] automatically
  • Blames the XPath when the failure is the wrapping
  • Thinks the node count is fixed by the path text
  • Reads XPath positional predicates as 0-based