skip to content

In JMeter, what does the XPath Extractor offer that the XPath2 Extractor does not?

level: middleimportance: should knowfreq 38%

answer

  1. One of them predates the other
  2. The difference is mostly parser configuration
  3. Only one can chew malformed markup
  4. Namespaces live in different places entirely

basics

~20 s

The Tidy option set. Only the older XPath Extractor can run a malformed HTML body through Tidy first, along with the DTD, namespace and whitespace toggles that go with its parser. The XPath2 Extractor has none of them.

solid answer

~50 s

JMeter ships two XML extractors. The **XPath2 Extractor** is the one the manual has recommended since JMeter 5.0: it evaluates on Saxon, takes namespace prefixes per element as `prefix=namespace` lines in its **Namespaces aliases list** box, and caches compiled queries under `xpath2query.parser.cache.size` (default 400). Its fields stop there — name, query, Match No., Default Value, Namespaces, and *Return entire XPath fragment instead of text content*. The legacy **XPath Extractor** carries a whole parser control panel the newer one dropped: **Use Tidy (tolerant parser)** with **Quiet**, **Report errors** and **Show warnings**, plus **Use Namespaces**, **Validate XML**, **Ignore Whitespace** and **Fetch External DTDs**. Its namespace prefixes come from a properties file named by `xpath.namespace.config`, not from the element. So the only real reason to keep the old element is a tag-soup HTML body needing Tidy — and JMeter's own manual says use the CSS Selector Extractor for HTML instead.

go deeper

for a junior

Recall that JMeter ships two XPath extractors, that the newer one is the recommended default, and that only the older one can be told to tolerate broken markup.

for a middle

Explain what the extra checkboxes on the older element configure, and the two different places the two elements read namespace prefixes from.

for a senior

Show why a plan that works on one machine returns defaults on another, and that a Tidy error and an unreachable external DTD are the two extraction failures here that turn a sample red — one per parser mode — while parse and query errors leave it green.

for a principal

Own the migration call: whether the remaining reason to keep the legacy element is real, or whether the body should be read by a different element family entirely.

Assumed version: Apache JMeter 6.0.0. ## Two elements, one job, different eras Both read values out of an XML or (X)HTML body and store them in a JMeter variable. The **XPath Extractor** is the original. The **XPath2 Extractor** arrived later and is what the component reference points you at, with an explicit note that since JMeter 5.0 you should prefer it. They serialise as `XPathExtractor` and `XPath2Extractor` in a `.jmx`, with property prefixes `XPathExtractor.` and `XPathExtractor2.` respectively — a naming quirk worth knowing if you ever edit a plan by hand. ## The field-by-field difference | Control | XPath Extractor | XPath2 Extractor | |---|---|---| | Name of created variable | yes | yes | | XPath Query | yes | yes | | Match No. (0 for Random) | yes | yes | | Default Value | yes | yes | | Return entire XPath fragment | yes | yes | | Use Tidy (tolerant parser) | **yes** | no | | Quiet / Report errors / Show warnings | **yes** | no | | Use Namespaces | **yes** | no | | Validate XML | **yes** | no | | Ignore Whitespace | **yes** | no | | Fetch External DTDs | **yes** | no | | Namespaces aliases list (in the element) | no | **yes** | Everything in the old element's extra column is parser configuration. The new element does not expose it because its parser does not need most of it. ## What Use Tidy actually does Ticking **Use Tidy (tolerant parser)** runs the response through Tidy to turn broken HTML into well-formed XHTML before any query is evaluated. That conversion is not lossless, and the manual is explicit about two consequences: - **All element and attribute names are lowercased.** A query written against `//DIV` stops matching. - **Improperly nested elements get corrected.** A document whose real structure is `ul/font/li` becomes `ul/li/font`, so a path that mirrored the broken markup no longer describes the document you are querying. The checkbox also gates the others: Quiet, Report errors and Show warnings apply only when Tidy is on, while Use Namespaces, Validate XML, Ignore Whitespace and Fetch External DTDs apply only when it is off. ## Namespaces: two different mechanisms This is the difference that decides most real migrations. 1. **XPath2 Extractor** — type the aliases into the element itself, one `prefix=namespace` per line, in the *Namespaces aliases list* box. The declaration travels with the plan and different elements can use different prefixes. 2. **XPath Extractor** — write the prefixes into a separate properties file and point at it from `user.properties` with `xpath.namespace.config=namespaces.properties`. It is global, it is outside the `.jmx`, and it is the thing people forget when a plan moves to a different machine. The documented alternative is to rewrite queries as `//*[local-name()='tagname' and namespace-uri()='uri-for-namespace']`, which works but reads badly. ## Failure behaviour differs too Neither element is silent about a broken document, and neither behaves like the JSON Extractor, which never touches the sample result: - **XPath Extractor** — a Tidy error adds an assertion result to the sample **and marks the sample unsuccessful**, and with Tidy off an unreachable external DTD does the same. Other parse and query errors add an assertion result without failing the sample. - **XPath2 Extractor** — any evaluation error adds an assertion result without failing the sample. So an XML extractor can turn a sample red where a JSON one cannot. Worth knowing before you conclude a red sample means the server misbehaved. ## What to actually pick - **XML body, well-formed** — XPath2 Extractor. Per-element namespaces, a compiled-query cache, and a supported path forward. - **HTML body** — neither, if you can help it. JMeter's own component reference says the CSS Selector Extractor is the correct and performing solution for HTML, and tells you not to use XPath for it. - **HTML body you cannot restructure the plan around** — the XPath Extractor with Use Tidy is the only option in the pair that will parse it at all, with the lowercasing and re-nesting caveats above. The useful framing at interview is that the old element's extra fields are not features you are giving up; they are the cost of an older parser. The one genuine capability the newer element lacks is tolerant HTML parsing, and JMeter would rather you did not use an XPath extractor for HTML in the first place.

  • A plan using the XPath Extractor works locally and returns defaults on a colleague's machine. What would you check?
    The namespace configuration. That element reads its prefix-to-namespace mappings from a properties file named by `xpath.namespace.config` in `user.properties`, so the mappings live outside the `.jmx` and do not travel with it. Moving the element to the XPath2 Extractor, whose aliases are typed into the element, removes the class of problem.
  • Why does the manual steer you away from XPath for HTML bodies at all?
    Because HTML is usually not well-formed XML, so an XPath extractor either refuses it or has to run Tidy first — and Tidy rewrites the document, lowercasing names and re-nesting misplaced elements, so the tree you query is not the tree that arrived. The CSS Selector Extractor reads the markup as delivered and is what the component reference recommends.
  • Does an extraction failure in these elements affect the sample's result?
    It can. A Tidy error in the XPath Extractor adds an assertion result and marks the sample unsuccessful, and with Tidy off an unreachable external DTD does the same; other parse and query errors in either element add an assertion result but leave the sample successful. That is a meaningful contrast with the JSON Extractor, which logs the error and stores the default without recording anything on the sample.

saying these in an interview costs you the question

  • Thinks the newer element simply dropped features for no reason
  • Believes tolerant parsing is available on both XML extractors
  • Assumes namespace prefixes are configured the same way on both
  • Uses an XPath extractor for HTML when a CSS one is recommended
  • Forgets that Tidy lowercases names and re-nests misplaced elements