In Appium, why is an xpath locator slow on Android and on iOS?
answer
- no XPath engine on the device
- a document has to be manufactured
- cost per find, not per expression
- Android XPath 2, iOS XPath 1.0
basics
~20 sNeither platform evaluates XPath natively, so each xpath find first serialises the accessibility hierarchy into an XML document - built by the UiAutomator2 server on Android and by WebDriverAgent on iOS - and matches the expression against that copy.
solid answer
~50 sA browser answers an XPath query out of a document it already holds. A phone holds no document: Android has an accessibility node tree and Apple platforms have XCTest element snapshots, and neither ships an XPath engine. So the driver manufactures one per find. On Android the UiAutomator2 driver's on-device server walks the node tree, serialises it to XML and evaluates the expression there; on iOS `WebDriverAgent` takes an element snapshot, builds its own XML and evaluates XPath 1.0 over it. **The cost is paid per find call and scales with the size of the tree, not the length of the expression.** The two are not equally expensive either: Android evaluates XPath 2 by default, with the UiAutomator2 setting `enforceXPath1` as the fallback switch, while the XCUITest driver's docs warn an `xpath` lookup can be up to ten times slower than an `-ios predicate string` or `-ios class chain` one.
go deeper
Be ready to say that an xpath locator in Appium is not free: the driver turns the screen into an XML document first, on Android and on iOS alike, before it matches anything.
Explain who builds that document on each platform - the UiAutomator2 server on Android, WebDriverAgent on iOS - and that it is rebuilt on every find call rather than cached between them.
Show you have measured it: name which finds in a run use xpath, how big the tree is on those screens, and which cheaper per-platform strategy you moved the hot ones to.
Own the trade-off between one portable expression and two platform-specific locators, and say how you keep that cost visible in the run rather than settling it as a style rule.
## The browser intuition, and why it does not transfer In a browser, an XPath query is answered by an engine that already owns the document it is querying. The DOM is the browser's own data structure, it is live, and an XPath implementation sits directly on top of it. The query is a read. On a phone there is no document. What the automation stack has is an **accessibility hierarchy**: on Android a tree of accessibility nodes the platform maintains for assistive technology, and on Apple platforms a tree of element snapshots that XCTest hands to a test process. Neither of those trees ships an XPath engine, and neither is queryable by an XPath expression as it stands. So before an `xpath` expression can match anything, something has to **manufacture a document for it** - walk the tree, write it out as XML, and evaluate the expression against that copy. That manufacturing step is the whole cost model, and it is why the advice carried over from browser work barely moves the number here. ## What each driver does on every find call - **Android, UiAutomator2 driver.** The find is answered by the driver's on-device server, which walks the accessibility node tree, serialises it into an XML document, and evaluates the expression over that document. `xpath` is one of the six locator strategies that driver declares, alongside `id`, `class name`, `accessibility id`, `css selector` and `-android uiautomator`. - **iOS, XCUITest driver.** The find is answered by `WebDriverAgent`, which obtains an XCTest element snapshot, builds its own XML from it, and evaluates XPath 1.0 over that. `xpath` is one of the eight native strategies that driver declares. - **Android, Espresso driver.** Its on-device server owns the find routes, and that server's strategy set includes `xpath` too - so the common claim that the Espresso driver has no XPath strategy is simply wrong. - **Every single time.** The serialisation is per find call. Two consecutive `xpath` finds build two documents; the driver hands the second one nothing cached from the first. ## Where the cost actually comes from Three things drive it, and none of them is the expression: 1. **How much of the screen is in the tree.** A tram fare-inspection screen showing forty scanned-ticket rows serialises forty rows' worth of nodes whether the expression matches one of them or all of them. 2. **How deep the tree runs.** Nesting multiplies nodes, and every node contributes its attributes to the document. 3. **What it costs to read the tree at all.** Here the platforms part company: the Android server reads a node tree it can already see, while on Apple platforms the snapshot has to be produced through XCTest and crosses into the application under test. ## The two platforms side by side | | Android, UiAutomator2 driver | iOS, XCUITest driver | |---|---|---| | Who builds the document | the on-device UiAutomator2 server | `WebDriverAgent` | | Built from | the accessibility node tree | an XCTest element snapshot | | XPath dialect | XPath 2 by default | XPath 1.0 | | Dialect switch | the `enforceXPath1` setting | none in the settings reference | | Shared scope switch | `limitXPathContextScope` | `limitXPathContextScope` | | Cheaper same-lane strategies | `id`, `accessibility id`, `-android uiautomator` | `-ios predicate string`, `-ios class chain` | The XCUITest driver's own documentation is blunt about the consequence: an `xpath` lookup there can run **up to ten times slower** than the equivalent `-ios predicate string` or `-ios class chain` lookup. That warning is why those two iOS-only strategies exist in the first place. ## What this means for a suite - Treat every `xpath` find as a request for a whole document, not a cheap read. - Count finds, not expressions: one `findElements` call that returns every matching row pays for one document, while forty single finds pay for forty. - Expect the same test to degrade unevenly - a per-row `xpath` loop over a fare-inspection ticket list is merely wasteful on Android and can dominate the run on iOS. - Use the per-platform strategies on the hot paths and keep `xpath` for the places where one portable expression genuinely earns its cost. - Do not lengthen a relative path into an absolute one hoping to help; the same document still gets built. ## The one thing that is not the fix Shortening the expression. Evaluating the expression is the cheap half; building the document is the expensive half, and it happens identically whether the expression is two steps or twenty. The levers that do move the number are how often you ask, how much tree each ask has to cover, and which strategy answers the hot paths. Android has one further lever that really is about the expression - `enforceXPath1` picks which engine evaluates it - but even that does not skip the serialisation, and iOS has no equivalent knob at all.
- Does a second xpath find reuse the document the first one built?No. Serialisation happens per find call on both platforms, so two consecutive `xpath` finds pay for two documents. That is why one `findElements` call returning every match is cheaper than a loop of single finds, and why the gap is far wider in an iOS lane than an Android one.
- Which Appium locator strategies avoid the serialisation cost on each platform?On Android the UiAutomator2 driver's `id`, `accessibility id` and `-android uiautomator` are matched by the on-device server without an XML document. On iOS the XCUITest driver's `-ios predicate string` and `-ios class chain` are resolved by `WebDriverAgent` against the element tree, which is exactly why its docs recommend them over `xpath`.
A browser is asked where one door is in a floor plan it already has open. Appium has to redraw the whole floor plan before it can point at anything.
saying these in an interview costs you the question
- Says XPath is slow because the expression is long
- Assumes the device evaluates XPath natively, as a browser does
- Thinks the driver caches one page source across finds
- Treats Android and iOS xpath cost as identical
- Claims Appium has no xpath strategy on iOS