Why does ZAP ship two browser-driven crawlers, and what can spiderClient see that spiderAjax cannot?
answer
- Two add-ons, two ways of watching
- One sees traffic, one reads the page
- Crawljax versus a ZAP browser extension
- The shipped help names a recommendation
basics
~20 sZAP's spiderAjax add-on drives a browser with a bundled Crawljax but only watches HTTP traffic on a proxy it starts itself. The client add-on's spiderClient adds a ZAP browser extension that reads the DOM, so it finds content spiderAjax cannot.
solid answer
~50 sBoth drive a real browser, but they look at different things. The `spiderAjax` add-on embeds Crawljax, which clicks elements through a WebDriver session while ZAP watches the traffic on a local proxy it starts per browser — ZAP's own shipped help says the AJAX spider "is not able to directly access the DOM". The `client` add-on's `spiderClient` runs a first-party ZAP browser extension inside the page, and that extension reads the DOM and streams the components and events it finds back to ZAP, so a route that exists only as a JavaScript click handler still becomes a node ZAP knows about. That is why the shipped AJAX spider help now says it "is no longer the recommended option for crawling modern web apps, we now recommend that you use the Client Spider instead". Neither crawler is in a plain install; each needs its add-on plus `selenium`.
code
yaml · 9 linesjobs:
- type: spiderClient # the client add-on's DOM-aware crawl
parameters:
context: myapp
url: https://example.com
- type: spiderAjax # the spiderAjax add-on's Crawljax crawl
parameters:
context: myapp
url: https://example.comgo deeper
Know that ZAP has more than one crawler and that two of them drive a real browser, so a single-page app needs one of those rather than the link parser.
Explain the mechanism: spiderAjax drives Crawljax and watches a proxy, spiderClient gets DOM reports from a ZAP browser extension inside the page. Say which add-on owns each.
Show you check what is installed and which crawler a run actually used before reading its URL count, and that you keep the cheap link parser alongside the browser crawl rather than instead of it.
The interesting tradeoff is cost against coverage: browser crawls need browser processes, wall-clock time and memory on every pipeline that runs them, and the coverage they buy is uneven across applications.
## The problem both crawlers exist to solve The traditional `spider` add-on reads a response body and pulls out the things that look like links. Give it a single-page app whose navigation never appears to a link parser — a shell page, a JavaScript bundle, and a router that swaps views on click — and it finds the shell and stops. Nothing in the markup names the other screens, because the other screens are produced by running code, not by serving documents. So ZAP ships two crawlers that run the application instead of reading it. Both launch a real browser through the `selenium` add-on, both click things, and both record what the browser requests. They differ in **where they stand while they watch**. ## spiderAjax: Crawljax, seen through a proxy The `spiderAjax` add-on bundles **Crawljax**, a third-party AJAX crawler, as a library. Crawljax drives the browser through WebDriver: it loads a page, builds a model of the DOM states it can reach, clicks the elements it has been told to click, and treats each resulting state as a node to explore further. ZAP's view of that is entirely external. For each browser it launches, `SpiderThread` asks the `network` add-on to start a local HTTP server, points the WebDriver at it, and registers a handler on it. Every request the browser makes flows through that handler, which is how the URLs reach the history and the Sites tree. What the handler does **not** get is the page: it sees request and response bytes, never the live object graph the running application built. ZAP's own help states it plainly — the AJAX spider "works by launching browsers, clicking links, and filling in fields" but "is not able to directly access the DOM". ## spiderClient: a browser extension that reports the DOM The `client` add-on takes the other route. It ships a first-party **ZAP browser extension** that runs inside the same browser as the page. The extension can read the DOM, and it streams what it finds — components, links, events — back to ZAP over a callback URL that ZAP generates for the session. The add-on's own internals page is blunt about the dependency: without that extension, "this add-on will not be able to do anything". Because the reporting channel is inside the page, `spiderClient` learns about an element whether or not clicking it ever produced a request. Its help says it "has access to the DOM via the ZAP Browser Extension which means that it can find content which the AJAX Spider cannot find", and that it "out performs the older AJAX spider in all known cases". ## The two, side by side | | `spiderAjax` (the `spiderAjax` add-on) | `spiderClient` (the `client` add-on) | |---|---|---| | crawl engine | a bundled Crawljax library | the add-on's own spider | | what ZAP observes | traffic on a local proxy it starts per browser | traffic **plus** DOM reports from the browser extension | | DOM access | none directly, by the project's own statement | yes, through the extension running in the page | | plan job name | `spiderAjax` | `spiderClient` | | project's steer | "no longer the recommended option" | "the recommended option for automatically crawling modern web apps" | ## What this means when you run it from a pipeline 1. **Name the crawler, not "the spider".** Three different things in this program answer to that word: the traditional `spider` add-on, `spiderAjax`, and `spiderClient`. A plan job, a wrapper flag and a log line all mean whichever one you configured. 2. **Check the add-ons are there.** The traditional `spider` add-on is part of a default install; **neither browser crawler is**, and both also need `selenium`. A plan that names `spiderClient` on an install that lacks the `client` add-on does not fall back to a simpler crawl — the job is unknown. 3. **Budget for a browser.** Both crawlers start real browser processes. A container image published without a browser cannot run either of them, however valid the plan is. 4. **Do not treat a browser crawl as a replacement.** The AJAX spider's help recommends combining it with the traditional spider, and the wrapper scripts describe the browser crawl as running "in addition to the traditional one". The link parser is cheap and finds things a click-driven crawl never reaches; the browser crawl finds what the parser cannot see. ## One nuance worth carrying "The AJAX spider cannot see the DOM" is true of the AJAX spider. It is not the whole story about an AJAX spider **run**: the `client` add-on's browser extensions can be loaded into the AJAX spider's browsers, but only if that spider's `enableExtensions` parameter is turned on, and it ships off. When it is on, the `client` add-on watches for URLs the DOM referenced that ZAP never requested, and after the crawl finishes it requests the in-scope ones itself. That is a post-hoc sweep bolted onto the crawl, not the crawl seeing the page — which is exactly why the project stopped recommending the arrangement.
- What has to be installed before a `spiderClient` job will run?The `client` add-on, which carries both the spider and the ZAP browser extension, and the `selenium` add-on that launches the browser. The traditional `spider` add-on is part of a default install; neither browser crawler is, so a plain install crawls without a browser unless you add them.
- Can the AJAX spider get DOM-derived URLs at all?Indirectly. The `client` add-on's browser extensions can be loaded into the AJAX spider's browsers, but only when that spider's `enableExtensions` parameter is on, and it defaults to off. With it on, the `client` add-on collects URLs the DOM referenced but ZAP never requested and, once the crawl ends, requests the in-scope ones itself.
- Does a browser crawl make the traditional spider redundant?No. The AJAX spider's own help recommends combining it with the traditional spider, and the wrapper scripts run the browser crawl *in addition to* the link parser. Parsing responses is far cheaper and reaches plain documents, sitemap entries and anything not behind a click; the browser crawl covers what only running the app reveals.
The AJAX spider watches a room from the doorway: it sees everyone who walks in and out, and infers the furniture. The client spider has someone standing inside describing the walls.
saying these in an interview costs you the question
- Says the AJAX spider reads the DOM directly
- Calls the client spider just a faster AJAX spider
- Assumes either browser crawler ships in a default install
- Treats a browser crawl as replacing the traditional spider
- Thinks Crawljax is ZAP code rather than a bundled library