skip to content

Do ZAP's spiderAjax and spiderClient crawls only read pages, or do they type into forms and submit them?

level: middleimportance: nice to knowfreq 34%

answer

  1. A browser crawl interacts, it does not just read
  2. One crawler has the dials, the other has none
  3. The default is to type into fields
  4. The flag selects a rule source, not a limit

basics

~20 s

Both browser crawls type into forms and submit them. The spiderAjax job's randomInputs parameter, on by default, fills inputs with random data; the client spider fills every displayed input and textarea before each click, with no switch at all.

solid answer

~40 s

Neither browser crawl is read-only. The `spiderAjax` job exposes `randomInputs`, which defaults to **on** and tells the bundled crawler to put random values into input elements, and `clickDefaultElems`, which selects which click rules apply — `true` uses the crawler's own default set of anchors, buttons and submit-or-button inputs, `false` uses the job's `elements` list instead. The `client` add-on's spider has **neither parameter**. Before any non-passive click it locates every displayed `input` and `textarea` on the page, clears it and types a value supplied by ZAP's value provider, then clicks. So a browser crawl fills registration forms, contact forms and search boxes and presses the buttons next to them, and only one of the two crawlers gives you a way to stop it.

go deeper

for a junior

Take away that these crawls are not read-only: they fill in fields and press buttons, so the target has to be an environment where that is acceptable and one you are permitted to test.

for a middle

Explain the mechanics: randomInputs and clickDefaultElems are AJAX spider parameters, and the client spider fills every displayed input and textarea before each click with no switch of its own.

for a senior

Show that you weigh the tradeoff rather than reflexively turning interaction down, since a crawl that stops submitting forms stops reaching whatever is behind them.

for a principal

The judgement is where these runs are allowed to point. If no crawler setting can make a run non-destructive, the guarantee has to come from the environment, and that is a standing decision rather than a per-pipeline one.

## The assumption worth demolishing "Crawling is read-only, scanning is where the writes happen" is a comfortable model and it is wrong for these two crawlers. A crawl that drives a real browser has to interact with the application to find anything — that is the entire point of using a browser — and interacting means clicking controls and filling in the fields next to them. ## What `spiderAjax` exposes Two job parameters govern the interaction: - **`randomInputs`** — on by default. It tells the bundled crawler to insert random data into input elements before acting on them. - **`clickDefaultElems`** — on by default, and its name misleads. It does not mean "restrict the crawl"; it selects **which source of click rules is used**. Left on, the bundled crawler uses its own default set: anchors, buttons, and `input` elements whose `type` is submit or button. Turned off, ZAP supplies the job's `elements` list instead. There is a subtlety in that second one. In an automation plan, turning `clickDefaultElems` off without supplying an `elements` list hands the crawler an empty rule set — and the crawler's own behaviour on an empty set is to fall back to its defaults. So flipping the flag alone changes nothing. The `elements` key is the half that does the work, and the flag only unlocks it. ## What `spiderClient` exposes Nothing equivalent. The `client` add-on's spider declares no `randomInputs` and no `clickDefaultElems`, and the job accepts neither. Its behaviour is fixed: before any non-passive click, it finds every displayed `input` and `textarea` on the page, clears each one and types a value obtained from ZAP's value provider, then performs the click. Form submissions go through the same path. | | `spiderAjax` | `spiderClient` | |---|---|---| | fills inputs | yes, when `randomInputs` is on (the default) | yes, always | | switch to stop it | `randomInputs: false` | none in the job | | which elements get clicked | `clickDefaultElems` plus `elements` | not configurable by tag | | deterministic don't-click list | `excludedElements` | none in the job | | values used | random data from the bundled crawler | values from ZAP's value provider | The asymmetry runs one way throughout: the older crawler is the configurable one, and the newer crawler trades those dials for better coverage. ## Why a QA engineer should care 1. **The crawl writes.** Pointed at an application with a sign-up form, a browser crawl will fill it and submit it, repeatedly, from several browsers at once. Rows appear. Emails may be sent. Workflows may start. 2. **Shared environments are the risk.** A crawl that creates junk in a shared test environment is an incident for whoever else is using it that afternoon, long before anything security-related is involved. 3. **"Baseline" does not mean "safe".** A packaged run that sends no attacks is still a run that types into fields and presses buttons. The passive-versus-active distinction is about what analyses the traffic, not about whether the crawl touches the application. 4. **It is not an authorisation.** Every one of these clicks is a request to a real system. A browser crawl belongs only against a target you were explicitly permitted to test, and "it is only the crawl phase" is not a mitigation you can offer afterwards. ## Turning it down, and what that costs On the AJAX spider job you can set `randomInputs: false`, restrict clicking through `clickDefaultElems: false` plus a narrow `elements` list, and name specific controls in `excludedElements`. Every one of those reduces reach: an application whose next screen only appears after a valid form submission will not be explored if the crawl stops filling forms. That is the honest tradeoff, and it is a decision about the environment rather than about the tool. On the client spider there is no such dial, so the levers are different ones: scope the crawl tightly, exclude the URLs behind destructive actions, and choose an environment you are content to see written to. If a crawl genuinely must not write, that constraint has to be met outside the crawler — by where you point it, not by how you configure it. ## The sentence to carry Both browser crawlers interact with the application by design. One lets you turn the typing off and the other does not, so the safe default assumption for either is: **this run will fill in forms and press buttons, and you should only point it somewhere that is fine.**

  • Does setting clickDefaultElems to false on its own narrow the crawl?
    No. In a plan it hands the bundled crawler an empty rule set, and an empty set makes that crawler fall back to its own defaults — anchors, buttons and submit-or-button inputs. The change only takes effect once you also supply the `elements` list the flag exists to unlock.
  • Can you stop the client spider filling in forms?
    Not through the job. It has no `randomInputs` equivalent and fills every displayed input and textarea before each non-passive click. If writes are unacceptable, the control has to be the environment and the scope you point it at, not a crawler setting.
  • Where do the values typed into fields come from?
    The AJAX spider's bundled crawler inserts random data when `randomInputs` is on. The client spider asks ZAP's value provider for each field, passing the field's name, its existing value and its control type, so what it types is shaped by the field rather than purely random.

saying these in an interview costs you the question

  • Calls a crawl read-only because it sends no attacks
  • Assumes both crawlers expose a randomInputs switch
  • Reads clickDefaultElems as a restriction rather than a rule source
  • Points a browser crawl at a shared environment without asking
  • Thinks a passive-only run cannot change application state