skip to content

Does ZAP's traditional `spider` add-on submit the HTML forms it finds, and what does `postForm` change?

level: juniorimportance: should knowfreq 58%

answer

  1. discovery is not read-only
  2. two switches, not one
  3. non-POST falls through to GET
  4. a ValueGenerator fills the blanks
  5. a field's own value wins

basics

~20 s

Yes - the traditional spider submits forms while it crawls. processForm, on by default, is the master switch; postForm, also on, covers only POST-method forms, so turning it off still leaves GET forms built into URLs and requested.

solid answer

~40 s

ZAP's traditional `spider` add-on does not just collect links. `SpiderHtmlFormParser` builds a submission for every `<form>` it parses, and queues it as a request. `processForm` (default on) decides whether forms are handled at all; `postForm` (default on) decides only whether actions whose resolved method is POST get sent. Anything that is not POST - including a form with no `method` attribute - falls through to the GET branch, where the fields are appended to the action as a query string and requested. Field values come from a `ValueGenerator`: a field's own `value` attribute is submitted verbatim, and empty fields get type-appropriate placeholders. That is how a hidden record id or action name is replayed exactly as the page shipped it.

code

yaml · 5 lines
yaml
- type: spider
  parameters:
    url: https://example.com
    processForm: true   # handle forms at all (default)
    postForm: false     # suppress POST actions only; GET forms still submitted

go deeper

for a junior

Know that the crawl writes. The traditional spider fills in and submits forms, so pointing it at a shared environment can create, change or delete real records.

for a middle

Separate the two switches. processForm turns form handling on or off entirely; postForm suppresses only actions that resolve to POST, and everything else is still submitted as a GET request.

for a senior

Expect to explain how you keep an unattended crawl off destructive endpoints: the spider's own exclude list is evaluated by its fetch filter before any request is created, so a match is never sent.

for a principal

Own the question of whether an unattended crawl may submit forms against a given environment at all, and where that permission is recorded - in the plan, in the environment definition, or in the target's data lifecycle.

## The crawl writes ZAP's traditional crawler is the **`spider` add-on**. It left the core program and core kept no implementation behind it - only the seams the add-on plugs into, such as the spider request initiator, the spider history types and the session's exclude-from-spider list. What the add-on does is fetch a response and hand it to a chain of parsers. One of those parsers does something the word "crawl" does not suggest. **`SpiderHtmlFormParser` builds a submission from each `<form>` it parses and queues it as a new request** - subject to two options, neither of which is off by default. A discovery run against an application carrying a "delete account" form, an "export everything" form or a "send invitation" form will submit them, with values it invented, unless you stop it. Calling the spider read-only is a common and consequential wrong model of this tool. ## The two switches, and the gap between them | option | default | what it governs | what it leaves alone | |---|---|---|---| | `processForm` | on | whether forms are parsed and submitted at all | ordinary link extraction from the rest of the page | | `postForm` | on | only actions whose resolved method is POST | GET-method actions, which are still built and requested | `processForm` is the master switch: turn it off and the form parser returns immediately, having read no form on the page. `postForm` is narrower than its position beside it suggests. The parser resolves a method for each action it derives from a form and applies the `postForm` test to that resolved method. Everything that is not POST falls through to the GET branch: - a form with `method="post"` becomes a POST body, and that one is suppressed when `postForm` is off; - a form with `method="get"` has its fields appended to the action as a query string, and is always requested; - a form with **no** `method` attribute is treated as GET; - a form with a method string the parser does not recognise is also treated as GET. There is no "unsupported method, skip it" branch. Switching `postForm` off narrows what the crawl writes; it does not make the crawl read-only. ## One form is not one request The parser expands a single form into one action per submit control, honouring a button's `formaction` and `formmethod` attributes when `formmethod` names GET or POST. Three consequences follow: 1. a form with three submit buttons aimed at three actions yields three submissions, not one; 2. a POST form whose button declares `formmethod="get"` yields a GET submission - which `postForm` never suppresses; 3. a GET form whose button declares `formmethod="post"` yields a POST submission, which `postForm` does suppress. So the option is evaluated per resolved action, not per `<form>` tag, and reasoning about it from the markup alone will mislead you. ## Where the submitted values come from Values are produced by a **`ValueGenerator`**, and the default implementation is a core class the add-on calls into. Its first rule is the one that matters: **if the field already carries a `value`, that value is submitted verbatim.** - a hidden field holding a record identifier, an action name or a state token is replayed exactly as the page shipped it; - an empty text, password or search field gets a fixed placeholder string; - a `number` or `range` field gets its declared `min`, failing that its `max`, failing that a small default; - `email`, `url`, `color`, `tel` and the whole date and time family get synthetic values of the right shape; - a file field gets a placeholder file name rather than a real upload. The requests that result are not noise. They are plausible enough for the application to accept and act on them, which is exactly why an unattended crawl can leave a trail of created, modified or deleted records behind it. ## What to do about it in a pipeline - Decide per environment whether an unattended crawl may submit anything at all. `processForm: false` is the honest setting for a shared or production-like target. - If you want form coverage without writes, treat `postForm: false` as a partial measure and read the GET list above before you rely on it. - Where one endpoint must never be touched, use the spider's own **exclude list**. The spider merges the extension's list, the session's exclude-from-spider regexes and the global exclusion regexes, and its fetch filter rejects a match before a task is created - so nothing is sent, which is not true of every exclusion mechanism in this program. - Crawl as an account whose authorisation is scoped to what you are willing to have exercised, rather than as an administrator. The interview version of this question is a check on whether you have actually watched a crawl run, or only read its name.

  • In ZAP's spider, is the `postForm` check applied per form or per submit button?
    Per resolved action. `SpiderHtmlFormParser` expands one `<form>` into an action per submit control, honouring each button's `formaction` and `formmethod` when that names GET or POST, and tests `postForm` against the resolved method. A POST form with a button declaring `formmethod="get"` therefore still produces a GET submission, which `postForm` does not suppress.
  • Does ZAP's spider submit a form that has no `method` attribute?
    Yes. Only a resolved method of POST takes the POST path. Everything else - a missing `method`, or a value the parser does not recognise - falls through to the GET branch, where the fields are appended to the action as a query string and the URL is requested. There is no skip case.

saying these in an interview costs you the question

  • A spider only sends GET requests, so it cannot change data
  • Turning off postForm makes the crawl read-only
  • The spider leaves form fields empty when it submits
  • Only forms with an explicit method attribute get submitted
  • Hidden fields are stripped before the spider submits a form