skip to content

In ZAP's `spider` job, what does `handleParameters` actually control - and what does it not?

level: seniorimportance: should knowfreq 45%

answer

  1. it decides what counts as duplicate
  2. it feeds the visited set
  3. query kept, named, or dropped
  4. use_all is the default

basics

~20 s

Only the already-visited key. use_all, ignore_value and ignore_completely decide how much of a query string counts when the spider asks whether it has seen a URL before. They do not decide which parameters anything later attacks.

solid answer

~40 s

`handleParameters` feeds exactly one thing: the string the spider uses to decide a URL is a duplicate. `UrlCanonicalizer` turns a canonical URI into that representation - `use_all` keeps the whole query, `ignore_value` keeps parameter names and drops their values, `ignore_completely` drops the query altogether - and the spider's controller concatenates the method, that string, the canonical headers and a normalised body into the identifier it checks against its visited set. The consequence is coverage, not depth. On `use_all`, the default, a faceted listing or a calendar mints a fresh page per value and the crawl runs until a depth, children or duration limit stops it. On `ignore_completely` you get one sample per path, cheaply, and a parameter-specific behaviour is never recorded at all.

go deeper

for a junior

Know the name and the direction: this option controls when the spider decides it has already been somewhere, which is why a crawl can either loop on query strings or skip straight past them.

for a middle

Walk the three values and say what each keeps in the visited key - the whole query, the parameter names only, or nothing - and name the default.

for a senior

Connect the setting to an outcome you have seen: a crawl that never finished on a faceted listing, or one that returned fast and left a parameter-specific behaviour unrecorded, and say how you told which it was.

for a principal

The judgment to own is how much crawl budget a pipeline may spend and what coverage claim its result may support, given that this one dial trades those against each other.

## What the option is wired to There are three values, spelled `use_all`, `ignore_value` and `ignore_completely` in an automation plan and in upper case in the add-on's own enum. It is tempting to read them as a statement about which parameters matter. They are not. **The option feeds exactly one thing: the string the spider uses to recognise a URL it has already visited.** The spider keeps a set of resource identifiers. Before it creates a task for a discovered resource it builds an identifier and checks the set; a hit means the resource is dropped with a "already visited" note and no request is sent. The identifier is not just a URL. It is the **request method**, then the cleaned URI representation this option shapes, then a canonical rendering of the headers the resource was found with, then a normalised form of the body. That last detail explains behaviour people find surprising: the same path reached by a link and by a form POST are two visits, and two POSTs to one action with different bodies are two visits as well. ## The three values | value | what survives into the visited key | effect on the crawl | |---|---|---| | `use_all` (default) | the whole URI, query string and all | every distinct value pair is a new page; maximum coverage, unbounded growth | | `ignore_value` | the path plus the parameter **names** | one visit per parameter *shape*; `?sort=asc` and `?sort=desc` collapse | | `ignore_completely` | the path only | one visit per path; the entire query dimension disappears from the crawl | The add-on's own shipped plan template describes the option in precisely these terms - how query string parameters are used when checking if a URI has already been visited - so the documentation is not the source of the confusion. ## What it is not - It does **not** decide which parameters an attack phase later exercises. That is a separate configuration on a separate component, and setting this option will not turn cookie or header testing on or off. - It does **not** strip the query from the requests the spider sends. A URL that was discovered is requested as it was found. What changes is whether a *second* URL differing only in its query is recognised as already seen. - It does **not** act alone. Depth, maximum children per node and maximum duration are separate brakes, and on a large parameterised site you usually need this option *and* one of those. ## Choosing it in a pipeline 1. **Start at the default and measure.** If a crawl finishes well inside its duration limit and the URL count looks like the application, leave `use_all` alone - it gives the richest history for anything that runs afterwards. 2. **Move to `ignore_value` when the count is dominated by one parameter.** Faceted search, pagination, sort orders, calendars and session-scoped tokens in the query are the usual culprits. You keep one representative request per parameter shape. 3. **Reserve `ignore_completely` for smoke-shaped runs.** It is cheap and predictable, and it is the setting most likely to make a later phase miss something, because only one sample per path ever reaches history. ## The knock-on effect that is easy to miss Everything downstream works from what the crawl recorded. Collapse the query dimension and a defect that only appears for one parameter value is not "rated low" - it is **absent**, because no request that would have revealed it was ever made. Conversely, leaving `use_all` on a site that mints parameters will burn the whole crawl budget in one corner of the application and leave the rest unvisited by the time a limit finally stops the run - or leave it running indefinitely if you set no duration or children limit at all. Both failure modes produce a run that completes normally. Neither announces itself. The way to tell them apart is to look at how many URLs the crawl added and where they clustered, not at whether the job reported success. ## The interview shape An interviewer asking this is usually checking one thing: whether you know the difference between a **discovery** control and an **attack** control. A candidate who answers "it decides which parameters get tested" has merged two phases that are configured separately, run on different engines and fail in different ways. Say what the option keys, say what it changes about the request stream, and then say what it does not touch.

  • What else besides the cleaned URL goes into the spider's visited-set identifier?
    The request method, a canonical rendering of the headers the resource was found with, and a normalised form of the body. That is why the same path reached by a link and by a form POST counts as two visits, and why two POSTs with different bodies are not collapsed into one.
  • With `ignore_completely` set, does the spider still send the query string it found?
    Yes. The option only shapes the string used for the duplicate check. The URL that was discovered is requested as found, query and all. What changes is that a second URL differing only in its query is treated as already visited and never requested.

saying these in an interview costs you the question

  • handleParameters chooses which parameters get attacked
  • ignore_completely stops the spider sending query strings
  • The visited key is just the URL the spider found
  • Changing handleParameters is only a performance setting
  • Two POSTs to one action with different bodies count as one visit