Your ZAP plan's `spider` job finished clean but barely crawled - what should have caught that?
answer
- completed is not the same as reached
- a counter of URLs added
- the default test is desktop-only
- info records, error stops
basics
~20 sA stats test you wrote yourself on automation.spider.urls.added. The spider job publishes that counter, but a headless plan with no tests block checks nothing, and the test in the shipped templates has an informational failure level.
solid answer
~40 sThe `spider` job increments `automation.spider.urls.added` as it adds URLs, and a `stats` test on that counter is the check that catches a thin crawl. Two things stop it catching yours by accident. The job does define a default test, but that default is attached only when a human creates a plan in the desktop New Plan dialog - a YAML plan loaded and run headless never gets it, so a plan with no `tests` block runs no test at all. And the example test in the shipped `spider-min.yaml` and `spider-max.yaml` templates ships with `onFail: 'info'`, which records a line and leaves the outcome alone. Write the test, set a floor you can defend, and set `onFail: error`.
code
yaml · 9 lines- type: spider
parameters:
url: https://example.com
tests:
- type: stats
statistic: automation.spider.urls.added
operator: '>='
value: # an integer floor, from a known-good run
onFail: error # the templates ship 'info', which never blocksgo deeper
Know that a plan can finish clean without the crawl having reached much, and that the number of URLs the crawl added is the first thing to look at when the results seem too quiet.
Name the counter, and explain why a headless plan checks it only if the plan itself declares a test - the job's own default is attached by the desktop plan dialog, not by the run.
Expect to design the gate: pick a floor you can defend from a known-good run, set the failure level so a shortfall stops the pipeline, and say what you do when the floor legitimately moves.
Own the policy question - which pipeline stages may block a release on a coverage floor, who owns the number, and how it is revised without decaying into a rubber stamp.
## The counter the job publishes ZAP's automation framework keeps a set of named counters during a run, and jobs publish into it. The traditional `spider` job publishes one: **`automation.spider.urls.added`**, incremented as the crawl adds URLs it has not seen before. That single number is the closest thing a plan has to a statement of how much of the application the run actually reached. It matters because a crawl that reached almost nothing looks, from the outside, exactly like a crawl of a small clean application. The job completes. Whatever runs afterwards works over the history the crawl produced. If that history is a login page and a favicon, the later phases find nothing - and "found nothing" reads like good news. ## The default test is a desktop feature The `spider` job does something most jobs do not: it defines a default test on itself, a `stats` test on that counter with a `>=` operator. It is easy to assume this runs everywhere. **It does not.** The only caller of that method is the automation add-on's **New Plan dialog** - the desktop window a human uses to create a plan by picking jobs from a list. The dialog attaches the test to each job it adds, and the test is then written into the YAML you save. Load a plan from a file and run it headless and nothing calls that method. **A `spider` job whose plan declares no `tests` block is checked by nothing at all.** This is the same shape as several other safeguards in this program: the net is hung under the desktop rather than under the pipeline, and the unattended path - which is the one a build server takes - walks past it. ## And the shipped example would not block either The example test in the job's `spider-min.yaml` and `spider-max.yaml` templates, which is what most people copy, is correct in shape and wrong in severity for a gate: | field | shipped value | what a gate wants | |---|---|---| | `type` | `stats` | unchanged | | `statistic` | `automation.spider.urls.added` | unchanged | | `operator` | `>=` | unchanged | | `value` | a starting number the template tells you to replace | a floor derived from a known-good run | | `onFail` | `info` | `error`, so a shortfall stops the plan | `info` records an informational line in the plan's progress and changes nothing about the outcome. That is a reasonable default for a tool that must not break everyone's first run, and a poor one for a pipeline that treats a clean plan as evidence. ## Why the crawl is where this bites A successful outcome says a run *completed*, not that it *reached* anything, and the crawl is where that distinction is created, because every later phase consumes what it produced. A crawl can come back near-empty because: 1. the URL was wrong, or redirected off the host, so the seed produced nothing to follow; 2. authentication did not take, and every page behind the login form was the login form; 3. a context or an exclusion rule removed the application, leaving only its shell; 4. the application is client-rendered, and this crawler found the shell and stopped; 5. a depth, children or duration limit fired long before coverage was reached. All five end the same way: a job that completed, and a URL count far below what the application should yield. ## The superseded parameters Older plans set `failIfFoundUrlsLessThan` or `warnIfFoundUrlsLessThan` on the job. Those fields still parse, and the job's verification step emits a warning naming the statistic test as their replacement - **and then nothing else happens.** They are read into fields the run never consults. A plan you inherited that relies on one has not been gating anything for some time. ## What the job's tests can actually be The shipped `spider` templates carry a comment saying only a `stats` test is supported for now. **That comment understates the code.** The framework's test loader accepts `stats`, `alert`, `url` and `monitor` tests on any job, gating only `monitor` behind a per-job capability flag - and the spider job declares that it supports monitor tests. That makes a `url` test available, and it is often the better gate. It resolves the URL you name against the **site tree** the crawl fills, with optional request and response regexes, so "the account settings page was reached" survives an application growing or shrinking, where a raw number does not. So the working checklist is short: - **Declare the `stats` test in the plan file.** Never rely on the job attaching one for you. - **Derive the floor from a known-good run**, not from a round number someone liked. - **Set `onFail: error`** on the run that is allowed to block, and leave `info` for exploratory runs. - **Add a `url` test** for one or two paths that must always be reached, so a structural failure is named rather than inferred from a count. - **Re-derive the floor when the application changes shape**, and record why it moved - a threshold nobody maintains becomes a rubber stamp within two releases.
- Does a ZAP `spider` job with no `tests` block run any test at all?Not in a headless run. The job defines a default statistic test, but the only caller is the desktop New Plan dialog, which attaches it when a human creates the plan. A plan loaded from YAML and run with the automation command line gets exactly the tests the file declares.
- What happens if you still set `failIfFoundUrlsLessThan` on a ZAP spider job?The job's verification step warns that the field has been replaced by the statistic test, and nothing else changes - the value is parsed into a field the run never reads. Move the intent into a `stats` test on `automation.spider.urls.added`.
saying these in an interview costs you the question
- A plan with no tests block still checks the crawl by default
- The template's URL-count test would fail a plan on a thin crawl
- Set failIfFoundUrlsLessThan to gate the spider job
- A spider job can only carry stats tests
- A clean plan outcome proves the crawl reached the application