skip to content

Traditional Spider

The `spider` add-on fetches and parses responses without ever running JavaScript, mining links out of HTML, forms, sitemaps and robots.txt. Interviewers ask what it therefore never sees.

on this pageshow

questions

5

In ZAP's traditional `spider` add-on, what does `parseRobotsTxt: true` do with a `Disallow` line?

level: middleimportance: must knowfreq 68%

answer

  1. a hint list, not a rulebook
  2. the file is fetched unprompted
  3. Allow and Disallow share one path
  4. no User-agent grouping at all
  5. the shipped help says it outright

basics

~20 s

It mines it. With parseRobotsTxt on, the spider seeds /robots.txt at the host root, then reads Disallow and Allow lines through the same code path and queues each path as a new target. It never obeys them.

solid answer

~40 s

`parseRobotsTxt` (default on) does two things, and neither is obedience. First, adding a seed also adds `/robots.txt` at that seed's host root, so the file is requested even when nothing on the site links to it. Second, `SpiderRobotstxtParser` strips comments and markup, then matches `Disallow:` and `Allow:` lines and hands both to the *same* `processPath`, which trims a trailing `*`, resolves what is left against the host root and queues it as a target. There is no `User-agent:` handling at all, so a rule written for one crawler reads like any other. ZAP's own shipped help says it outright: the Spider does not follow the rules specified in the robots.txt file.

code

yaml · 5 lines
yaml
- type: spider
  parameters:
    url: https://example.com
    parseRobotsTxt: true    # default: seeds /robots.txt, then mines it for URLs
    parseSitemapXml: true   # default: seeds /sitemap.xml the same way

go deeper

for a junior

Remember the direction. The option mines robots.txt for URLs; it does not make the crawl obey the file. Treat that file as a list of places the site would rather nobody looked.

for a middle

Describe both effects: the robots file is added as a seed at the host root, and Disallow and Allow lines are turned into targets by the same code, with no User-agent handling anywhere.

for a senior

Expect to explain the operational consequence on a shared or pre-production host: the crawl walks straight into the paths robots.txt advertises, so the boundary has to live in the spider's exclude list, which is checked before the request is made.

for a principal

The tradeoff to own is whether an unattended crawl may treat a site's own exclusion hints as a target list at all, and how that permission is recorded for each environment it runs against.

## What the option actually switches on `parseRobotsTxt` lives on ZAP's traditional **`spider` add-on** and defaults to on. Its name suggests a parsing preference. It is really two behaviours, and the first one surprises people: 1. **It seeds a request.** When a seed URL is added to a spider run, the add-on also adds `/robots.txt` at that seed's host root to the seed list. The file is fetched because the option is on, not because anything linked to it. 2. **It mines the response.** The robots parser claims a response whose path is exactly `/robots.txt` and turns the file's contents into new crawl targets. Turning the option off removes both halves: no seed is added, so the spider itself never asks for the file. The sitemap option is wired the same way for `/sitemap.xml`. ## How the file is read The parser walks the body line by line. For each line it cuts everything from a `#` onward, strips any HTML tags, trims the result, and skips the line if nothing is left. Then it applies two case-insensitive patterns: | line | what the parser does | |---|---| | `Disallow: /admin/` | queues `/admin/` against the host root | | `Allow: /public/` | queues `/public/` against the host root - **the same code path** | | `Disallow: /admin/*` | trims the trailing `*`, queues `/admin/` | | `Disallow:` (empty) | produces an empty path and is skipped | | `User-agent: *` | matches neither pattern; ignored entirely | | `Sitemap: https://example.com/sm.xml` | matches neither pattern; ignored entirely | The two branches call one shared method. **A `Disallow` and an `Allow` are indistinguishable to this parser** - both are just a path worth visiting. The parser then marks the message fully consumed so no fall-back parser re-reads it. ## What it therefore never does - It never groups rules by crawler. There is no `User-agent:` handling, so the notion of "the rules that apply to me" does not exist in this code. - It never reads a `Sitemap:` directive. A sitemap advertised at a non-root path from robots.txt is not picked up; the sitemap option seeds `/sitemap.xml` at the root instead, and nowhere else. - It handles only a **trailing** asterisk. A wildcard in the middle of a pattern is carried into the URL literally, producing a target that usually 404s. - It claims only the exact path `/robots.txt`. A robots file served from a subdirectory is handled by whichever parser would ordinarily claim that content type. ## The exemption nobody expects The spider's parse filter normally drops a response before parsing if it exceeds the maximum parse size, or if its content type is not text. **`robots.txt` is checked ahead of both** - alongside `sitemap.xml`, SVN and Git metadata files and `.DS_Store`. So a robots file served as `application/octet-stream`, or one far past the size cap, is still parsed in full. If you have ever wondered why a crawl exploded after touching a machine-generated robots file with thousands of lines, this is why. ## Why this is the fact to carry into a pipeline A robots.txt is written by someone listing the paths they would rather the world did not index: staging areas, admin consoles, export endpoints, old upload directories. That makes it, from a crawler's point of view, an unusually high-value index of exactly the material nobody advertises. 1. The spider requests the file unprompted and turns every listed path into a target. 2. Those targets are recorded in history like any other crawl result. 3. Whatever runs afterwards on that history - a passive sweep, an active run - works from a list the site owner meant to keep quiet. None of that is a defect. It is the documented behaviour of a security testing tool, and the project says so in its own help. What it means for you is that **robots.txt is not a boundary**. If a path must not be touched on a shared environment, express it in the spider's own exclude list, which its fetch filter evaluates before a request is ever created, or scope the run so the path is out of reach. ## The ownership footnote The traditional spider is not in ZAP's core. It moved out to the `spider` add-on and, unusually for this project, **left no deprecated shim behind** - there is no spider implementation package in core at all. Core did keep the seams the add-on plugs into: the spider request initiator, the spider history types, and the session-level exclude-from-spider list with its desktop panels. So "nothing survives in core" is too strong; what survives is sockets, not an engine.

  • Does turning `parseRobotsTxt` off stop ZAP's spider requesting `/robots.txt`?
    Yes. The root-file seed is added only inside that check, so with the option off the spider generates no request for the file. `parseSitemapXml` gates `/sitemap.xml` identically. The file can still reach history by another route, such as a link on the site or a request you made yourself through the proxy.
  • What does ZAP's robots parser do with `Disallow: /admin/*`?
    It trims the trailing `*`, trims whitespace, and queues `/admin/` resolved against the host root. Only a trailing asterisk is handled: a wildcard in the middle of a pattern is carried into the URL literally. A bare `Disallow:` with no path yields an empty string and is skipped.
  • Is `/robots.txt` subject to the spider's maximum parse size?
    No. The parse filter returns early for `robots.txt`, `sitemap.xml`, SVN and Git metadata and `.DS_Store`, ahead of both the response-size check and the text-content-type check. A robots file served with a binary content type, or one far past the cap, is still parsed in full.

A robots.txt is a staff-only sign at the end of a corridor. The traditional spider does not read it as an instruction; it reads it as a directory of corridors worth walking down.

saying these in an interview costs you the question

  • parseRobotsTxt makes the spider respect robots.txt
  • ZAP only reads robots.txt if something links to it
  • Only Disallow lines are mined; Allow lines are skipped
  • The spider honours the User-agent group aimed at it
  • A Sitemap line in robots.txt tells the spider where the sitemap is
  • A huge or binary robots.txt is dropped by the parse-size limit
open as a page

Does ZAP's traditional `spider` add-on submit the HTML forms it finds, and what does `postForm` change?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Yes - the traditional spider submits forms while it crawls. processForm, on by default, is the master switch; postForm, also on, covers only POST-method forms, so turning it off still leaves GET forms built into URLs and requested.

open as a page

ZAP's traditional spider never executes JavaScript - so what does it extract from a `.js` response?

level: middleimportance: should knowfreq 60%

basics

~20 s

Absolute URLs only. A script body is text but not HTML, so the spider's fall-back text parser scans it with a regular expression for http and https strings. A relative path assembled in code is invisible to it.

open as a page

In ZAP's `spider` job, what does `handleParameters` actually control - and what does it not?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Only the already-visited key. use_all, ignore_value and ignore_completely decide how much of a query string counts when the spider asks whether it has seen a URL before. They do not decide which parameters anything later attacks.

open as a page

Your ZAP plan's `spider` job finished clean but barely crawled - what should have caught that?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

A stats test you wrote yourself on automation.spider.urls.added. The spider job publishes that counter, but a headless plan with no tests block checks nothing, and the test in the shipped templates has an informational failure level.

open as a page