In ZAP's traditional `spider` add-on, what does `parseRobotsTxt: true` do with a `Disallow` line?
answer
- a hint list, not a rulebook
- the file is fetched unprompted
- Allow and Disallow share one path
- no User-agent grouping at all
- the shipped help says it outright
basics
~20 sIt mines it. With parseRobotsTxt on, the spider seeds /robots.txt at the host root, then reads Disallow and Allow lines through the same code path and queues each path as a new target. It never obeys them.
solid answer
~40 s`parseRobotsTxt` (default on) does two things, and neither is obedience. First, adding a seed also adds `/robots.txt` at that seed's host root, so the file is requested even when nothing on the site links to it. Second, `SpiderRobotstxtParser` strips comments and markup, then matches `Disallow:` and `Allow:` lines and hands both to the *same* `processPath`, which trims a trailing `*`, resolves what is left against the host root and queues it as a target. There is no `User-agent:` handling at all, so a rule written for one crawler reads like any other. ZAP's own shipped help says it outright: the Spider does not follow the rules specified in the robots.txt file.
code
yaml · 5 lines- type: spider
parameters:
url: https://example.com
parseRobotsTxt: true # default: seeds /robots.txt, then mines it for URLs
parseSitemapXml: true # default: seeds /sitemap.xml the same waygo deeper
Remember the direction. The option mines robots.txt for URLs; it does not make the crawl obey the file. Treat that file as a list of places the site would rather nobody looked.
Describe both effects: the robots file is added as a seed at the host root, and Disallow and Allow lines are turned into targets by the same code, with no User-agent handling anywhere.
Expect to explain the operational consequence on a shared or pre-production host: the crawl walks straight into the paths robots.txt advertises, so the boundary has to live in the spider's exclude list, which is checked before the request is made.
The tradeoff to own is whether an unattended crawl may treat a site's own exclusion hints as a target list at all, and how that permission is recorded for each environment it runs against.
## What the option actually switches on `parseRobotsTxt` lives on ZAP's traditional **`spider` add-on** and defaults to on. Its name suggests a parsing preference. It is really two behaviours, and the first one surprises people: 1. **It seeds a request.** When a seed URL is added to a spider run, the add-on also adds `/robots.txt` at that seed's host root to the seed list. The file is fetched because the option is on, not because anything linked to it. 2. **It mines the response.** The robots parser claims a response whose path is exactly `/robots.txt` and turns the file's contents into new crawl targets. Turning the option off removes both halves: no seed is added, so the spider itself never asks for the file. The sitemap option is wired the same way for `/sitemap.xml`. ## How the file is read The parser walks the body line by line. For each line it cuts everything from a `#` onward, strips any HTML tags, trims the result, and skips the line if nothing is left. Then it applies two case-insensitive patterns: | line | what the parser does | |---|---| | `Disallow: /admin/` | queues `/admin/` against the host root | | `Allow: /public/` | queues `/public/` against the host root - **the same code path** | | `Disallow: /admin/*` | trims the trailing `*`, queues `/admin/` | | `Disallow:` (empty) | produces an empty path and is skipped | | `User-agent: *` | matches neither pattern; ignored entirely | | `Sitemap: https://example.com/sm.xml` | matches neither pattern; ignored entirely | The two branches call one shared method. **A `Disallow` and an `Allow` are indistinguishable to this parser** - both are just a path worth visiting. The parser then marks the message fully consumed so no fall-back parser re-reads it. ## What it therefore never does - It never groups rules by crawler. There is no `User-agent:` handling, so the notion of "the rules that apply to me" does not exist in this code. - It never reads a `Sitemap:` directive. A sitemap advertised at a non-root path from robots.txt is not picked up; the sitemap option seeds `/sitemap.xml` at the root instead, and nowhere else. - It handles only a **trailing** asterisk. A wildcard in the middle of a pattern is carried into the URL literally, producing a target that usually 404s. - It claims only the exact path `/robots.txt`. A robots file served from a subdirectory is handled by whichever parser would ordinarily claim that content type. ## The exemption nobody expects The spider's parse filter normally drops a response before parsing if it exceeds the maximum parse size, or if its content type is not text. **`robots.txt` is checked ahead of both** - alongside `sitemap.xml`, SVN and Git metadata files and `.DS_Store`. So a robots file served as `application/octet-stream`, or one far past the size cap, is still parsed in full. If you have ever wondered why a crawl exploded after touching a machine-generated robots file with thousands of lines, this is why. ## Why this is the fact to carry into a pipeline A robots.txt is written by someone listing the paths they would rather the world did not index: staging areas, admin consoles, export endpoints, old upload directories. That makes it, from a crawler's point of view, an unusually high-value index of exactly the material nobody advertises. 1. The spider requests the file unprompted and turns every listed path into a target. 2. Those targets are recorded in history like any other crawl result. 3. Whatever runs afterwards on that history - a passive sweep, an active run - works from a list the site owner meant to keep quiet. None of that is a defect. It is the documented behaviour of a security testing tool, and the project says so in its own help. What it means for you is that **robots.txt is not a boundary**. If a path must not be touched on a shared environment, express it in the spider's own exclude list, which its fetch filter evaluates before a request is ever created, or scope the run so the path is out of reach. ## The ownership footnote The traditional spider is not in ZAP's core. It moved out to the `spider` add-on and, unusually for this project, **left no deprecated shim behind** - there is no spider implementation package in core at all. Core did keep the seams the add-on plugs into: the spider request initiator, the spider history types, and the session-level exclude-from-spider list with its desktop panels. So "nothing survives in core" is too strong; what survives is sockets, not an engine.
- Does turning `parseRobotsTxt` off stop ZAP's spider requesting `/robots.txt`?Yes. The root-file seed is added only inside that check, so with the option off the spider generates no request for the file. `parseSitemapXml` gates `/sitemap.xml` identically. The file can still reach history by another route, such as a link on the site or a request you made yourself through the proxy.
- What does ZAP's robots parser do with `Disallow: /admin/*`?It trims the trailing `*`, trims whitespace, and queues `/admin/` resolved against the host root. Only a trailing asterisk is handled: a wildcard in the middle of a pattern is carried into the URL literally. A bare `Disallow:` with no path yields an empty string and is skipped.
- Is `/robots.txt` subject to the spider's maximum parse size?No. The parse filter returns early for `robots.txt`, `sitemap.xml`, SVN and Git metadata and `.DS_Store`, ahead of both the response-size check and the text-content-type check. A robots file served with a binary content type, or one far past the cap, is still parsed in full.
A robots.txt is a staff-only sign at the end of a corridor. The traditional spider does not read it as an instruction; it reads it as a directory of corridors worth walking down.
saying these in an interview costs you the question
- parseRobotsTxt makes the spider respect robots.txt
- ZAP only reads robots.txt if something links to it
- Only Disallow lines are mined; Allow lines are skipped
- The spider honours the User-agent group aimed at it
- A Sitemap line in robots.txt tells the spider where the sitemap is
- A huge or binary robots.txt is dropped by the parse-size limit