ZAP's traditional spider never executes JavaScript - so what does it extract from a `.js` response?
answer
- fetch and parse, never execute
- non-HTML text falls to one parser
- that parser requires a scheme
- relative paths in code are invisible
basics
~20 sAbsolute URLs only. A script body is text but not HTML, so the spider's fall-back text parser scans it with a regular expression for http and https strings. A relative path assembled in code is invisible to it.
solid answer
~40 sThe traditional `spider` add-on fetches and parses; it runs nothing. A `.js` response is text but not HTML, so no dedicated parser claims it and `SpiderTextParser` picks it up as the fall-back. That parser does one thing: it runs a regular expression over the body for absolute `http://` or `https://` URLs and queues every hit. A hard-coded absolute endpoint in a bundle is therefore found, while `fetch('/api/v1/orders')` is not, because the pattern requires a scheme and there is no base-URL resolution on that path. The same limit applies to CSS and any other non-HTML text. Anything that exists only after a script has run needs a crawler that drives a browser.
go deeper
The one-line version: this crawler downloads and reads text, it never runs a page. If a link only exists after JavaScript has run, this crawler will not find it.
Be able to say which parser handles what - markup and comments by the HTML parser, forms by the form parser, and any other text response by a fall-back that matches only absolute URLs.
Expect to justify a coverage claim rather than assert one: show how you compared the crawl's URL list against the application's real routes, and name the parts you know this crawler cannot reach.
The call to own is how much crawl coverage a release gate may assume, and what evidence - not which tool - is allowed to stand behind that assumption.
## Fetch and parse, never execute ZAP's traditional crawler - the **`spider` add-on** - has one loop: request a URL, hand the response to a chain of parsers, queue whatever URLs those parsers report, repeat until a limit stops it. There is no JavaScript engine and no DOM anywhere in that loop. Every claim about what this crawler can and cannot see follows from which parser claims a response and what that parser matches. ## Which parser claims a response A parser is asked whether it can handle a response; the first that says yes consumes it, and a fall-back catches whatever is left. | response | parser | what comes out | |---|---|---| | HTML | the HTML parser | `href`, `src`, `action`, `cite`, `data`, `srcset` and friends, relative URLs resolved against the base | | HTML | the form parser | a built submission per form action | | the exact path `/robots.txt` | the robots parser | every `Disallow` and `Allow` path | | a sitemap XML file | the sitemap parser | the listed locations | | a redirect | the redirect parser | the target of the redirect | | **any other text response** | **the fall-back text parser** | **absolute `http`/`https` URL strings only** | | a non-text response | none | nothing is parsed | A JavaScript bundle, a CSS file, a JSON document and a plain text file all land in that last text row. ## What the fall-back text parser actually matches The fall-back applies one regular expression to the body and queues every capture. The pattern is narrow in ways that decide your coverage: - it **requires a scheme** - `http://` or `https://`; - it stops the match at whitespace, quotes, angle brackets, brackets and braces, so a URL glued into an expression often truncates; - it does **no base-URL resolution** on that path, because a relative string never matches in the first place; - it is case-insensitive and does not care that the surrounding file is code. So the practical rule is: **a URL your bundle spells out in full is found; a URL your bundle builds is not.** `fetch("/api/v1/orders")` yields nothing. `fetch(BASE + "/orders")` yields nothing. A hard-coded `https://…` analytics or API endpoint becomes a *candidate* URL, which the spider's fetch filter then accepts or rejects on scope - and the scope it checks against is derived from the hosts you seeded. ## What this means for coverage 1. **Server-rendered navigation is covered well.** Anchors, frames, images and even links in HTML comments - comment parsing is on by default - come out of the HTML parser with relative URLs resolved properly, and forms out of the form parser beside it. 2. **Client-rendered navigation is not covered at all.** A route a client-side router registers at runtime, a menu injected after load, an endpoint reached only from an event handler: none of it exists in the bytes the spider parsed. 3. **Partial coverage is the dangerous case.** The crawl returns a healthy-looking number of URLs from the server-rendered shell, and the application's real surface is elsewhere. Nothing in the output announces the gap. ## Two limits that compound it - **Size.** A response larger than the spider's maximum parse size is fetched but not parsed. Large bundles are exactly the files most likely to exceed it. - **Content type.** A response the program does not consider text is not parsed either. Both checks have a documented exception list - `robots.txt`, `sitemap.xml`, SVN and Git metadata and `.DS_Store` are examined ahead of both - so the exemptions run in the direction of *more* crawling, not less. ## Where the boundary is The honest summary for an interview is three sentences. The traditional spider reads what the server sent. It extracts links from markup properly and from any other text only when they are written out as absolute URLs. Anything that requires the page to run belongs to a crawler that drives a real browser, which is a different component with different tradeoffs. The follow-on discipline matters more than the fact: **do not claim coverage you have not measured.** Compare the URL list a crawl produced against the routes your application actually defines, or against the traffic a functional test suite generates through the proxy. If a route is missing from the crawl, everything downstream of the crawl is blind to it too, and no severity count will tell you so.
- Does ZAP's traditional spider parse links out of HTML comments?Yes, when `parseComments` is on, which is its default. Valid markup inside a comment is treated as an ordinary source of links, so a block someone commented out before release is still crawled. Turning the option off leaves comment contents unread.
- Why is an absolute URL in a bundled script crawled while a relative one is not?The fall-back text parser's pattern requires a scheme: it matches `http://` or `https://` followed by a run of characters that are not whitespace, quotes or brackets. There is no base-URL resolution step on that path, so a bare `/api/orders` string never becomes a candidate URL at all.
saying these in an interview costs you the question
- The spider ignores JavaScript files completely
- It finds any URL string in a script, relative ones included
- It renders the page first and then reads the resulting DOM
- Links inside HTML comments are never crawled
- Every response is parsed, whatever its size or content type