In ZAP's browser-driven crawler jobs, what does scopeCheck Strict do that Flexible does not?
answer
- Two settings, one proxy decision
- One answers for the resource, one forwards it
- Neither of them crawls a third party
- A starved crawl looks like a small site
basics
~20 sStrict blocks out-of-scope resources at the crawler's own proxy: ZAP answers the browser with a canned out-of-scope response instead of the real asset. Flexible lets them load and records them as third-party, but still refuses to crawl them.
solid answer
~40 sBoth browser crawlers route their browser through a local server they start themselves, and `scopeCheck` decides what that server does with a request for something out of scope. Under **Strict** the handler overrides the response — the browser gets a canned out-of-scope body, not the resource — so a third-party script, stylesheet or font simply does not arrive. Under **Flexible** the request passes through, the response is recorded as a third-party resource, and the crawler still declines to *crawl* it: it loads, it is not explored. The distinction matters because a modern application's own code is frequently served from somewhere else. Under Strict the page can fail to initialise at all, and a crawl that found one page looks exactly like a crawl of a site with one page.
go deeper
Know that this setting decides what happens to a request for something outside the crawl's scope, and that one of its two values stops the browser from ever receiving the real resource.
Explain the mechanism: the crawler routes the browser through its own local server, and Strict makes that server answer with a canned out-of-scope response instead of forwarding.
Recognise the failure signature — clean exit, tiny URL count, no error — and know the fix is usually to widen the context to your own asset host rather than to loosen the check globally.
The real decision is what your pipelines are permitted to fetch. Flexible lets runs reach every third party a page references, which is a policy question about your own outbound traffic, not just a crawl setting.
## Where the check happens Neither browser crawler watches the network from outside. Each launches a browser and points it at a local HTTP server it started for that browser, and every request the page makes goes through a handler on that server. `scopeCheck` is that handler's policy for a request whose URL is not in scope. Scope here is resolved the usual way — a subtree prefix if one was given, otherwise the crawl's context, otherwise the target host — followed by the crawl's exclusion regexes. What `scopeCheck` decides is not *what counts as out of scope* but *what to do about it*. ## The two behaviours | | `Strict` | `Flexible` | |---|---|---| | request for an out-of-scope resource | blocked at the crawl's own proxy | forwarded to the real host | | what the browser receives | a canned out-of-scope response from ZAP | the actual response | | how it is recorded | as a temporary, out-of-scope message | as a third-party resource | | is it crawled? | no | no | The last row is the one people miss. **Flexible is not "crawl everything".** It widens what the browser may *load*; it does not widen what the crawler will *explore*. A third-party host is still never treated as somewhere to go looking for pages. ZAP's own option help puts it the same way: Flexible "allows out-of-scope resources (e.g. third-party scripts and stylesheets) to pass through the proxy so the browser can load them, but does not crawl them". The implementations differ slightly in shape. The AJAX spider leaves its bundled crawler free to consider any URL under Strict and relies on the proxy handler to refuse the traffic; under Flexible it does the opposite, teaching the crawler to skip out-of-scope states while letting the traffic through. The client spider expresses the same policy by switching its handler between allowing everything and filtering. The observable behaviour is the one in the table. ## Why this decides whether a crawl works Take the case this leaf exists for: a single-page app whose navigation never appears to a link parser. The document served at the target is a shell. The router, the views and most of the interactivity arrive in a JavaScript bundle, and in a real deployment that bundle is very often on a content-delivery host, with the data behind an API on a third hostname. Under Strict with only the application's own host in scope: 1. the browser asks for the bundle; 2. the crawl's proxy handler decides the bundle's host is out of scope; 3. ZAP answers with its out-of-scope response instead of the script; 4. the application never initialises, so there is nothing to click; 5. the crawl finishes quickly, reports very few URLs, and **raises no error**. That last point is what makes it dangerous in a pipeline. A crawl that was strangled by its own scope policy produces the same shape of result as a crawl of a genuinely small site: a clean run, a small number of URLs, no failure. ## What to do about it - **Set `scopeCheck` explicitly** in the job rather than inheriting it. The two crawler jobs default it differently, so which behaviour you get otherwise depends on which crawler ran. - **Put the application's own assets in scope** when they live on another hostname. Widening the context to include your own delivery host is a scope decision about assets you own, and it keeps Strict usable. - **Watch the URL count, not just the exit status.** A browser crawl that suddenly finds an order of magnitude fewer URLs than last week is the signal; a green run is not evidence that the application was reached. - **Do not reach for Flexible as a general fix without thinking about what it means.** It permits the browser to fetch from hosts you were not authorised to touch — analytics, fonts, trackers, whatever the page references. Those are ordinary requests to third parties, made by your pipeline, and Flexible is the setting that allows them. ## The precise claim, stated carefully It is easy to over-generalise this into "ZAP blocks out-of-scope requests", and that is not a safe sentence. What is true here is narrow and specific: **the browser crawlers' own proxy handler, under Strict, overrides the response for an out-of-scope request.** That is a property of those crawlers' handlers, not a general property of scope or exclusion elsewhere in the program. Keep the claim attached to the mechanism that makes it true.
- Does Flexible mean the crawler explores third-party hosts?No. Flexible only lets the browser *load* out-of-scope resources so the page renders; the crawler records them as third-party and never treats them as somewhere to explore. Widening what may load and widening what is crawled are separate things, and `scopeCheck` only moves the first.
- How do you keep Strict and still crawl an app whose bundle is on another host?Bring the asset host into scope. A context can include more than one hostname, so adding the delivery host you own lets the bundle load while everything else stays blocked. That is usually better than switching to Flexible, which permits fetches from every third party the page references.
- What does a Strict-starved crawl look like in a pipeline?Like a successful crawl of a very small site: no error, a quick finish, and a handful of URLs. Nothing in the exit status distinguishes it, so the signal you need is the URL count compared with previous runs, plus a look at whether the crawl's messages contain out-of-scope entries for your own assets.
saying these in an interview costs you the question
- Says Flexible makes the crawler explore third-party hosts
- Thinks Strict merely stops recording out-of-scope requests
- Assumes a blocked resource produces a run failure
- Reads a low URL count as proof the site is small
- Generalises this to how all ZAP exclusions behave