In OWASP ZAP, where can a URL be excluded, and which part of a run does each place affect?
answer
- more than one place says no
- one belongs to the context, two do not
- one of them needs an add-on installed
- the same key name in two different blocks
- exclude-from-scan versus exclude-from-spider
basics
~20 sThree places: a context's exclude list, which subtracts from scope; the session's per-subsystem lists for the proxy, the active scan, the crawlers and websockets; and the network add-on's global exclusions, read by the proxy, the crawlers and the active scanner.
solid answer
~40 sA **context** exclude regex subtracts from scope, so it reaches everything derived from scope — a context-bound crawl, the passive engine when restricted to in-scope traffic, and the active scanner's scope check. The **session** holds separate lists keyed per subsystem: exclude-from-proxy read by the `network` add-on's local server, exclude-from-scan merged into core's active-scan exclude list, exclude-from-spider merged into every crawler's, and one for websockets. The `network` add-on's **global exclusions** are a third list, read by the proxy handler, the crawler fetch filter and the active scanner; without that add-on the session returns an empty list. Only the context layer and the session's scan list are expressible in a plan — the latter through `activeScan-config`, whose key is confusingly also called `excludePaths`.
code
yaml · 8 linesenv:
contexts:
- name: app
urls: [ "https://example.com/" ]
excludePaths: [ ".*/logout.*" ]
jobs:
- type: activeScan-config
excludePaths: [ ".*/admin/.*" ]go deeper
Recall that excluding a URL is not a single setting: a context has its own exclude list, and the session and the network add-on hold separate ones.
Name which component reads which list, and say why a context exclude reaches more of a run than a session list does.
Diagnose the mismatch cases: an exclusion that a plan cannot express, a job that silently replaces a list, and an add-on whose absence turns a whole layer into an empty list with no warning.
Decide which layer your teams are allowed to use, and accept the cost of the choice — the layer that travels in the plan is reviewable, and the layer that catches third-party noise is the one your plans cannot carry.
## Three places, not one "Exclude this URL" is not one setting in OWASP ZAP. It is three, they are stored in different places, they are read by different components, and only one of them is part of a context. ### 1. A context's exclude list Regexes held by the `Context` object itself, written as `excludePaths` in a plan or added from the site tree. They subtract from **scope**, and therefore from everything that derives from scope: a crawl bound to that context, the passive engine when it is restricted to in-scope traffic, and the active scanner's scope check. They are matched with the query string already removed. ### 2. The session's per-subsystem URL lists Core's session holds several lists of regexes keyed by a type constant in its URL table. Each is read by exactly one kind of consumer: - **exclude from proxy** — read by the `network` add-on's local server handler as traffic passes. - **exclude from scan** — merged into the exclude list that core's active-scan engine is given. - **exclude from spider** — merged into the exclude list handed to the traditional crawler, and to both browser-driven crawlers. - **exclude from websocket** — read by the `websocket` add-on. These are *not* context settings. They belong to the session, they apply whatever context a job names, and they are ordinarily set from the desktop window or over the control API. ### 3. Global exclusions, in the `network` add-on A separate list again, owned by the add-on's `GlobalExclusionsOptions`. Core has no implementation of its own: `Session.getGlobalExcludeURLRegexs()` calls a supplier, and **returns an empty list when no supplier has been installed**, which is the case in any build without that add-on. The list is also not saved in the session. It is read in three places: the proxy handler, the crawler's fetch filter, and the active scanner's exclude list. The shipped defaults are worth knowing about because they are easy to misread. The add-on ships a prepared set — media and document extensions, stylesheets and scripts, and a handful of well-known third-party hosts — and **every one of them is shipped disabled except the Windows-Update entry**, which ships enabled. Only enabled entries are returned to callers. ## Which one each consumer actually reads | consumer | context excludes | session list | global exclusions | |---|---|---|---| | local proxy handler (`network` add-on) | no | exclude-from-proxy | yes | | traditional and browser crawlers (`spider`, `spiderAjax`, `client`) | yes, via the context it is bound to | exclude-from-spider | yes | | active-scan engine (core `Scanner`) | yes, via the scope check | exclude-from-scan | yes | | passive engine (`pscan` add-on) | yes, via scope, when restricted to it | no | no | Read the table by column and the shape becomes clear: the **context** layer is the one every scanning component consults once a context is bound to the job, and the **global** layer is the only one that reaches the proxy as well as the scanners. ## What a plan can and cannot express This is where a pipeline reader is most often caught out. - A context's `excludePaths` is expressible in a plan, in the `env.contexts` block. - Of the session lists, **only exclude-from-scan is reachable from a plan**, through the `activeScan-config` job — whose key for it is also called `excludePaths`. Two keys, the same name, the same help text in the shipped template, two entirely different layers. - **No job sets global exclusions.** They are an add-on option: a configuration key, or the desktop window. A plan cannot turn one on. And one sharp edge on the job that can write a session list: it writes it **unconditionally**. Its data object holds an empty list when the key is absent, and running the job replaces whatever the session held. Against a fresh one-shot run that is invisible, because the list was empty anyway; against a long-lived daemon, or a saved session loaded with `-session`, adding an `activeScan-config` job with no `excludePaths` key **clears the exclusions that were already there**. ## How to choose 1. If the rule should govern what the run treats as its target, put it in the **context**. It is the only layer that travels inside a plan and the only one every scanning component consults. 2. If the rule must keep a specific endpoint away from the active scanner only, the session exclude-from-scan list is the right layer, and `activeScan-config` is how a plan reaches it. 3. If the rule is about noise on the wire — third-party hosts, static assets — that is the **global** layer, and you are accepting that it needs an add-on, cannot be written in the plan, and will be silently absent in a build that does not carry that add-on.
- Why can adding an `activeScan-config` job remove exclusions you had already set?Because the job writes the session's exclude-from-scan list unconditionally from its own data, and that data is an empty list when the `excludePaths` key is absent. Running the job therefore replaces whatever was there. It is invisible in a one-shot run that started with an empty session, and it bites against a long-lived daemon or a session loaded from disk.
- What happens to global exclusions in a build without the `network` add-on?They are not merely inactive — they do not exist. Core keeps no implementation; the session asks a supplier for the list and returns an empty one when no supplier has been installed. Nothing warns you, so a plan that assumed those exclusions were trimming traffic simply scans more than it did before.
- Which shipped global exclusions are on out of the box?Almost none. The add-on ships a prepared set covering media and document extensions, stylesheets and scripts, and several third-party hosts, and they are all shipped disabled apart from the Windows-Update entry. Only enabled entries are returned, so installing the add-on does not by itself change what a scan reaches.
saying these in an interview costs you the question
- Thinks a context exclude also keeps the request out of the proxy's records
- Assumes the shipped global exclusions are active because they ship configured
- Says an automation plan can turn a global exclusion on
- Treats the session exclude-from-scan list and a context excludePaths entry as one setting
- Expects global exclusions to work in a build without the network add-on
- Believes the exclude-from-spider list also keeps the active scanner away