What does each of ZAP's three packaged scan scripts run against a target, and how do they differ?
answer
- three scripts, three amounts of traffic
- one of them never attacks
- one of them never crawls
- the API scan starts on API-Minimal
basics
~20 szap-baseline.py crawls the target and passively inspects the traffic, never attacking. zap-full-scan.py adds an active scan with all rules enabled. zap-api-scan.py never crawls: it imports an API definition and attacks only what the import reached.
solid answer
~40 sAll three scripts ship in the core repository's `docker/` directory and drive one ZAP process, but they run different work. `zap-baseline.py` runs the traditional spider and then waits for the passive scan to drain; it never starts an active scan, so it sends no attack traffic of its own. `zap-full-scan.py` runs the same crawl and then an active scan under the `Default Policy` with every active rule enabled, and waits for passive inspection on top. `zap-api-scan.py` does not crawl at all: you give it `-f openapi`, `-f soap` or `-f graphql` plus a definition, the import issues the requests that populate the site tree, and an active scan then runs under the narrower `API-Minimal` policy unless you pass `-S`, which skips it.
code
bash · 8 lines# crawl and watch only - no attack traffic of its own
zap-baseline.py -t https://example.com
# the same crawl, then an active scan with every rule enabled
zap-full-scan.py -t https://example.com
# no crawl at all - the definition supplies the requests
zap-api-scan.py -t https://example.com/openapi.json -f openapigo deeper
Know which script attacks and which does not: the baseline crawls and watches, the full scan attacks, and the API scan attacks whatever its definition described.
Be able to name the sequence each script runs and the policy it starts from, including that the API scan begins on a narrower policy than the full scan and that it never crawls.
Explain how the choice of script bounds the traffic an unattended run puts on a deployed environment, and name the inputs that widen that bound without anyone editing the job.
Own the rule for which environments each script may be pointed at, and make the narrower script the default so that widening it is a visible decision rather than a copied flag.
## Three scripts, one program ZAP's **packaged scan scripts** are three Python programs that ship inside the project's published container images and live in the core repository's `docker/` directory: `zap-baseline.py`, `zap-full-scan.py` and `zap-api-scan.py`. Each is a wrapper. It starts one ZAP process, drives it through a fixed sequence of work, prints a per-rule summary and returns an outcome the caller can branch on. None of them is a separate engine: the crawl, the passive inspection and the active attacks are all the running program's, and the script only decides which of them happen, in what order, and against what. The difference that matters when you schedule one is **what traffic each puts on the wire**, because that is what decides whether you are allowed to point it at a given environment at all. ## What each one actually runs | script | discovery | active scan | policy it starts from | |---|---|---|---| | `zap-baseline.py` | traditional spider, optionally a browser crawler with `-j` | **none** | — | | `zap-full-scan.py` | traditional spider, optionally a browser crawler with `-j` | yes, every active rule enabled | `Default Policy` | | `zap-api-scan.py` | **no crawl at all** — an import of the definition you pass | yes, unless `-S` | `API-Minimal` | ## The baseline: crawl, then watch `zap-baseline.py` fetches the target through the ZAP process, runs the traditional spider, waits for the passive queue to drain, and summarises. Its sequence contains no call to the active scanner at all, which is the property the name is really about: - it generates traffic — a crawl is requests — but the requests are ordinary ones; - passive rules inspect whatever the process handled, so findings come from observation, not from crafted input; - the spider is capped by `-m` at a short default duration — but the wait for passive inspection that follows has no limit of its own unless you set `-T`, so the crawl is bounded and the run is not; - `-j` adds a browser-driven crawl on top, which increases reach without changing the passive-only character of the run. ## The full scan: the same crawl, plus attacks `zap-full-scan.py` repeats that crawl — with `-m` defaulting to **no** limit rather than a short one — and then starts an active scan against the `Default Policy`. Two consequences follow that people routinely get wrong: 1. **Passive inspection does not stop.** Passive rules see every request and response the process handles, so the crawl *and* the attack traffic feed them. The wrapper waits for that queue before it summarises and reports passive rules beside active ones. "Passive is the baseline's job" is false. 2. **The policy is the wide one.** The wrapper names `Default Policy` and, when you hand it a rule configuration file, explicitly enables every scanner in that policy before turning individual ones off. A full scan is opt-out, not opt-in. ## The API scan: the definition is the discovery `zap-api-scan.py` has no spider step anywhere in it. `-f` is mandatory and accepts only `openapi`, `soap` or `graphql`; `-t` is the definition — a URL or a file — rather than a site root. The import is what issues requests, so the site tree it builds is exactly what the document described, and anything the document omits is not scanned. Then: - the active scan runs under `API-Minimal`, a shipped policy narrower than the full rule set, which the wrapper copies into the container before the run; - `-S` skips that active scan entirely, leaving an import plus passive inspection; - supplying a rule configuration file switches the run onto `Default Policy` with every scanner enabled — a file people reach for to quiet output also widens the attack; - `-O` overrides the host the definition names, and `--schema` supplies a GraphQL schema separately. ## What the three have in common Everything outside the work itself is shared, which is part of why they get confused. All three take `-t` for the target and rewrite a path-bearing one down to the host root before scanning; all three accept the same report options and the same rule configuration file; all three start their ZAP process with the same fixed start options; and all three print a per-rule summary before returning an outcome. **The sequence in the middle is the only part that differs — and it is the part that decides what the run does to somebody's environment.** ## How to say it in an interview Name the traffic, not the label. "The baseline crawls and watches; the full scan crawls and then attacks with everything on; the API scan skips discovery, replays the definition you gave it, and attacks the narrower policy unless you widen it." That sentence answers the question that is actually behind it — which of these may I schedule against which environment — and it is the frame the rest of this subject hangs on.
- Does `zap-full-scan.py` still produce passive findings, or only active ones?Both. Passive rules inspect whatever traffic the process handles, so the crawl and the attacks both feed them. The wrapper waits for the passive queue to drain before summarising and lists passive rules beside active ones.
- What happens if you hand `zap-api-scan.py` a rule configuration file?It stops using `API-Minimal`. Any parsed rule line switches the run onto the `Default Policy` with every active scanner enabled, and the file's exclusions are then applied by turning individual rules' thresholds off. A file reached for to quiet output widens the attack.
- Can `zap-baseline.py` be made to send attack traffic through its own options?Its sequence has no active-scan step — it fetches, crawls, waits for passive inspection and summarises. `-z` does append arbitrary options to the ZAP process it starts, so the boundary is the script's steps rather than a refusal built into the program.
saying these in an interview costs you the question
- The baseline scan is just a faster full scan
- All three packaged scripts spider the target first
- Passive scanning only happens in the baseline script
- The API scan runs the same rule set as the full scan
- The full scan is the baseline plus a longer crawl