How would you choose which packaged ZAP scan script a nightly pipeline runs against a deployed environment?
answer
- start from what you are permitted to send
- coverage is the second question
- one script never attacks
- a config file can widen the API scan
basics
~20 sChoose by the traffic you are authorised to put on that environment, not by coverage. The baseline only crawls and inspects; the full scan attacks everything it found; the API scan attacks only what its definition described.
solid answer
~40 sStart from what you are permitted to do to the environment, because the script is the only thing bounding it — none of the three sets any offence-limiting mode on the process it starts, so an unattended run is unrestricted by default. A baseline run is the one you can point at a shared or near-production environment: it crawls and inspects but starts no active scan. A full scan enables every active rule and by default puts no time limit on its crawl, so it belongs against an environment you own and can restore. An API scan's reach is exactly the definition you feed it, which makes it the most predictable of the three — until someone hands it a rule configuration file, which promotes it onto the full rule set.
go deeper
Know that picking a script is a decision about what you are sending at someone's environment, and that the baseline is the one that sends no attacks.
Be able to name, for each script, the default that bounds its run and the flag or input that widens it.
Explain why an unattended run is the case that most needs the narrow script, and how you would detect a job that has quietly widened.
Own the per-environment rule for which script may run and who can change it, and design so that widening is a reviewed decision rather than an added flag.
## The choice is about traffic, not coverage The instinct is to pick the script with the most findings. That is the wrong first question for a scheduled run, because the scripts differ most in **what they do to the system under test**, and that difference is what your authorisation to run them is granted against. Coverage is the second question and it only arises once the first is settled. ## What each script commits you to | script | traffic it adds | default bound | what widens it | |---|---|---|---| | `zap-baseline.py` | a crawl, and passive inspection of it | the crawl is capped by `-m` | `-j` adds a browser crawl | | `zap-full-scan.py` | a crawl, then every active rule | **no crawl limit by default** | nothing needs to — it is already wide | | `zap-api-scan.py` | the definition's requests, then a narrow active scan | `API-Minimal`, or `-S` for none | a rule configuration file promotes it to the full rule set | ## The guardrail that is not there The program does have a mode that can restrain what a scan will do, and **no packaged script sets it**. Nothing in the three restrains what the scan may do beyond the sequence the script runs, the target you named, and any scope you configured yourself. Two things follow: 1. **The script's own sequence is the guardrail.** Choosing the baseline is not a preference about thoroughness; it is the enforcement mechanism for "this environment may be crawled but not attacked", because that script has no active-scan step to switch on by accident. 2. **Authorisation is an operating decision, not a tool setting.** Nothing in the wrapper checks that the target is yours. That check exists in your change process or it does not exist. ## The target argument is not a fence A related surprise belongs in the same decision: all three wrappers rewrite a target that carries a path down to the host root before pointing the crawl or the active scan at it. Scheduling a full scan against a subpath of a shared host does not confine it to that subpath. If the environment is shared, that is by itself a reason to choose the narrower script. ## Things that widen the choice after you made it - **a rule configuration file on the API scan** — added to silence noisy findings, it switches the run from the narrow policy to every active rule enabled; - **`-j` on the baseline or the full scan** — more reach, and on the full scan more attack surface; - **an unbounded crawl** — the full scan has no default time limit, so a nightly slot it used to fit in can stop being enough as an application grows; - **moving the step** — the same command behaves differently run inside the image and run from the runner's host, so "we did not change anything" is not always true. ## How to make the decision durable 1. Write down, per environment, which script may run against it and who may change that. 2. Make the narrow script the default for anything shared, and require the wide one to name the environment it owns. 3. Review the flags, not just the script name — the widening inputs above are all flags. 4. Put the scan somewhere it can be contained, since the process it starts exposes an unauthenticated control interface for the duration. 5. Re-ask the question when the application changes shape: a service that grew an API is a candidate for the definition-driven script, which is the only one whose reach you can read off a document. ## Where the narrow choice costs you Be honest about the trade rather than presenting the baseline as free. A passive-only run reports what is visible in ordinary traffic — missing headers, exposed information, cookie handling — and cannot find the classes that only appear when input is deliberately malformed. Scheduling it nightly against a shared environment and the wide one less often against an environment you own is usually the right shape, but it is a shape, not a way of getting both. Say which findings you are choosing not to look for, and where in the plan you do look for them. ## The answer an interviewer is listening for Not "full scan, because it finds more". The answer that lands is: *what am I allowed to send at this environment, which script's sequence enforces that, and what could change it without anyone editing the job?* That ordering is the judgment; the per-script mechanics are what you use to justify it.
- A team wants the full scan nightly against a shared staging environment. What do you ask?Who else depends on that environment overnight, whether it can be restored, and whether the crawl has a time limit at all — the full scan's has none by default. If the answers are 'others', 'not easily' and 'no', the baseline is the nightly job and the full scan gets an environment of its own.
- What makes the API scan's reach more predictable than the other two?It has no crawl. The requests come from the definition you supply, so its coverage is something you can read off a document and review, rather than an emergent property of a crawler. That predictability ends if a rule configuration file promotes it to the full rule set.
- Does running unattended in CI make a scan safer?No — it removes the person who would have stopped it. Nothing in the packaged scripts limits what the run may do beyond the target and any scope you configured, so an unattended run is the case that most needs the narrow script and a contained network.
saying these in an interview costs you the question
- Always run the full scan; more coverage is strictly better
- A scan is safe because it runs unattended in CI
- The target URL you pass is the scan's boundary
- The API scan is narrow, so its policy cannot change
- The scripts refuse to attack a host you did not authorise