What do ZAP's packaged scan scripts do with a `-t` target URL that includes a path?
answer
- the target is a start, not a fence
- a path in -t is dropped
- it resets to the host root
- the scripts' own comment says why
basics
~20 sThey truncate it to the scheme and host and work from there. Targeting an application's subpath does not confine the crawl or the attacks to that subpath; the wrapper deliberately resets to the host root for backwards compatibility.
solid answer
~40 sEach wrapper rewrites a target that carries more than the scheme's own slashes down to everything up to and including the first slash after the host, and uses that as the crawl and active-scan target. The scripts' comments say why: the URL may include a valid path, but they *always reset to spider the host*, kept for backwards compatibility. The plan-writing path does the same and puts both the URL you passed and the host root into the plan's scope. So `-t` picks a starting point and a host, not a boundary — the exact URL is fetched once, and then the run works the host. Confining a scan to part of a site has to come from a context or an exclusion list instead.
go deeper
Remember that the target you pass is where the scan begins, not where it stops; a path in the URL does not keep the scan inside that path.
Be able to describe the rewrite, say that all three wrappers and the plan path perform it, and note that the exact URL is still fetched once as a seed.
Explain the blast-radius consequence on a shared or proxied host, and name what does confine a run instead of the target argument.
Treat per-application scoping on shared hosts as a design constraint on where scans may run, not something a target URL can be trusted to deliver.
## What `-t` actually selects Every packaged scan script takes its target with `-t`, and each of them rewrites one that carries a path. The rule is the same in all three: if the URL contains more slashes than the scheme's own, everything from the first slash after the host onward is dropped, and the result — scheme, host, trailing slash — is what the crawl and the active scan are pointed at. `https://example.com/app/reports` becomes `https://example.com/`. The scripts are candid about it. Their comments read *the url can include a valid path, but always reset to spider the host*, with backwards compatibility given as the reason. This is not a bug and it is not conditional on anything you pass. ## Where it happens in each script | script | where the rewrite happens | what ends up aimed at the host root | |---|---|---| | `zap-baseline.py` | once, after fetching the target, before the crawl | the crawl | | `zap-full-scan.py` | twice — before the crawl, and again before the active scan | both steps | | `zap-api-scan.py` | before the active scan | the active scan (discovery came from the definition) | | the plan-writing path | while assembling the plan | the crawl job, with both URLs left in scope | Two details are worth pulling out of that table. The full-scan wrapper truncates a second time immediately before the active scan, so even if something had narrowed the target in between, the attack runs against the host. And the plan-writing path keeps *both* the URL you passed and the host root in the plan's scope rather than replacing one with the other, so the widening is recorded in the artefact rather than hidden in the wrapper. ## The URL you passed is still requested One detail stops this from being a pure widening. The baseline and full-scan wrappers fetch the exact URL you gave through the ZAP process *before* the rewrite, so that page enters the site tree and seeds the crawl. The path is honoured as a starting point and discarded as a boundary. That is the sentence to carry: **a start, not a fence.** ## Why it matters for a scheduled run - **Scoping by target does not work.** A nightly job aimed at one application's subpath on a shared host will crawl, and on the full scan attack, the whole host. - **Multi-tenant hosts are the sharp case.** Several applications behind one hostname means one job's target selects all of them. - **A reverse proxy makes it worse**, since paths that route to entirely different services are reachable from the same host root. - **The widening is quiet.** Nothing in the default output announces the rewrite — the API scan records it at debug level only, and the other two do not log it at all — so you see the crawl results and infer it. ## What actually confines a run The confinement has to come from somewhere that is genuinely a boundary rather than a starting point — a context with include and exclude rules loaded with `-n`, or the exclusion lists the program maintains. Both of those are subjects of their own, and the point here is only the negative one: `-t` is not among them. Note too that on the baseline wrapper, passing `-n` is itself one of the options that sends the run down the legacy execution path, so scoping a run and choosing how it executes are not independent decisions. ## Getting the answer right under questioning The weak answer is "I set the target to the subpath, so that is what it scans". The strong answer names the rewrite, gives the reason the script itself gives, notes that the exact URL is still fetched once as a seed, and then moves to what does confine a run. An interviewer asking this is usually checking whether you have read the wrapper or only its usage text. There is a broader habit behind it. A packaged wrapper's arguments are a convenience layer over a program that has its own, richer notion of what is in and out of bounds, and a convenience layer will usually simplify in the direction of *doing more*, because that is the direction in which nothing visibly breaks. When you are relying on an argument to keep a scan away from something, check what the wrapper does with it before you rely on it — and prefer a mechanism whose whole purpose is exclusion to one whose purpose is choosing a starting point.
- Is the path you passed requested at all?Yes. The baseline and full-scan wrappers fetch the exact URL you gave through the ZAP process before rewriting anything, so that page enters the site tree and seeds the crawl. The rewrite only changes what the crawl and the active scan are aimed at afterwards.
- Does the plan-writing path behave differently?It performs the same truncation, and it puts both the URL you passed and the host root into the plan's scope while giving the crawl job the host root. The widening is in the plan too, not only in the API-driven path.
- So how do you confine a packaged scan to part of a site?With a context's include and exclude rules, loaded with `-n`, or with the program's exclusion lists — not with the target argument. On the baseline wrapper, `-n` also sends the run down the legacy execution path, so the two choices are linked.
Giving -t a path is like writing a street address on a parcel when the round is dispatched by postcode: that exact door gets its visit, and then the van works the whole postcode.
saying these in an interview costs you the question
- Pointing -t at a subpath keeps the scan inside it
- The target URL is the scan's scope boundary
- Only the baseline wrapper rewrites the target
- The generated plan scopes the run to the exact URL you passed
- A trailing path in -t is ignored and never requested