In ZAP's zap-baseline.py wrapper, what does -j add, and which browser crawler does it now run?
answer
- It adds a crawl, it does not replace one
- Two flags, two separate decisions
- The usage text disagrees with itself
- Newer crawler is the current default
basics
~20 sThe -j flag adds a browser-driven crawl on top of the traditional spider. It now runs the client spider by default, and --ajax-spider switches back to the AJAX spider. The wrappers' own usage text still claims the opposite default.
solid answer
~40 s`-j` does not replace the traditional spider — it adds a browser crawl after it, in both `zap-baseline.py` and `zap-full-scan.py`. Which browser crawler runs is a second decision: the scripts initialise their internal flag to the **client** spider and only switch to the AJAX spider when you pass `--ajax-spider`; `--client-spider` states the default explicitly. Read the usage text and you will be told the opposite — the `-j` line still ends "(default: Ajax spider)" while the `--client-spider` line two rows below is marked "(default)". The code settles it, and the stale line is in both scripts. In CI, pass the crawler flag explicitly so a pipeline reader never has to work out which one ran.
code
bash · 5 lines# browser crawl on top of the traditional spider, crawler named explicitly
zap-baseline.py -t https://example.com -j --client-spider
# same run, but with the older crawler - and its different defaults
zap-baseline.py -t https://example.com -j --ajax-spidergo deeper
Remember that -j adds a browser crawl on top of the traditional spider, and that a second flag decides which browser crawler runs. Pass that second flag rather than relying on a default.
Be able to say why the usage text is wrong and how you would check: the generated plan names the job, and the code initialises the flag to the newer crawler.
Point out that the choice silently changes scope handling, logout avoidance, depth and browser count, and that a wrapper cannot express most job parameters — which is the argument for a hand-written plan.
Decide where the team draws the line between a wrapper invocation everyone can read and a checked-in plan nobody can run by accident, and who owns the defaults when a tool changes one under you.
## What `-j` actually turns on The packaged scan scripts always run the traditional `spider` add-on. `-j` adds a **second, browser-driven crawl** afterwards; the scripts' own usage text describes it as "use the modern spider in addition to the traditional one". Nothing is removed by passing it. Both `zap-baseline.py` and `zap-full-scan.py` accept the flag; the third wrapper does not spider at all, so the flag has nothing to attach to there. ## Which browser crawler `-j` picks The wrappers keep two separate pieces of state: - one that says *whether* to do a browser crawl at all — set by `-j`; - one that says *which* browser crawler — initialised to the client spider, set to the AJAX spider by `--ajax-spider`, and set back by `--client-spider`. So `-j` alone gives you the **client** spider. `-j --ajax-spider` gives you the AJAX spider. The wiring is the same on both execution paths the baseline script has: when it writes an automation plan it emits either a `spiderAjax` or a `spiderClient` job, and on the older path it calls either the AJAX spider or the client spider through the control API. ## The help text contradicts itself, and the code wins Inside one usage block you can read both of these: - `-j` … "use the modern spider in addition to the traditional one **(default: Ajax spider)**" - `--client-spider` … "use the Client spider when -j is specified **(default)**" Only one of those can be true, and it is the second. The stale parenthesis on the `-j` line survives in **both** wrappers. A related trap sits next door: the wrapper that imports a definition instead of spidering still *accepts* `-j` in its option string and has no handler for it, so passing it there is neither an error nor an effect. This is worth more than the fact itself: a shipped help string is documentation of what the program once did, and it is not evidence about what it does now. The same caution applies to the YAML job templates you copy plan fragments from — a comment in a template is a comment, not a validator. ## Why the choice is not cosmetic The generated job block sets only the target URL and a duration. Everything else is left at whichever add-on's own default applies — and the two crawlers **do not share defaults for the settings they share by name**. Flipping `--ajax-spider` therefore changes more than the engine: | setting | what `spiderAjax` defaults to | what `spiderClient` defaults to | |---|---|---| | `scopeCheck` | `Strict` | `Flexible` | | `logoutAvoidance` | off | on | | `maxCrawlDepth` | the deeper of the two | the shallower of the two | | `numberOfBrowsers` | roughly one per core | roughly half that, capped | So a pipeline that "just added `--ajax-spider` to compare" has also switched how out-of-scope resources are handled, whether the crawl avoids clicking logout controls, how deep it goes and how many browsers it starts. None of that appears in the command line you changed. ## How to run it so the next reader can tell 1. **Pass the crawler flag explicitly**, even when it is the default: `-j --client-spider` reads unambiguously six months later, and it survives a change of default. 2. **Do not read the crawler choice out of the usage text.** If you need to know what a version in front of you does, run it with `--plan-only` on the baseline script and read the plan it writes, or read the job list in the output. 3. **Pin the dials you care about** rather than inheriting them. The wrappers cannot express every job parameter; a hand-written plan can, and that is the argument for moving off the wrapper once the scan matters. 4. **Remember what `-j` costs.** It starts real browsers. On a shared runner that is the difference between a crawl that finishes inside the stage's time budget and one that is killed halfway. ## The usual misreading People assume `-j` means "AJAX spider" because the letter and the old help line both suggest it, and because that was true for a long time. It is not true now, and the failure is silent: you get a crawl, you get URLs, and the wrong crawler's defaults are quietly in force. If a scan's coverage changed between two pipeline runs and nobody touched the application, the crawler selection and its inherited defaults are the first place to look.
- Why is naming the crawler explicitly worth the extra flag?Because the default has already changed once, and because the two crawlers do not share defaults for `scopeCheck`, `logoutAvoidance`, crawl depth or browser count. An explicit `--client-spider` or `--ajax-spider` records which set of defaults a stored result was produced under, which is what you need when two runs disagree.
- How would you confirm which crawler a wrapper run actually used?On the baseline script, generate the plan without running it and read the job list — you will see a `spiderAjax` or a `spiderClient` job. Failing that, the run's own output names the crawler as it starts. Do not infer it from the usage text, which still carries a stale default.
saying these in an interview costs you the question
- Says -j runs the AJAX spider because the help says so
- Thinks -j replaces the traditional spider rather than adding to it
- Assumes swapping crawler flags changes only the engine
- Thinks -j does something on the wrapper that never spiders
- Treats a shipped help string as evidence of current behaviour