skip to content

In ZAP's active-scan engine, what actually runs in parallel: the rules, the hosts, or the nodes?

level: seniorimportance: should knowfreq 35%

answer

  1. two pools, stacked and multiplied
  2. nodes run together, rules do not
  3. one site, one host process
  4. host rules report their own completion

basics

~20 s

Hosts and nodes, not rules. Scanner runs several host processes at once, bounded by scanner.hostPerScan; inside one host process, scanner.threadPerHost worker threads run one rule's node scans together, and the next rule waits for them.

solid answer

~40 s

There are two nested pools and one deliberate serialisation. `Scanner` creates one `HostProcess` per site and runs them on a pool sized by `scanner.hostPerScan`. Inside a host process, `scanner.threadPerHost` sizes a second pool that runs the individual node scans. What is **not** parallel is the rules: for a per-node rule, `HostProcess` dispatches every node, then blocks until all those threads finish before it marks the rule complete and asks for the next one — `AbstractAppPlugin`'s own Javadoc says as much. Host-wide rules are the exception; the engine dispatches one and moves straight on, which is precisely why `AbstractHostPlugin` overrides `notifyPluginCompleted` to report its own completion while `AbstractAppPlugin`'s override is deliberately empty.

code

bash · 6 lines
bash
# the two nested pools, set as config keys on the command line
zap.sh -config scanner.hostPerScan=<sites-at-once> \
       -config scanner.threadPerHost=<threads-per-site>

# peak in-flight attack requests is roughly the product of the two;
# adding rules to the policy lengthens the scan, it does not widen it.

go deeper

for a junior

Know that a scan runs many requests at once, and that the parallelism is across pages within a single rule rather than across rules. One site gets one host process with its own pool of workers.

for a middle

Explain the two nested pools and the serialisation between rules, and say why the two base classes differ in who signals completion. Both facts come out of the same scheduling loop.

for a senior

Tune the per-host threads when a shared environment is suffering, and read a slow tail as the last rule working through its node list rather than as a hang. Cutting rules shortens a scan without widening it.

for a principal

The engine bounds itself per scan and knows nothing about the other pipelines pointed at the same environment. Deciding the organisation-wide ceiling, and where it is enforced, is a decision no ZAP setting makes for you.

## Two pools, stacked Concurrency in ZAP's active scan is configured in two places and composes by multiplication. | level | class | sized by | what a thread carries | |---|---|---|---| | across sites | `Scanner` | `scanner.hostPerScan` | one whole `HostProcess` — every rule against one site | | inside a site | `HostProcess` | `scanner.threadPerHost` | one node scan: one rule against one recorded message | `Scanner` derives the site key from the start nodes — scheme, host and port together — so `http://` and `https://` on the same name are two host processes, not one. Each process clones the scan policy's rule set, builds its own worker pool and opens its own HTTP sender. The peak number of in-flight attack requests against a whole scan is therefore roughly the product of the two settings, not either one alone. That product is the number to reason about before pointing a pipeline at a shared environment. ## What is deliberately *not* parallel Rules do not run alongside each other. For a rule that fires per node, `HostProcess`: 1. dispatches that rule against every message id in the in-scope list, onto the worker pool; 2. blocks until every one of those worker threads has finished; 3. marks the rule completed; 4. only then asks its plugin factory for the next rule. `AbstractAppPlugin`'s Javadoc states the contract directly: multiple threads will be executed, but the rule must complete before another can start. So the concurrency you get is **within one rule, across nodes** — never across rules. Adding more rules to a policy lengthens a scan roughly linearly; it does not raise the peak load on the target. ## The one exception, and why the base classes differ Host-wide rules break the pattern. `HostProcess` dispatches an `AbstractHostPlugin` against its single message and returns immediately, without waiting. The rule may therefore still be running when the next rule is dispatched. That is exactly why the two base classes differ in a second way beyond fire count: - `AbstractHostPlugin.notifyPluginCompleted` calls `parent.pluginCompleted(this)` — the rule tells the engine it is done, because nobody is waiting for it; - `AbstractAppPlugin.notifyPluginCompleted` is an empty method with a comment saying the parent will wait for it — the engine already knows. Both are invoked from the same place: `AbstractPlugin.run()` calls `notifyPluginCompleted` in a `finally` block, so the signal fires whether the rule succeeded or threw. **The base class decides both how often a rule runs and who declares it finished**, and those two facts are the whole difference between the classes. ## Ordering is not free either The engine's plugin factory hands out rules one at a time and will hand out none while a rule's declared dependencies are still incomplete; when that happens the scheduling loop sleeps briefly and asks again. So a rule can be pending not because the pool is busy but because something it depends on has not finished. Two timing caps sit alongside this — a per-rule limit and a whole-scan limit — and **both default to unlimited**. If a rule in your run is skipped with a max-rule-time reason, someone configured that; it is not a default. ## What this means when you are operating it 1. **To lower the load on one target, lower `scanner.threadPerHost`.** That is the setting that governs concurrent requests to a single site. Raising `scanner.hostPerScan` only helps when the scan genuinely covers several sites. 2. **A one-site scan does not get faster by enabling fewer rules — it gets shorter.** Since rules are serialised, cutting the rule list cuts total time without changing peak pressure. 3. **Expect a long tail.** The last rule in the list runs alone against every node, at full pool width, with nothing else in flight. A scan that looks stalled near the end is usually a slow rule finishing its node list. 4. **Budget by the product.** Several pipelines scanning the same shared environment multiply again, across processes ZAP knows nothing about. The engine bounds itself per scan, never per environment; that ceiling is yours to impose. 5. **Attribute before you tune.** Turning on the request header that names the sending rule makes the target's own logs show which rule generated the burst, which turns a concurrency argument into a measurement.

  • Why does AbstractHostPlugin report its own completion when AbstractAppPlugin does not?
    Because the engine waits for one and not the other. It blocks on a per-node rule's worker threads, so it already knows when that rule is finished; it dispatches a host rule and moves straight on, so only the rule itself can say. Both signals fire from a `finally` block in `AbstractPlugin.run()`.
  • Does adding more rules to a policy increase the load on the target?
    It increases total requests and total time, not peak concurrency. Per-node rules are serialised — the engine finishes one across every node before starting the next — so the width of the traffic is set by the per-host thread pool, not by how many rules are enabled.
  • Why might a rule sit pending while the worker pool is idle?
    The engine's plugin factory will not hand out a rule whose declared dependencies have not completed. When nothing is ready it sleeps briefly and asks again. A pending rule therefore means an unmet dependency at least as often as it means a busy pool.

saying these in an interview costs you the question

  • Says every enabled rule runs concurrently against the target
  • Thinks enabling more rules raises the peak request rate
  • Assumes one thread setting bounds the whole scan
  • Treats the two host processes for http and https as one
  • Believes the engine waits for host-wide rules to finish