In ZAP, how does the pscan add-on's passive engine get the messages that its rules scan?
answer
- the scan trails the traffic
- a table is read, not a wire
- a cursor over ascending history ids
- PassiveScanController walks, PassiveScanTask scans
- records to scan is the backlog
basics
~20 sZAP's pscan add-on runs a PassiveScanController that walks the History table by record id rather than the live wire. For each record it queues a PassiveScanTask, which re-reads the message and runs every enabled passive rule.
solid answer
~40 s`ExtensionPassiveScan2` in the `pscan` add-on starts a `PassiveScanController` thread. It does not receive messages from the proxy — the proxy listener's only job is to interrupt the thread so it wakes early. The controller keeps a cursor over the **History table**, fetches the next `HistoryReference` by id, and if the record passes the `scanOnlyInScope` check submits a `PassiveScanTask` to a fixed worker pool. The task re-reads the message from the database, parses the response body once, and calls each enabled rule that applies to it. The gap between the last recorded id and the cursor, plus the tasks still running, is `getRecordsToScan()` — the backlog the `pscan` API exposes as `recordsToScan` and the engine publishes as the `stats.pscan.recordsToScan` high-water mark.
go deeper
Hold on to the shape: traffic is written to a history table first, and the passive engine reads that table afterwards. The scan trails the browsing rather than happening during it.
Walk the chain out loud: history record, controller cursor, per-record task, rule callbacks. Then name the backlog number, getRecordsToScan, and say it is last-id minus scanned-id plus running tasks.
Show that you treat the backlog number as ambiguous. Zero means drained or means no engine, and a run should prove the engine was alive rather than trusting the count, especially before a report job reads the alert store.
Own the standard for what a passive sweep is allowed to certify across teams. A cursor that never goes back means coverage is a property of capture order, so the honest contract is over traffic recorded during the run, not over the application.
## The pipeline, end to end Nothing hands a passive rule a message directly. The path is always the same, and it is worth learning as a chain because every surprising behaviour in a passive sweep is a property of one link: 1. Something puts traffic on the wire — the local proxy, a spider job, a definition import, a script, a replayed request. 2. That traffic is **recorded in the History table** as a row with an ascending integer id and a history *type* saying where it came from. 3. `ExtensionPassiveScan2`, in the **`pscan` add-on**, runs a `PassiveScanController` on its own thread. The controller holds a cursor over history ids and walks forward. 4. For each record it reaches, the controller fetches the `HistoryReference` and — if `scanOnlyInScope` is off, or the session says the record is in scope — submits a `PassiveScanTask` to a fixed-size worker pool. 5. The task re-reads the message from the database, parses the response body once into an HTML `Source`, and runs every enabled rule against it. The decoupling in step 3 is the part people get wrong. The add-on does register a proxy listener, but its `onHttpResponseReceive` does nothing except call `psc.responseReceived()`, and *that* method's entire body is `this.interrupt()`. The listener never passes the message along; it only pokes a sleeping thread so it looks at the table sooner. When there is nothing new, the controller sleeps and re-reads the last history id on waking. ## The cursor, and what it never goes back for The controller is a one-way walk. That single property explains several behaviours that otherwise look like bugs: - **Opening a saved session does not re-scan it.** On start the controller sets its cursor to the last history id and steps past it, with the comment *"Prevent re-scanning of existing message"*. Everything already in the table is behind the cursor and will never be offered to a rule. - **A skipped record is skipped for good.** If `scanOnlyInScope` is on and a record is out of scope, the controller simply does not submit it and the cursor moves on. Turning the setting off later does not bring that record back. - **`clearQueue` moves the cursor forward, never back.** The `pscan` API action of that name, and the button behind it, jump the cursor to the last history id and abandon the pending tasks. It empties the queue *without* scanning what was in it. - **Enabling a rule mid-run only affects what comes next.** The records already behind the cursor were scanned by whichever rules were enabled at the time. ## Measuring the backlog The number that matters to an automated run is `getRecordsToScan()`, computed as *(last history id − last scanned id) + running tasks*. It surfaces in these places: | surface | what it is | who reads it | |---|---|---| | `ExtensionPassiveScan2.getRecordsToScan()` | the live count | the `passiveScan-wait` automation job's poll loop | | the `pscan` API's `recordsToScan` view | the same count over HTTP | scripts driving ZAP as a daemon | | the `stats.pscan.recordsToScan` high-water mark | a statistic the controller republishes each pass | the desktop footer counter, and anything reading stats | One trap is worth memorising: `getRecordsToScan()` returns **zero** when passive scanning is not enabled or the controller was never started. Zero therefore means *either* "drained" *or* "there is no engine running", and nothing in the number distinguishes them. ## Inside one record's task The per-record task is where the remaining defaults bite. It re-reads the message, builds the shared HTML `Source`, and then, for every rule that is enabled and applies to this record's history type, copies the rule, binds a per-message helper to the copy, and calls the request-side and response-side callbacks. Around that loop: - The **response** callback fires only when the response actually came from the target host. - A configured maximum body size skips whichever half exceeds it, incrementing a `stats.pscan` counter instead of scanning. - An exception from a rule is caught, logged with the record id and URL, and the loop continues with the next rule. - A rule that takes an unusually long time on one record is reported as a warning naming the rule, the URL and the response content type — which is how you find the rule that is costing you the backlog. ## Where this lives, and why it matters The engine described above — controller, task, options, API component, automation jobs — is the **`pscan` add-on**. Core still contains classes with several of the same names, and those are `@Deprecated` shims kept so older add-ons compile. What did *not* move is the rule contract: `PassiveScanner` and `PluginPassiveScanner` are live, non-deprecated core types. So "where does passive scanning live" needs both halves, and a claim that names only one of them will be wrong about the other.
- Why does ZAP's passive backlog counter reading zero not prove the passive sweep is finished?`ExtensionPassiveScan2.getRecordsToScan()` returns zero whenever passive scanning is disabled or the controller was never started, as well as when the queue really is empty. A run that switched the engine off through the `pscan` API gets the same zero as a run that drained properly, and any wait built on that number returns immediately in both cases.
- If you load a saved ZAP session, will its recorded traffic be passively scanned?No. On starting, the controller sets its cursor to the last history id and steps past it, explicitly to avoid re-scanning existing messages. Everything already in the table is behind the cursor. Only traffic recorded after the controller starts is offered to the rules, so a session reopened for a fresh report will not regenerate passive findings by itself.
- What does the pscan API's clearQueue action actually do to pending messages?It jumps the cursor to the last history id and shuts the pending tasks down, so the queued records are discarded rather than scanned. Tasks already running finish; nothing waiting is picked up. It is a way to stop a runaway backlog, not a way to flush it, and every finding those records would have produced is lost.
The engine is a clerk working down a numbered logbook, not a listener standing at the door. It takes the next entry by number, and whatever it has not reached yet is the backlog someone has to let it finish.
saying these in an interview costs you the question
- Says the proxy hands each response straight to the passive rules.
- Thinks reopening a saved session re-runs passive scanning over its history.
- Reads a zero backlog count as proof the passive sweep completed.
- Claims clearQueue flushes the queue by scanning everything in it.
- Assumes the passive engine still lives in core because core has the class names.