Beyond issuing requests, what does a ZAP `openapi` job write into the history, Sites tree and context?
answer
- four places, not one
- history entry plus a Sites tree node
- recorded as user-sent, not spider
- the context gains an include regex
- `{petId}` becomes a data-driven node
basics
~10 sIt records every message it sent as a history entry typed TYPE_ZAP_USER, adds each to the Sites tree, bumps an openapi.urls.added counter, and appends an include regex per imported path to the named context.
solid answer
~30 sEach message the import sends goes through a `HistoryPersister`, which creates a `HistoryReference` and calls `getSiteTree().addPath(...)`. The history type is `TYPE_ZAP_USER` — the type used for hand-sent messages — not `TYPE_SPIDER`, which the add-on uses only when it is driven from the spider path. It also increments an `openapi.urls.added` statistic. The part people miss is the context write: for every imported path the converter calls `Context.addIncludeInContextRegex`, turning `/pets/{petId}` into `/pets/[^/?]+`, and registers a `Variant` so the Sites tree groups those URLs under one data-driven node instead of one node per id.
go deeper
Know that an import leaves records behind: entries in the history and nodes in the Sites tree. That is what later steps read from.
Be able to name all four effects and say which history type is used and why. The context write is the one that separates a read of the docs from a read of the behaviour.
Show you verify the import produced something before trusting anything downstream, and that you know the plan's context is mutated by the job rather than fixed by the plan.
Consider what it means for reproducibility that a job rewrites the scope definition the plan declared, and where you would want that recorded in the run's artefacts.
## Four places the import leaves a mark An import is usually described as "it fills the history". It does, but that is one of four side effects, and the other three change how the rest of the run behaves. | what changes | mechanism | why it matters | |---|---|---| | the history table | `HistoryPersister` creates a `HistoryReference` per message | this is what the passive engine drains | | the Sites tree | `getSiteTree().addPath(historyRef, message)` | this is what you can select and attack | | a statistic | an `openapi.urls.added` counter is incremented | gives you something to assert on | | the named context | `Context.addIncludeInContextRegex(...)` per imported path | the import edits your scope definition | ## The history type, and why it is not the spider's The persister picks its type from the initiator the requestor was built with: - driven from the import paths, the requestor is constructed with the **manual-request** initiator, so messages are recorded as `TYPE_ZAP_USER` — the type documented for messages a user sends by hand from a resend or manual-request dialogue; - only when the add-on is driven from the spider integration is the initiator the spider's, and only then are messages recorded as `TYPE_SPIDER`. This is worth getting right because it is easy to assume anything that populates the tree is "spider traffic". It is not, and a filter written on that assumption will hide the import's messages. It also matters that `TYPE_ZAP_USER` is one of the history types the passive rule base class treats as in scope by default, so imported messages **are** passively scanned — the type differs, the treatment does not. ## The context write, which nobody expects When the job names a context, the converter walks the imported operations and, for each path that the context does not already include and does not exclude, adds an include regex: - a path with no parameters is added quoted, so it matches literally; - a path such as `/pets/{petId}` is rewritten to `/pets/[^/?]+` and added as a regex. So after the job runs, the context you declared in `env.contexts` is not the context you wrote. That is usually what you want — the imported operations are now in scope for jobs that respect scope — but it means the plan's context block is a starting point, not a fixed definition, and two imports into the same context accumulate. ## The tree shape The add-on also registers a `Variant` whose only real job is `getTreePath`. Given a message, it matches the URL against the path templates it learned from the definition, and returns a tree path where the parameter segment is marked as a data-driven node. The practical effect: `/pets/1`, `/pets/2` and `/pets/9999` collapse into a single node in the Sites tree instead of one node per identifier. That is not cosmetic. It is the difference between an active scan attacking one node once and attacking one node per identifier it happens to have seen, and between a readable tree and an unusable one. It also only works for messages whose method and path actually match an imported operation, so traffic that arrived some other way is unaffected. ## What this means when you drive it from a plan Three habits follow: 1. **Assert the import worked.** A run where the import silently produced nothing still proceeds, and the later jobs quietly scan an empty tree. The statistic the import bumps is the cheapest evidence that it did something. 2. **Do not hand-maintain include regexes for imported paths.** The job will add them; writing them yourself as well only creates drift. 3. **Expect the history to look hand-sent.** When you are reading the record afterwards, the import's messages are typed as user-sent, and that is the correct reading — a person, via a plan, asked for exactly those requests. The job also reports what it did. The add-on collects the history references it created into a results object, and the job writes the count of URLs added into the plan's progress output, so the run log carries the number even when nothing tested it. The sibling definition jobs persist the same way with their own listeners, and the `soap` importer is the bluntest of them: it writes the user-sent type unconditionally, with no initiator test at all. If you are building a filter or a report that separates crawl traffic from import traffic, key it on the job that ran, not on the history type.
- Why does the Sites tree not fill with one node per identifier after an import?The add-on registers a `Variant` that supplies a tree path derived from the definition's path templates. Segments that were `{petId}` in the document are marked as data-driven nodes, so every value collapses into one node. It applies only to messages whose method and path match an imported operation.
- If the import records messages as user-sent rather than spider traffic, do passive rules still see them?Yes. The passive rule base class treats `TYPE_ZAP_USER` as one of the history types it applies to by default, alongside proxied and crawler traffic. The type changes how the record is labelled, not whether it is scanned.
- What happens if two openapi jobs in one plan name the same context?Both add include regexes to it, and they accumulate — the second import does not replace the first one's entries. It only skips a path the context already includes, so duplicates are avoided but the context keeps growing across jobs.
saying these in an interview costs you the question
- Says imported messages are recorded as spider traffic
- Thinks the import leaves the context untouched
- Believes imported messages are skipped by passive rules
- Assumes the Sites tree gets one node per path parameter value
- Treats a completed import as proof that URLs were actually added