Why does ZAP's `openapi` import reach endpoints that its `spider` add-on never finds?
answer
- a crawl only reaches what links to it
- an API response links to nothing
- the import is not a dry parse
- one request per declared operation
- `Requestor` calls `sendAndReceive`
basics
~20 sA definition lists every operation whether or not anything links to it, and the openapi add-on sends a real request for each one. A link-following crawl can only enqueue what a response already points at.
solid answer
~40 sA crawler builds its request list from what it reads: anchors, form actions, script-referenced URLs, and whatever a browser-driven crawl actually exercises. An API answers with JSON that references nothing, so an endpoint no front end calls is invisible to it. The `openapi` add-on takes the other route: it converts each declared operation into a `RequestModel` and its `Requestor` calls `sendAndReceive` on every one, synthesising parameter and body values from the schema. That fills the history table and the Sites tree directly, which is what the active and passive engines then work from. The import is not a dry parse — it is the traffic.
go deeper
Remember the one-line difference: a crawl follows links, an import reads a list. Be able to say that the import itself sends the requests rather than just reading the file.
Explain the chain — document to operation models to request models to sendAndReceive to history and Sites tree — and say where the values in those requests come from.
Show that you treat the import as live traffic against a target you were authorised to write to, and that you assert the import reached something rather than trusting a green run.
Frame the tradeoff: scan coverage now equals document coverage, so the definition becomes a testing artefact the team has to maintain, and undocumented routes become a blind spot nobody sees.
## Why a crawl comes back nearly empty on an API A link-following crawl is a closure operation. ZAP's `spider` add-on fetches a seed URL, parses the response for things that look like further URLs, enqueues them, and repeats. The browser-driven crawlers (`spiderAjax` and `spiderClient`) widen the input — they drive a real browser, so requests a single-page front end makes are recorded too — but the principle does not change: **a crawl reaches only what something in the system points at.** An HTTP API usually points at nothing. A JSON response is data, not navigation. So the endpoints that stay invisible are exactly the ones you most want tested: - operations no current front end calls — admin routes, an older API version still deployed, a mobile-only path; - non-`GET` operations, because there is no form to submit and no link to follow; - paths whose only caller is another service. ## What the import does instead The `openapi` add-on does not discover; it *enumerates and issues*. The chain is short and worth knowing by name: 1. `SwaggerConverter` reads the document and produces an `OperationModel` for every operation under every path. 2. Each model becomes a `RequestModel` — a method, a URL, headers and a body, with values synthesised from the schema. 3. `Requestor.run(...)` loops those models and calls `HttpSender.sendAndReceive` on each. 4. A `HistoryPersister` listener turns each message into a history record and adds it to the Sites tree. Step 3 is the whole point. **Loading the definition is what puts traffic on the wire**: the `openapi` job has no separate "now go and request these" step and no switch that suppresses the sending. The `soap` and `postman` jobs follow the same shape with their own converters. The `graphql` job differs in one respect worth knowing — its generated operations sit behind a switch of their own — but its import still fetches the schema over the wire. ## Which operations, and with what values Every operation the document declares is converted, not a safe subset: | what the document declares | what the import issues | |---|---| | `get`, `head`, `options` | issued | | `post`, `put`, `patch` | issued, with a generated body | | `delete` | issued — nothing filters it out | | a path parameter such as `/pets/{petId}` | filled from the schema before sending | Values come from the schema in a fixed order of preference — a declared `default`, then the first `enum` member, then a supplied `example`, then a value invented for the declared type or format. That is enough to make most requests well-formed, and it is also why an import against a live environment can create, modify and delete records. Treat the import as a write to the target, because on any document containing state-changing operations it is one. ## What bounds it Two dials, and neither limits anything unless you set it: - `maxMessages` on the job caps how many request models are generated before any are sent. It defaults to `0`, which means no limit. - The context named in the job's `context` parameter can suppress a request, but only through its **exclude** regexes. Being outside the context's include list does not stop anything. So the honest default position is: the import sends one request per declared operation, per server URL the document names, until it runs out of operations. ## What the import still cannot tell you Three limits are worth stating plainly, because the reach of an import is easy to over-read. - **It only knows what the document says.** An operation the team never documented is an operation nothing requested, and there is no crawl depth that recovers it. Coverage of the scan has become coverage of the file. - **A request reaching the server is not a request the server accepted.** Generated values are plausible, not meaningful, so an endpoint with real validation answers many of them with a client error. The URL still lands in the tree, which makes "it was imported" weaker evidence than it looks. - **Nothing about the import is incremental.** Run the job twice and every operation is issued twice; there is no memory of what a previous run already sent. Fetching the document is itself traffic, too. Point the job at a URL rather than a file and the add-on requests that URL through the same sender, so the definition fetch lands in the history alongside the operations it produced. ## How this reads in a pipeline Two practical consequences for someone running this from CI. First, if the import fails or the document is empty, everything downstream has nothing to work on and the run can still finish looking healthy — so assert that the import reached something (the add-on bumps an `openapi.urls.added` counter you can test on) rather than trusting a green exit. Second, the coverage of the whole scan is now the coverage of the document: an operation the team forgot to document is an operation nothing tested, and no amount of crawl depth will recover it.
- Does the import issue state-changing methods, or only safe ones?Every operation the document declares is converted, including `post`, `put`, `patch` and `delete`. There is no filter for safe methods and no dry-run switch. Against a live environment the import will create and delete records, which is why the target has to be one you were permitted to write to.
- If a definition is imported, is there still any reason to crawl the same host?Yes, when the host serves more than the API — a console, documentation, static assets or a login flow. The definition covers only what it declares, so anything outside it still needs a crawl. The two request lists are complementary, not alternatives.
- Where do the parameter and body values in the generated requests come from?From the schema, in preference order: a declared `default`, then the first `enum` member, then a supplied `example`, then a value synthesised for the declared type or format. They are plausible rather than meaningful, so an endpoint with real validation may reject many of them.
A crawl can only walk corridors that somebody hung doors on. A definition is the floor plan: it lists every room, including the ones with no door from the lobby.
saying these in an interview costs you the question
- Thinks importing a definition only parses it and sends nothing
- Believes a deep enough crawl would find the same API endpoints
- Assumes only GET operations are issued by the import
- Expects destructive operations such as DELETE to be filtered out
- Thinks the imported URLs must be crawled afterwards before they can be scanned