skip to content

In an OWASP ZAP automation plan, what do a context's urls and includePaths entries become?

level: middleimportance: should knowfreq 46%

answer

  1. one key is a convenience, two are literal
  2. something is appended to only one of them
  3. checked as one thing, used as another
  4. a bare URL under includePaths covers itself only
  5. the appended wildcard is what makes it recursive

basics

~20 s

Both become include regexes on the context, but by different routes. Each urls entry is checked as a URI and then gets a wildcard appended; each includePaths entry is added exactly as written, so a bare URL there matches only itself.

solid answer

~40 s

When the plan builds the context, each `urls` entry has plan variables substituted, is checked as a URI, and is then added as an include regex with `.*` appended — which is what makes "everything under this URL" work. Each `includePaths` and `excludePaths` entry is added **raw**, with nothing appended. Since the context layer applies a full match, a bare URL written under `includePaths` matches that one URL and nothing beneath it. The append is also a plain string concatenation, so a `urls` entry is validated as a URI and then used as a regular expression: every unescaped `.` is a wildcard, and a `+`, `(` or `?` in the path changes what the pattern means. No key in the block is ever matched against a query string.

code

yaml · 7 lines
yaml
contexts:
  - name: shop
    urls:
      - https://example.com/shop
    includePaths:
      - https://example\.com/api
      - https://example\.com/api/.*

go deeper

for a junior

Recall that both keys end up as include regexes on the context, and that only urls gets a wildcard appended for you.

for a middle

Explain why a bare URL under includePaths matches only itself, and why an unescaped dot in either key is a wildcard.

for a senior

Treat the block as reviewable code: the entry is checked as a URI and used as a regex, so validation passing tells you nothing about what it will match.

for a principal

Decide how scope blocks are reviewed and owned across repositories, given that a wrong one produces a quiet green run rather than a failure anyone is paged for.

## Two keys that look alike and are not A context block in an OWASP ZAP automation plan carries `urls`, `includePaths` and `excludePaths`. The shipped template describes `urls` as "a mandatory list of top level urls, everything under each url will be included" and `includePaths` as "an optional list of regexes to include". That is the whole difference, and it is easy to read past. When the plan creates the context, the automation add-on does this: 1. For each entry in **`urls`**, substitute any plan variables, check that the result parses as a URI, then append `.*` and add it as an **include regex**. 2. For each entry in **`includePaths`**, substitute variables and add it as an include regex **exactly as written**, with nothing appended. 3. For each entry in **`excludePaths`**, substitute variables and add it as an exclude regex, likewise exactly as written. So `urls` is a convenience that builds a regex for you, and the two `*Paths` keys are raw regex lists. A bare URL written as an `includePaths` entry therefore matches **that URL and nothing beneath it**, because the context layer applies a full match and no wildcard was added. ## The append is unescaped Step 1 is a string concatenation. The entry is checked as a *URL* and then used as a *regex*, and nothing between those two steps escapes it. Consequences: - Every `.` in the host and path is a **wildcard**. `https://example.com/` becomes `https://example.com/.*`, which also matches `https://exampleXcom/` and any other single character in those positions. Usually harmless, occasionally not. - A URL containing regex metacharacters means something other than itself. A `+` in a path is a repetition operator; a `(` starts a group and may make the pattern fail to compile; a `?` makes the preceding character optional. - Putting a query string in a `urls` entry is doubly wrong: the `?` changes the pattern's meaning, and the context layer removes the query from the URL before matching anyway, so the part you were trying to pin can never be compared. ## What is validated, and what is not | key | validated as | used as | |---|---|---| | `urls` | a URI — an entry that does not parse is reported as an error | a regex, after `.*` is appended | | `includePaths` | a regular expression | a regular expression | | `excludePaths` | a regular expression | a regular expression | The mismatch on the first row is the whole trap: an entry can pass the check it is given and still be a different pattern from the one you meant, because the check and the use are asking different questions of the same string. ## Writing the block so it does what it says ``` contexts: - name: shop urls: - https://example.com/shop includePaths: - https://example\.com/api/.* excludePaths: - https://example\.com/shop/logout.* ``` Three habits make this reliable: - **Let `urls` do the work** for the ordinary "this site and everything under it" case. It is the only key that appends the wildcard, and it is the one the template says is mandatory. - **Escape the dots** in anything you write into `includePaths` or `excludePaths`, and end it with `.*` if you mean "and everything beneath". Neither happens for you. - **Keep the query out of all three.** No key in this block is matched against a query string. ## The variable substitution happens first All three keys run through the plan's variable replacement before anything else, so an entry may be assembled from `env.vars` or from the environment. That is what makes one plan serve several deployments, and it is also a second place for a scope block to go quietly wrong: a variable that resolves to an empty string leaves an entry that is still a legal URI or a legal regex, just a different one from the one you read in the file. Whatever you conclude from reading the block, the pattern that was actually installed is the substituted one. ## Why this is worth a code review A scope block is the one part of a plan whose defects are invisible in the run's output. A wrong job name is an error; a wrong URL is a connection failure; a wrong include regex is a scan that finishes green having visited almost nothing. The scan reached fewer URLs, reported fewer findings, and exited the same way a good run does. Reviewing the block means reading each entry and asking which of the three translations above it is going through — and that is a question about the key it is written under, not about the string itself.

  • A `urls` entry parsed fine but the crawl reached almost nothing. What would you check first?
    That the entry means what it looks like as a regular expression. It was validated as a URI and is then used as a pattern with `.*` appended, so metacharacters in the path are live: a `+` is a repetition operator, a `(` may stop it compiling, and a `?` makes the preceding character optional. Read it back as a regex rather than as a URL.
  • Should scope live in `urls` or in `includePaths`?
    Use `urls` for the ordinary case — it is the mandatory key, and it is the only one that appends the wildcard that makes a prefix behave like one. Reach for `includePaths` when you need a pattern a URL cannot express, and then write it as a real regex: escaped dots, and a trailing `.*` if you mean everything beneath.
  • Why is a wrong scope block harder to notice than a wrong job name?
    Because it is not an error. A bad job name is reported and a bad target fails to connect, but an include regex that matches too little just produces a run that visited fewer URLs, reported fewer findings and exited the same way a good one does. The only signal is the count of what was reached.

saying these in an interview costs you the question

  • Thinks includePaths behaves like urls and covers everything beneath
  • Writes a plain URL under urls and expects the dots to be literal
  • Believes urls entries are prefixes rather than regexes with a wildcard appended
  • Puts a query string in a urls entry and expects it to narrow scope
  • Assumes an entry that passed validation must mean what it looks like