skip to content

Contexts and Scope

Several independent layers each decide what is in scope, they do not agree on how a regex is matched, and one of them throws the query string away before it looks.

on this pageshow

questions

6

In an OWASP ZAP context, what is an include or exclude regex actually matched against?

level: middleimportance: must knowfreq 62%

answer

  1. the pattern must describe everything
  2. something is discarded before the match
  3. everything from the first question mark
  4. matches(), not find(), and case-insensitive
  5. include is consulted before exclude

basics

~20 s

A full, case-insensitive match against the URL with everything from the first question mark onward removed. A prefix will not match because the pattern must describe the whole string, and a rule written against a query parameter can never fire.

solid answer

~40 s

Core's `Context` compiles every include and exclude entry with `Pattern.CASE_INSENSITIVE`, then calls `matcher(url).matches()` — a **full** match, so the pattern has to describe the whole string rather than a prefix. Before that, `isIncluded` and `isExcluded` both truncate the URL at the first `?`, so the query string is gone before any pattern sees it: an exclude such as `.*\?debug=true` cannot fire at this layer, and neither can anything keyed on a parameter name or value. That is also why the regexes the program generates for you — from a site-tree node, or from a plan's `urls` list — end in `.*`. Inclusion is evaluated first: `isInContext` returns false straight away if nothing includes the URL, so an exclude rule with no matching include changes nothing.

code

yaml · 8 lines
yaml
contexts:
  - name: shop
    urls:
      - https://example.com/shop
    excludePaths:
      - .*\?debug=true
      - https://example\.com/shop/logout.*
      - https://example\.com/shop/reports

go deeper

for a junior

Recall that a context holds two lists of regular expressions — include and exclude — and that both are matched against the whole URL rather than searched inside it.

for a middle

Be able to explain the three properties of the match: full match, case-insensitive, and query string removed before the pattern is applied. Show why a trailing .* is needed for a prefix.

for a senior

Explain how you would review a scope block before a scan rather than diagnose it after: read each entry, check that it describes the whole string, and check that the string it is handed still contains what it matches on.

for a principal

Own the convention. Decide whether your teams hand-write these regexes at all, or only ever generate them, and what a reviewer is expected to check when a scope change lands in a repository.

## What a context is, and what it holds A **context** in OWASP ZAP is a named bundle of settings that core's `Context` class owns: a list of **include** regular expressions, a list of **exclude** regular expressions, an `inScope` flag, and the authentication, session-management and technology settings that go with them. Everything the program later calls "scope" is derived from those lists. Whether you create the context by right-clicking a node in the site tree, by calling the control API, or by writing an `env.contexts` block in an automation plan, what lands in the context is the same thing: strings that are compiled into `java.util.regex.Pattern` objects. ## The three properties of the match Every include and exclude entry is compiled once, with `Pattern.CASE_INSENSITIVE`, and then applied in `isIncluded(String url)` and `isExcluded(String url)`. Three properties of that call decide whether your rule ever fires: 1. **It is a full match, not a search.** The code calls `matcher(url).matches()`. The pattern has to describe the entire remaining string. `/admin` does not match `https://example.com/admin` — you need something like `https://example\.com/admin.*`. 2. **It is case-insensitive.** The compile flag is set on every entry, so `/ADMIN` and `/admin` are the same rule here. That is not true of every exclusion layer in the program, which is why it is worth knowing it is true of this one. 3. **The query string is removed first.** Both methods begin by looking for the first `?` and, if there is one, truncating the URL there. The pattern never sees the query. That third property is the one that costs people an afternoon. A rule written to keep a destructive endpoint out of scope by its parameters — anything of the shape `.*\?action=delete.*`, or a pattern naming a parameter value — **cannot match at this layer**, because the text it is written against has already been thrown away. The rule is not wrong, and it does not error: it is simply evaluated against a string that no longer contains what it is looking for. ## Include first, exclude second `isInContext(url)` truncates the URL, asks `isIncluded`, and returns `false` immediately if nothing includes it. Only then does it consult `isExcluded`. Two consequences follow: - An exclude rule is **dead weight** unless an include rule already matched the same URL. Excluding something you never included changes nothing. - A context with an empty include list contains nothing at all. Adding excludes to it does not make it broader. ## Why the generated regexes end in `.*` When the program writes a pattern for you it writes a full-match-shaped one. Building a regex from a site-tree node escapes each path segment, joins them with `/`, strips the parenthesised parameter summary the tree appends to a node's name, and — when children are to be included — puts `.*` on the end. A plan's `urls` list is handled the same way: each entry has `.*` appended before it becomes an include regex. Those trailing wildcards are not decoration; they are what turns a prefix into a legal full match. If you hand-write an include rule and leave the `.*` off, you have included exactly one URL. ## Reading a rule before you trust it | you wrote | what the matcher sees | result | |---|---|---| | `https://example\.com/shop` | the URL, query removed | matches only that exact URL | | `https://example\.com/shop.*` | the URL, query removed | matches that URL and everything beneath it | | `.*\?debug=true` | the URL, query removed | never matches — the `?` and all after it are gone | | `.*/ADMIN/.*` | the URL, query removed | matches `/admin/` too — the compile flag is case-insensitive | | `https://example.com/shop.*` | the URL, query removed | matches, but the unescaped dots also match any character | The last row is the quiet one. A URL pasted straight into an include list is a **regular expression**, not a literal: every `.` is a wildcard, and `+`, `(`, `?` and `|` all carry their regex meaning. It usually still matches what you intended, which is exactly why nobody checks it. ## What this means in a pipeline For someone driving the program from CI, the practical form of all of the above is: scope is declared as full-match regexes over a **path**, and never over a query. If the thing you have to keep in or out is identified by a parameter rather than by a path, a context rule is the wrong tool and you need a different layer — one that is applied by a different component, with different matching rules. The check to run before a scan, not after it, is to read each include and exclude entry and ask two questions of it: *does this describe the whole string*, and *does the string it will be handed still contain the part I am matching on?* Three habits close almost all of it: - **End an include with `.*`** unless you genuinely mean one URL, because the match is anchored. - **Escape the dots** in anything you hand-write, because the entry is a pattern and not a literal. - **Never key a context rule on a query parameter**, because that text is removed before the pattern is applied and no error tells you so.

  • An include entry and an exclude entry both match the same URL. Which wins?
    The exclude. `isInContext` first asks whether any include pattern matches and gives up if none does; if one does, it returns the negation of the exclude check. So exclusion is applied on top of inclusion and always overrides it, but it is only ever consulted for a URL that was included in the first place.
  • Why does an include entry copied straight from the address bar usually still work?
    Because the characters that differ in meaning — mostly `.` — match themselves as well as everything else. An unescaped dot is a wildcard, so `example.com` also matches `exampleXcom`, which is broader than intended rather than narrower. It bites when the URL carries `+`, `(` or `?`, where the regex meaning changes what matches rather than merely widening it.
  • How do you keep a destructive endpoint out of a run when it is identified only by a query parameter?
    Not with a context rule — that layer never sees the query. You need an exclusion applied by a component that matches the whole URI, such as the session exclude-from-scan list or the `network` add-on's global exclusions, and you have to accept that those are configured outside the context block and matched by different rules.

It is a doorman who tears the stub off your ticket before reading the name. Whatever was printed on the stub cannot be part of the check, however carefully you wrote it there.

saying these in an interview costs you the question

  • Says a context regex matches a prefix of the URL
  • Writes an exclude against a query parameter and expects it to fire
  • Thinks a context's include and exclude lists are case-sensitive
  • Assumes a bare URL in includePaths also covers everything beneath it
  • Believes an exclude rule works without any include rule matching first
open as a page

In an OWASP ZAP automation plan, what do a context's urls and includePaths entries become?

level: middleimportance: should knowfreq 46%

basics

~20 s

Both become include regexes on the context, but by different routes. Each urls entry is checked as a URI and then gets a wildcard appended; each includePaths entry is added exactly as written, so a bare URL there matches only itself.

open as a page

In OWASP ZAP, what decides whether a URL is in scope, and what has no say in it?

level: middleimportance: should knowfreq 54%

basics

~20 s

Contexts decide it, and nothing else does. A URL is in scope when some context with its in-scope flag set includes it and no in-scope context excludes it. Exclusion crosses context boundaries, and scope has no storage of its own.

open as a page

In OWASP ZAP, where can a URL be excluded, and which part of a run does each place affect?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Three places: a context's exclude list, which subtracts from scope; the session's per-subsystem lists for the proxy, the active scan, the crawlers and websockets; and the network add-on's global exclusions, read by the proxy, the crawlers and the active scanner.

open as a page

Does the same exclusion regex behave identically in every OWASP ZAP layer that reads it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

No. Only the full-match anchoring is shared. A context matches case-insensitively with the query already removed, the proxy handler and the crawler match the whole URI case-sensitively, and the active scanner matches that same shared list case-insensitively.

open as a page

How would you standardise where OWASP ZAP scope is declared across many unattended pipelines?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Put the boundary in the plan's context block. It is the only scope layer a plan can express, so the only one that lives in the repository and gets reviewed. Bound the exceptions, and say what it cannot express: a query-keyed rule.

open as a page