skip to content

How does ZAP decide, part-way through a scan, that an authenticated user is still logged in?

level: middleimportance: must knowfreq 52%

answer

  1. carrying and checking are different objects
  2. two patterns, one strategy, one cadence
  3. polling, not per-response, is the default
  4. the indicator match is a search
  5. between polls the last verdict is reused

basics

~20 s

A context's verification method matches a logged-in or logged-out regex against traffic. Its strategy decides which traffic: the request, the response, both, or — the default — a separate poll request sent to a URL you nominate on a cadence you set.

solid answer

~40 s

Carrying the session and checking the session are different objects. Checking is the context's **verification method**, which holds a `loggedInRegex`, a `loggedOutRegex` and a strategy. The each-request, each-response and both strategies match those patterns against traffic the scan is already producing. The default strategy instead **polls**: it sends a separate request to a nominated URL, as the user, and matches the patterns against that response. A poll is not sent every time — a cadence value and a unit, either seconds or requests, govern how often, and between polls the previous verdict is reused and a counter recorded. In a plan the block is `authentication.verification`, and its `method` is one of `response`, `request`, `both`, `poll` or `autodetect`.

code

yaml · 6 lines
yaml
authentication:
  verification:
    method: poll                  # the default strategy
    loggedInRegex: 'Sign out'
    pollUrl: "https://example.com/api/me"
    pollUnits: requests

go deeper

for a junior

Know that the tool needs to be told what a logged-in page looks like — a pattern — and that by default it finds out by fetching a URL you nominate rather than by reading every response.

for a middle

Explain the strategies and what each matches against, that the match is a search rather than a full match, and that between polls the previous verdict is reused rather than re-derived.

for a senior

Reason about cost and honesty together: what the poll does to the target's logs and rate limits, how the cadence unit changes the size of the unchecked window, and how you would prove from counters that a run stayed authenticated.

for a principal

Own the default for the organisation — which strategy, which kind of poll endpoint is acceptable, and what evidence a pipeline must produce before an authenticated scan result is allowed to mean anything.

## The object that answers the question A context holds a **verification method** alongside its authentication and session-management methods. It owns four things: a **logged-in indicator** pattern, a **logged-out indicator** pattern, a **strategy** that says what text those patterns are matched against, and a **cadence** that applies only to the polling strategy. The strategy is the part people skip, and it is the part that decides how much the check costs and how honest it is. ## The five strategies | plan `method` | what the patterns are matched against | extra traffic | |---|---|---| | `request` | the request header and body the tool just built | none | | `response` | the response header and body just received | none | | `both` | request and response, all four parts | none | | `poll` | a separate request sent to a nominated URL, as the user | one request per cadence | | `autodetect` | nothing — the verdict is deferred to an add-on's detection rules | none | **Polling is the default.** That matters more than it sounds, because a poll strategy needs a poll URL and the default value of that URL is empty. It also matters because the poll costs traffic the crawl did not ask for. The matching itself is a **search**, not a full match: the pattern is run against each candidate string and any hit anywhere counts. A logged-in pattern of `Log Out` therefore matches a script bundle, a cached navigation fragment, or an error page that still renders the header — which is a different rule from the full-string matching used for context include and exclude rules, and mixing the two up produces indicators that look reasonable and never fire correctly. ## What a poll actually is A poll is a real HTTP request to the target. Specifically: 1. A message is built for the poll URL. A method and a body may be set, and additional headers may be attached; supplying a body with no explicit method makes it a POST. 2. The request is marked as belonging to the user, so the session-management method stamps the session onto it exactly as it would any other request. 3. It is sent through the tool's own sender, subject to the same rate limiting as scan traffic, and it is added to history tagged as verification traffic so you can find it afterwards. 4. Its response is matched against the indicator patterns, and the verdict, the poll timestamp and the request counter are stored on the user's authentication state. Two practical consequences follow. The poll reaches the application, so it is visible to the application's own logging and rate limits. And because it carries the session, **a poll URL that itself refreshes a session will keep the session alive** — the health check becomes part of the treatment. ## The cadence, and the gap it opens The cadence has a value and a unit, and the unit is either **seconds** or **requests**. With the polling strategy, each check first asks whether the previous poll said *logged in* and whether the cadence has elapsed. If the answer is *yes, and not yet*, the tool returns *still authenticated* **without checking anything** and increments an `assumed-in` counter. That is the correct design — polling every request would double the traffic — but it is also the window in which a session can die unnoticed. The unit choice changes the shape of that window: a requests-based cadence scales with how busy the scan is, while a seconds-based one does not, so a fast scan under a seconds cadence can push a great many requests into a single gap. ## Where the answer is recorded Every verdict increments one of five counters under `stats.auth.state.`, recorded against the site: - `loggedin` — an indicator matched positively; - `assumedin` — the cadence had not elapsed, so nothing was checked; - `unknown` — a logged-out pattern was configured and did not match; - `loggedout` — the verdict was negative, and a re-login will be attempted; - `noindicator` — no pattern was configured at all, so the check was skipped and *authenticated* returned. Those five are the whole story of an authenticated run. **Four of the five return *authenticated*** — only `loggedout` is a negative verdict — and of those four, exactly one, `loggedin`, rests on a pattern having actually matched. For anyone running the tool unattended, reading those counters is how you find out whether a scan was really logged in, and it is far more reliable than reading the findings.

  • What is the difference between the request-based and seconds-based poll units in practice?
    A requests unit counts scan requests between polls, so the check rate scales with how busy the scan is. A seconds unit counts wall-clock time, so a fast scan can push far more requests into one gap. Requests-based cadence bounds the damage from a mid-scan logout more predictably.
  • Why does polling reach the application at all, rather than being an internal check?
    Because the only way to know the target still accepts the session is to use it. The poll is built as a real message, marked as the user's so the session is stamped on, sent through the normal sender under normal rate limiting, and recorded in history with a verification tag.
  • With an indicator configured, what happens if the polling strategy is selected but no poll URL is set?
    Building the poll throws, the failure is logged as a warning, and the check returns *not authenticated*. That verdict drives a fresh login attempt and a resend on every request, so the symptom is a scan that logs in constantly and crawls slowly rather than a scan that stops.
  • Is the verification method the same object as the session-management method?
    No. Session management stamps the session onto outgoing requests and strips it off again. Verification only produces a verdict about whether the session still works. They are configured separately on a context and either one can be right while the other is wrong.

A polling strategy is a wristband checked at the door on a schedule rather than at every bar. Between checks the tool assumes the band is still on.

saying these in an interview costs you the question

  • Assumes every response is checked against the indicators
  • Thinks the poll is an internal check that sends no traffic
  • Treats the indicator regex as matching the whole response
  • Says a poll URL is optional for the polling strategy
  • Cannot name where the verification verdict is recorded