skip to content

In ZAP's spiderAjax and spiderClient jobs, what does logoutAvoidance do and why is it not a guarantee?

level: seniorimportance: should knowfreq 44%

answer

  1. It is about words, not about sessions
  2. Matches visible text against a short list
  3. Different default on each crawler
  4. An icon or a translation defeats it

basics

~20 s

logoutAvoidance tries to stop the crawl clicking sign-out controls, matching an element's visible text against a short English word list. Text is all it reads, so an icon, a translated label or unusual wording slips past.

solid answer

~50 s

A browser crawl clicks things, and one of the things on an authenticated page is the control that ends the session. `logoutAvoidance` is the guard against that, and both crawlers implement it as a **text heuristic**. The `spiderAjax` job builds don't-click rules for anchor, span and button elements whose lowercased, space- and hyphen-stripped text contains `logout`, `logoff`, `signout` or `signoff`. The `client` job normalises each clickable component's visible text the same way and compares it against a shared list of logout indicators, applied to any component type. Neither inspects the response, the URL or your application's semantics. An icon-only control, a non-English label, a menu item reached by hovering, or a link whose text says something else entirely will still be clicked — and once the session ends, the rest of the crawl explores the anonymous application while reporting success.

code

yaml · 8 lines
yaml
- type: spiderAjax
  parameters:
    url: https://example.com
    logoutAvoidance: true        # off by default on this job
    excludedElements:
      - description: the account menu's sign-out link
        element: a
        xpath: "//a[@id='logout' and @class='logout']"

go deeper

for a junior

Remember that a browser crawl clicks whatever is on the page, including the sign-out control, and that this setting is the built-in attempt to avoid that rather than a promise it will not happen.

for a middle

Explain that both crawlers match on the element's visible text against a short English word list, and say which crawler ships it on and which ships it off.

for a senior

Show the operational habit: pin the real element where the job lets you, exclude the sign-out path from scope as a second layer, and verify afterwards that authenticated URLs are actually in the results.

for a principal

Decide what evidence your organisation requires that an authenticated scan stayed authenticated, because a heuristic built on button labels is not something a report should rest a coverage claim on.

## The failure it exists to prevent You authenticate the crawl, it starts exploring, and somewhere on the page is a sign-out control. The crawler clicks it, because clicking things is its job. From that moment the browser is anonymous, every subsequent page is the logged-out application, and the run still finishes cleanly. The URL count may even go **up**, because the marketing pages are usually bigger than the product. ## How each crawler implements the guard Both do the same kind of thing, and both do it on text. - **`spiderAjax`** constructs a set of don't-click rules from a small fixed vocabulary — `logout`, `logoff`, `signout`, `signoff` — crossed with three element names: anchor, span and button. Each rule is an XPath that lowercases the element's text and strips spaces and hyphens before testing for the fragment. The rules are handed to the bundled crawler as exclusions. - **`spiderClient`** normalises the visible text of each clickable component the same way — lowercase, spaces and hyphens removed — and checks it against a list of logout indicators shared across several add-ons. It is not restricted to particular tag names, so it covers controls the AJAX spider's rules would not match. | | `spiderAjax` | `spiderClient` | |---|---|---| | default | **off** | **on** | | what it matches on | element text, via XPath | component text, normalised in code | | which elements | anchor, span, button | any clickable component | | deterministic escape hatch | `excludedElements` in the job | none in the job | That last row is the awkward one. The crawler that ships the guard **off** is the one that also lets you name elements to never click; the crawler that ships it **on** has no equivalent job parameter. So on the newer crawler the heuristic is all you have. ## Why it is a heuristic and not a guarantee The list is short, English, and about words. Everything that is not a matching word gets clicked: 1. **An icon.** A control whose visible text is empty cannot match anything; the check bails out on blank text. 2. **Another language.** A button reading *Abmelden*, *Déconnexion* or *Cerrar sesión* is not on the list. 3. **Different wording.** *End session*, *Leave*, *Exit*, *Switch account* — all ordinary, none matched. 4. **Indirect routes.** A session can end without any element that says so: a link into an identity provider's sign-out flow, a control that revokes the current device, a "sign out everywhere" buried in settings. 5. **A false positive in the other direction.** A help article titled "How do I log out?" matches, and the crawler declines to click a perfectly ordinary link. That silently removes a branch of the site from the crawl — which is exactly the sort of coverage loss that appears when you compare the two crawlers and blame the engine. The guard is worth having. It is not a control you can build a claim on. ## What to do instead, in a pipeline - **Turn it on where it is off.** On the AJAX spider job it defaults off; set it rather than inheriting it, and write it down even on the job where it is already on. - **Name the elements you actually have.** The AJAX spider job accepts an `excludedElements` list where each entry can pin an element name plus an XPath, a text value, or an attribute name and value. That is deterministic: it matches your application's markup rather than a guess about wording. ZAP's own tests use a logout anchor as the worked example. - **Prove the session survived, don't assume it.** The cheap version is to check, after the crawl, that URLs only reachable while authenticated are present in what the run found. A run that reports plenty of URLs and none of the authenticated ones is the signature of a crawl that logged itself out early. - **Exclude the sign-out path from the crawl's scope as well.** A URL exclusion is a second, independent layer; it does not depend on what the control's label says. - **Expect to revisit it.** Application wording changes, a redesign turns a text link into an icon, and a guard built on words quietly stops matching. Nothing fails when that happens, which is why it needs a check that does not depend on the guard. ## The sentence to carry `logoutAvoidance` reduces the chance that a crawl ends its own session. It does not establish that the session lasted, it defaults differently on the two crawlers, and on the one where it defaults on there is no job-level way to name a specific element instead. Treat it as a useful default, and put the assurance somewhere that does not read the button's label.

  • Why is excludedElements stronger than logoutAvoidance?
    Because it matches your markup rather than a guess about wording. An entry can pin an element name together with an XPath, a text value, or an attribute name and value, so it identifies the control you actually have. It is a `spiderAjax` job parameter, though — the client spider job has no equivalent.
  • How do you tell that a crawl logged itself out partway through?
    Compare what it found against URLs only reachable while authenticated. A run that reports a healthy URL count with none of the authenticated paths in it, or one where the authenticated paths all appear early and then stop, is the signature. The exit status will not tell you, because logging out is not an error.
  • Can logoutAvoidance cost you coverage?
    Yes, in both directions. It matches on text, so any element whose label happens to contain one of the words — a help page titled "How do I log out?", for instance — is declined. That branch simply is not explored, and nothing reports it, so a coverage difference between the two crawlers is not automatically a difference in engine capability.

saying these in an interview costs you the question

  • Treats logoutAvoidance as a guarantee the session survives
  • Assumes it inspects the response rather than element text
  • Thinks it is on by default on both browser crawlers
  • Never checks whether authenticated pages appear in the results
  • Ignores that it can also decline harmless links