skip to content

A ZAP scan finished clean but every page reached was the login page. How do you diagnose the form login?

level: seniorimportance: must knowfreq 65%

answer

  1. a clean exit is not evidence of coverage
  2. the login step reports almost nothing
  3. two messages are recorded, read both
  4. compare the sent body against a real login

basics

~20 s

Read the two recorded authentication messages and compare the login request that was actually sent against a real browser login. The login step reports only preparation and send failures, so a rejected login leaves no error behind.

solid answer

~50 s

Start from the fact that **a failed login is silent**. Core's form method records a failure only if it could not prepare or could not send the request; a response of any status, including the login page again, counts as sent. So go to the recorded history and read both messages - the pre-flight `GET` of `loginPageUrl` (or of `loginRequestUrl` when that parameter is blank) and the prepared login request. Check the substituted body byte for byte: a misspelled `{%username%}` is posted literally, because substitution is a plain string replace and the configuration check never looks for the tokens. Then check the login response - a 200 carrying the form again is a rejection - and check the scanning job actually named a `user`. Whether the tool *noticed* it was logged out is a separate verification concern.

go deeper

for a junior

Learn where the authentication messages are recorded and how to read the request that was actually sent - most of these failures are visible in those bytes.

for a middle

Explain why the failure is silent: the login step records only preparation and send failures, so any HTTP response at all counts as a request that succeeded.

for a senior

Work the list in order - sent body, comparison against a real login, the response, whether the job named a user - rather than guessing, and describe what each step rules out.

for a principal

Own the structural fix: a run with zero authenticated coverage should be a failed pipeline, which means asserting on a behind-the-login signal rather than trusting the run's own verdict.

## Why a clean exit proves nothing here A scan that never authenticated does not look like a broken scan. The crawl succeeds - it just crawls the public surface. The rules run - they just run against the login page. Findings come back - there are simply very few of them. **The absence of authenticated coverage is silence, not an error**, and the wrapper's exit code has no opinion about it. Worse, the login step's own error surface is unusually narrow. Core's form-based login records a failure in two situations only: it could not **prepare** the request, or it could not **send** it. A login request that is sent and answered - with anything, including the login page again, a 401, or a 302 back to the login form - is a request that succeeded as far as that code is concerned. ## What the login actually does, so you know what to look at One authentication attempt is **two HTTP requests**, not one: 1. A **GET** first, of `loginPageUrl`. If that parameter is blank, it fetches `loginRequestUrl` instead. Its job is to obtain fresh cookies and, if an anti-CSRF token is present in the response and the anti-CSRF extension is available, a fresh token value. 2. The **prepared login request**: the configured body with the credential tokens substituted, the anti-CSRF token copied across if one was found, and the `Cookie` header explicitly cleared before sending. Both are added to the recorded history, which is what makes this diagnosable at all - you can read the exact bytes that went out. ## The failure modes, in the order worth checking - **A misspelled or missing credential token.** Substitution is a literal string replace, so `{%user%}` is simply not found and is posted as itself. The method's configuration check does not look for the tokens, so this is a fully valid configuration. - **The wrong method for the body.** Form and JSON are the same base class with different encoders; a JSON body configured as `form` has its credential URL-encoded into it. - **A `loginRequestUrl` that is not the form's real action.** Very common when the URL was copied from the browser address bar rather than from the request. - **An anti-CSRF token the tool did not recognise**, so the login post carries a stale one. The copy only happens when the pre-flight response actually contained a token the extension knows. - **A login the application will not accept from a non-browser client** at all - a signature computed in page JavaScript, or a step the replay does not reproduce. - **A credential variable that never resolved.** An unresolved `${...}` is warned about and then left in the string as literal text, so the login is sent with the placeholder itself as the username - and it is sitting in plain sight in the recorded request. ## What you see, and what it means | what the recorded login request shows | what it is telling you | |---|---| | a credential token still present in the body | the token was misspelled, so the replace found nothing | | a placeholder variable in the credential field | a plan variable never resolved; a warning was recorded | | a credential encoded for the wrong body format | the method does not match the body you configured | | only one recorded message, not two | the request could not be prepared or could not be sent | | a well-formed request and a 200 carrying the form | the application simply rejected the credentials | ## How to work it out, concretely 1. Open the two recorded authentication messages and read the **request body that was actually sent**. If you can see `{%username%}` or a misspelling of it in the bytes, you are done. 2. Compare that request, field for field, against a real login captured from a browser. The difference is usually one parameter, not the whole shape. 3. Read the **response** to the login request. A 200 carrying the login form again is a rejected login; a redirect to a dashboard is an accepted one. 4. Check whether the scanning job actually named a `user`. A job with no user runs anonymously no matter how correct the context's method is. 5. Only then look at whether the tool **noticed** - whether it was logged out again is a separate verification concern, configured separately, and it is the wrong place to start. ## Making the failure loud next time The structural fix is to stop treating "the plan completed" as the success signal. Decide, before the run, one request or one finding that can only exist behind the login, and assert on it - whether that is an authenticated URL appearing in the crawl or a count of pages reached. A run whose authenticated coverage is zero should be a red pipeline, and it will not be one by default, because nothing in the login path raises an error when it is simply ignored.

  • Why is the pre-flight GET there, and what breaks if the login page URL is wrong?
    It renews cookies and, when the response carries a token the anti-CSRF extension recognises, picks up a fresh value that is copied into the login request. If the parameter is blank the tool fetches the login request URL instead, which for a POST-only endpoint yields nothing useful; if it points at the wrong page, the login posts a stale or missing token and the application rejects it.
  • The credentials came from plan variables that were never set in the runner. What does that look like?
    The environment warns about the unknown variable and then leaves the literal placeholder text in place, so the login goes out with that text as the credential. It is among the easiest of these failures to spot, because the placeholder sits in plain sight in the recorded login request - provided you read the request rather than the configuration.
  • How would you make this failure fail the pipeline next time?
    Assert on something that can only exist behind the login - an authenticated URL appearing in the crawl, or a minimum count of pages reached - rather than treating plan completion as success. Nothing in the login path raises an error when it is simply ignored, so the assertion has to be yours.

saying these in an interview costs you the question

  • Treats a clean exit code as proof the scan was authenticated
  • Assumes a rejected login is recorded as an authentication error
  • Starts by rewriting the login body instead of reading the sent one
  • Forgets the scanning job must name a user to run logged in
  • Believes one login attempt is a single HTTP request