skip to content

Why does a Content-Security-Policy violation report name the original request URL rather than the redirect target that was actually refused?

level: seniorimportance: nice to knowfreq 27%

answer

  1. a report is an outbound channel
  2. would otherwise probe third-party endpoints
  3. pre-redirect URL, deliberately
  4. fragment and credentials removed, query kept
  5. non-HTTP(S) reduced to the scheme alone

basics

~20 s

Because the report would otherwise be an oracle: a page could point a request at a third-party endpoint and learn from its own violation report where that endpoint redirects. Reporting the pre-redirect URL keeps the record free of information the page could not already see.

solid answer

~50 s

A violation report is a channel that carries information out of the browser to a destination the page chose. Anything the report body reveals, the page's author learns. If `blocked-uri` named the post-redirect location, a page could aim a deliberately refused request at somebody else's endpoint and read the redirect target out of its own report, which is information the same-origin rules would never have handed it. So the specification reports the **pre-redirect** URL. The same reasoning drives the stripping rules: fragment, username and password are removed from reported URLs, and a URL whose scheme is not HTTP(S) is reduced to its scheme alone. Note what is *not* stripped: the query string survives, so a lookup URL carrying a prescription identifier arrives at your collector intact. And reports are unauthenticated POSTs, so anything in the feed is attacker-supplied data.

code

http · 5 lines
http
GET /widget.js HTTP/1.1
Host: metrics.example

HTTP/1.1 302 Found
Location: https://internal.metrics.example/t/8841/profile#session-9932

go deeper

for a junior

The point to carry away is that a violation record is deliberately vague in places. It describes what the page asked for, not where that request ended up, and it leaves some parts of a URL out on purpose.

for a middle

Explain the mechanism: the record is delivered to a destination the page chose, so anything in it is something the page's author learns. Then name the three rules, pre-redirect URL, removed fragment and credentials, scheme-only for non-HTTP(S).

for a senior

Show both directions. Outbound, the record is constrained so it cannot be used to probe somebody else's endpoint; inbound, the collector is an unauthenticated POST target whose contents are attacker-supplied and must be validated before anything is concluded from them.

for a principal

The call to own is where reports go and what the URLs in them contain. A feed of page URLs is a data flow, and choosing a destination for it is a retention and third-party decision, not a debugging convenience.

## A report is an outbound channel, not a log line It is tempting to think of a violation report as internal diagnostics that happen to travel over the network. It is not. The report is generated inside a browser that holds data from many origins, and it is delivered to a destination named by the page's own policy. Every field the specification allows into the body is a field the page's author gets to see. The design question behind the whole report format is therefore: *what can this record say without telling the page something it was not entitled to know?* That single question explains every rule below. ## The redirect rule When a resource is refused after the request has been redirected, the record names the URL that was **originally requested**, not where the redirect pointed. Without that rule, a page could: 1. Ship a policy that refuses a particular host. 2. Deliberately request a third-party endpoint on that host, one that redirects based on whether a visitor is signed in, or which tenant they belong to. 3. Read its own violation report and recover the redirect target. Step 3 is a cross-origin read dressed as a diagnostic, and it would work on any endpoint in the world, from any page, with no cooperation from the endpoint. Reporting the pre-redirect URL removes the channel entirely: the page learns only the URL it already wrote down itself. ## The stripping rules The same reasoning is applied to the URLs in the record before they are serialised: - The **fragment** is removed. A fragment never reaches a server, and treating it as reportable would make it leak to the collector when nothing else does. - The **username and password** components are removed, because credentials embedded in a URL are the one part of it that is unambiguously a secret. - A URL whose scheme is **not HTTP(S)** is reduced to **its scheme alone**, so a record can say that something with a non-network scheme was refused without describing what it was. | Part of the URL | Present in the record | |---|---| | scheme, host, port | yes | | path | yes | | query string | **yes** | | fragment | no, removed | | username, password | no, removed | | whole URL if the scheme is not HTTP(S) | no, scheme only | ## The half that surprises people: the query survives The stripping rules are a privacy floor against *other people's* data, not against your own. A prescription-status lookup served at `/status?rx=8841` reports `document-uri` with that query intact, on every violation, to whatever collector the policy names. If the collector is a third party, the identifier is now theirs too. If it is internal, it is now in a log store that was probably never assessed for that class of data. So the operational rule is: - Treat the reporting destination as a system that receives page URLs, and choose it accordingly. - Keep identifying detail out of URL paths and queries on pages that report, or accept that the collector inherits it. - Remember that the same URL travels in the `url` member of the delivery envelope as well as in the body. ## The other direction: what arrives is untrusted The rules above constrain what a browser *puts into* a report. Nothing constrains what somebody else *sends to* your collector. The endpoint is an unauthenticated POST target on the public internet: - Anyone can post a well-formed report describing a violation that never happened. - Anyone can post a badly-formed one, or a very large one, to see what your parser does. - A real browser can be induced to report against your collector from a page you do not control, if that page names your endpoint. Therefore the feed is **input**, with all that implies: validate the shape, cap the size, rate-limit per source, never interpolate a field into a query or a rendered page unescaped, and never treat a count of reports as a measurement of anything until you know where they came from. A record that says a policy was violated is a claim, and the collector cannot authenticate the claimant. ## The summary a senior candidate gives The report format is deliberately less informative than it could be, in exactly the places where being more informative would leak somebody else's data — and deliberately silent about the one place where it carries your own, the query string. Both halves matter when you decide where reports go and what you are willing to conclude from them.

  • What follows for how you treat the reports arriving at your collector?
    Treat them as untrusted input. The endpoint is an unauthenticated POST target, so anyone can send a well-formed record describing a violation that never happened, or a malformed one aimed at your parser. Validate the shape, cap the body size, rate-limit, and never render or interpolate a field without escaping it. A report is a claim you cannot authenticate.
  • Which part of a reported URL is most likely to carry data you did not intend to send?
    The query string, because it is not stripped. Fragment, username and password are removed, but a path and query go to the collector verbatim, in both the body and the delivery envelope's `url` member. On a lookup page whose URL carries an identifier, every violation sends that identifier to whoever runs the collector.
  • Why is a non-HTTP(S) URL reduced to just its scheme?
    Because such URLs frequently describe something local or opaque to the page's author, and the record only needs to say what class of thing was refused. Reporting the scheme alone keeps the record useful for deciding whether a directive needs to cover that scheme, without describing the specific target at all.

saying these in an interview costs you the question

  • Assumes the record names the final redirect target
  • Treats the collector's feed as authenticated, trustworthy evidence
  • Believes a fragment identifier reaches the collector
  • Thinks a non-HTTP(S) URL is reported in full
  • Assumes the query string is stripped along with the fragment
  • Reads the stripping as a client quirk rather than a specified rule