What request forgery does an authorization server or client invite when it fetches a URL another party supplied?
answer
- somebody else picks the destination
- the request starts inside your network
- discovery turns a string into a fetch
- resolve the address before you connect
- no redirects, no echoed body
basics
~20 sA URL a party hands you and your server then dereferences becomes a request made from inside your network, at a destination someone else chose. Aimed at an internal-only address it reaches services no external caller can.
solid answer
~50 sOAuth deployments dereference URLs that did not come from your configuration: an issuer identifier turned into the location of an authorization-server metadata document, the `jwks_uri` found inside such a document, URLs submitted when a client record is created. Each fetch is an outbound HTTP request your server makes on someone else's instruction, and it originates inside your network rather than outside it. An attacker who controls the value can aim it at addresses nobody outside can reach, at internal endpoints that act on a plain GET, or at a host that answers differently depending on what is listening - and echoing the response body or the upstream error back turns probing into exfiltration. Treat the URL as untrusted input at the point of fetch: allowlist which issuers you will dereference, require `https`, resolve the name and reject internal addresses before connecting, refuse redirects, bound size and time, and never surface the body or the upstream error.
code
pseudocode · 16 linesfunction fetch_metadata(url, allowed_issuers):
if url not in allowed_issuers:
reject "issuer not allowed"
if scheme of url is not "https":
reject "scheme not allowed"
address = resolve_once(host of url)
if address is loopback, private or link-local:
reject "internal address"
response = http_get(address,
host_header = host of url,
follow_redirects = false,
max_bytes = 65536,
timeout_seconds = 2)
if response.status is not 200:
reject "unexpected status"
return response.bodygo deeper
Know that a server fetching a URL someone else supplied is making a request from inside your network to a destination you did not choose.
Explain the concrete fetch points a discovery-driven deployment has, and why an internal-only service is reachable through them when it is not reachable from outside.
Show the guarded fetch in detail - allowlist, scheme, resolve-then-connect, no redirects, bounded size and time, nothing echoed - and say what each control leaves open.
Decide how much your platform discovers at all: dynamic dereferencing is a network-reach primitive, and for a handful of known partners static configuration removes the class instead of managing it.
## Where an OAuth deployment fetches a URL A deployment that configures every endpoint statically makes no such fetches. Discovery-driven ones make several: - an **issuer identifier**, from configuration or from a party, converted into the location of an authorization-server metadata document at the `/.well-known/oauth-authorization-server` path; - the **`jwks_uri`** found inside that document, fetched to obtain verification keys and re-fetched on rotation; - **URLs submitted when a client record is created**, which a server may dereference to display or validate something; - health checks and admin tooling that will happily fetch whatever an operator pastes in. The common shape is: *a value arrives from outside, and a server inside your network turns it into an HTTP request*. ## Why that fetch is the attacker's request An external attacker's own requests arrive at your edge and are filtered there. A request your server makes on their behalf does not. It starts inside, carries your source address, and may pass through no filter at all. That single change of origin is the whole vulnerability, and it is why input validation phrased as "escape it" misses the point - nothing is being injected into a string. A destination is being chosen. What the attacker gets, in rough order of severity: - **Reach** to services that are only listening internally and were never designed to be exposed. - **Effect**, where an internal endpoint performs an action on a plain GET. - **An oracle**, because a connection refused, a timeout and a well-formed error take different times and produce different failures, which maps your internal estate one probe at a time. - **Exfiltration**, if the fetched body, or the upstream error containing it, is returned to the caller or written to a log the caller can read. - **Amplification**, if the fetch happens per request rather than once and cached, so an attacker can make your server flood a third party. ## Controls, and what each actually closes | Control | What it closes | What it leaves | |---|---|---| | Allowlist of issuers you will dereference | Arbitrary destinations | Nothing, if the list is configuration rather than discovery | | Require the `https` scheme | Redirection to other protocols and cleartext fetches | Internal hosts that speak HTTPS | | Resolve the name and reject internal addresses | Internal reach by name or literal | A name that resolves differently on the second lookup | | Refuse redirects | A validated host handing you off to an internal one | Nothing, provided you refuse rather than re-validate loosely | | Cap response size and time | Amplification and resource exhaustion | Small, fast probes | | Never echo body or upstream error | Exfiltration through the response | Timing as a weaker oracle | Two of those rows deserve emphasis. **Refusing redirects** matters because a host that passes your check can answer with a redirect to one that would not; validating the first hop and then following whatever comes back is the same as not validating. And **re-resolution** matters because the address you checked and the address the connection uses can differ if the name is looked up twice - the safe pattern is to resolve once, validate the address, and connect to that address. ## A guarded fetch ```pseudocode function fetch_metadata(url, allowed_issuers): if url not in allowed_issuers: reject "issuer not allowed" if scheme of url is not "https": reject "scheme not allowed" address = resolve_once(host of url) if address is loopback, private or link-local: reject "internal address" response = http_get(address, host_header = host of url, follow_redirects = false, max_bytes = 65536, timeout_seconds = 2) if response.status is not 200: reject "unexpected status" return response.body ``` ## The judgement behind it The cheapest control is the one people skip: **do not make the fetch dynamic at all**. For a partner integration whose issuer you agreed in a contract, the endpoints belong in configuration, refreshed deliberately, with discovery used to detect drift rather than to decide where to connect. Dynamic dereferencing earns its keep in a deployment that genuinely meets parties it has never met; in one that has three known partners, it is a network-reach primitive handed to whoever can influence a string, in exchange for convenience nobody asked for.
- Why is refusing redirects on such a fetch more important than it first appears?Because the host you validated can answer with a redirect to one you would have rejected. If the client follows it, the check applied to the first hop bought nothing. Refusing redirects outright is simpler and safer than re-validating each hop, and a metadata document that requires a redirect to retrieve is a misconfiguration worth surfacing anyway.
- The host allowlist passed, but the name resolves to an internal address. What went wrong?Name-based allowlisting says nothing about where the name points. Resolve the name, validate the resulting address against your internal ranges, and connect to that address with the original host header - otherwise a second lookup can return a different answer to the one you checked.
- How do you keep the fetch from becoming a denial-of-service lever?Do not fetch per request. Cache the document and the key set with a sensible lifetime, refresh on a schedule or on a verification failure, cap response size and connect timeout, and serve the last known-good copy when a refresh fails rather than failing every request behind it.
- When should a deployment not dereference discovery URLs at all?When it deals with a known, small set of parties. Put the endpoints in configuration, change them deliberately, and use discovery only to detect that something drifted. Dynamic dereferencing is worth its risk when you genuinely meet parties you have never met before.
saying these in an interview costs you the question
- Calls it harmless because the server only performs a GET
- Allowlists the hostname but never checks the resolved address
- Validates the first hop and then follows redirects
- Returns the fetched body or upstream error to the caller
- Fetches the document on every request with no cache or timeout
- Thinks escaping the URL string is the relevant defence