skip to content

A scheduler mails technicians a link carrying the access token in the query string — which copies can you never recall?

level: middleimportance: must knowfreq 60%

answer

  1. the URL is the copied part
  2. logs, Referer, caches, history, previews
  3. no-store covers stored responses only
  4. specification says SHOULD NOT, not MUST NOT

basics

~20 s

A credential in a URL is copied into access logs at every hop, into the referring-page header on same-origin asset requests, into shared caches, and into history, bookmarks and pasted-link previews. The log copies are the ones nobody can recall.

solid answer

~50 s

A URL is the most widely copied part of an HTTP request, and every copy is made by software you did not write. The request line is logged verbatim by your server, by any reverse proxy and by any content delivery network in the path, then shipped into an index with a retention policy measured in weeks. The browser repeats the full URL in the `Referer` of same-origin asset requests. A shared cache may store the response keyed on it — that one, and only that one, is what `Cache-Control: no-store` addresses. The user's history, bookmarks and pasted links carry it too, and a chat system building a link preview will fetch the URL using the credential. The bearer-token specification says query-parameter carriage SHOULD NOT be used unless the header and the body are both impossible.

code

http · 11 lines
http
GET /shifts?access_token=eyJhbGciOi...Qssw5c HTTP/1.1
Host: scheduler.example
Accept: text/html

HTTP/1.1 200 OK
Content-Type: text/html
Cache-Control: no-store

GET /assets/turbine-map.js HTTP/1.1
Host: scheduler.example
Referer: https://scheduler.example/shifts?access_token=eyJhbGciOi...Qssw5c

go deeper

for a junior

Remember the shape: a URL gets copied into places you do not control, and encryption does not help because the copies are made at the ends, not on the wire. Naming logs and browser history is enough at this level.

for a middle

Enumerate the channels and say who creates each copy. Be exact about what Cache-Control: no-store covers, and quote the specification's SHOULD NOT at its real strength rather than upgrading it.

for a senior

Treat the credential as disclosed and go looking for the retention window, because that window is the real duration of the incident. Then design the case that forced the URL: a short single-use handle the landing page exchanges immediately.

for a principal

Decide the standing rule — no long-lived credential in a request line, anywhere, with a named exception process — and make it enforceable at the edge rather than in review, because this defect is reintroduced by whoever next needs a mailable link.

## A URL is not a private place A credential in the query string is not protected by the fact that the connection is encrypted. Transport encryption hides the request from anyone watching the wire between two hops; it does nothing about the hops themselves. And a URL is the single most widely **copied** part of an HTTP request: the request line is what proxies log, what caches key on, what browsers remember, what users paste, and what pages repeat. Every one of those copies is created by software nobody on your team wrote and cannot reach afterwards. The turbine maintenance scheduler mails technicians a deep link to a shift board, and someone puts the access token in the query string because a mail client cannot set a request header. That is exactly the situation the trade-off exists for, so the honest thing is to enumerate the channels rather than wave at "don't do that". ## The channels, one at a time | channel | who creates the copy | can you recall it? | |---|---|---| | access logs at every hop | your server, a reverse proxy, a content delivery network | no — logs are shipped, indexed and retained | | the `Referer` on same-origin subresource requests | the browser | no — repeated on every asset the page loads | | a shared intermediary cache | the cache, keyed on the full URL | only by preventing storage in the first place | | history and bookmarks | the browser, at the user's request | only the user can, and only on that device | | pasted links and their previews | chat and ticketing systems fetching the URL | no — and the fetch itself uses the credential | Two of those deserve a closer look. **Logs are the channel that actually causes incidents.** Not because anyone is careless, but because the request line is logged *by default, everywhere*, long before your own code runs. The token lands in a file on a proxy, in a log-shipping pipeline, in a searchable index with a retention policy measured in weeks, readable by everybody with support access. A redaction rule added afterwards removes nothing that was already written. **The referring-page channel is narrowed, not closed.** Browsers changed their default so that a request to a *different* origin carries only the origin, not the path or query — so a query-string credential no longer routinely walks out to an unrelated host. But a request back to your own origin still carries the full URL, so every script, image and font the page loads repeats the credential in a header, and each of those requests is itself logged. Treat the channel as reduced, and do not tell an interviewer it is gone. ## What the specification actually says The OAuth 2.0 bearer-token specification, RFC 6750, defines three ways to present an access token: the `Authorization` request header with the `Bearer` scheme, a form-encoded body parameter, and a URI query parameter named `access_token`. It does not forbid the third. It says it **SHOULD NOT** be used unless it is impossible to carry the token in the header or in the body, and it gives the reason: the high likelihood that the URL will be logged. Quote that strength exactly. "SHOULD NOT with a narrow escape hatch" is a different claim from "MUST NOT", and upgrading it is the kind of small inaccuracy that costs a candidate credibility. The escape hatch exists because some clients genuinely cannot set a header — and that is the case you are in when the link arrives by mail. ## The fix, and what it costs 1. **Move the credential to the header** wherever a client can set one. This removes the URL channel completely: headers are not part of the cache key, not repeated in the referring-page header, not bookmarked, and not in the request line a proxy logs by default. 2. **Where a header is impossible, move it to the body.** A form-encoded parameter is not logged by default and not cached, though it forces the request to stop being a `GET`. 3. **Where neither is possible**, do not send the long-lived credential at all. Send a single-use, short-lived, narrowly scoped handle that the landing page exchanges once and then removes from the address bar. The thing that leaks is then worth much less, for a much shorter time. `Cache-Control: no-store` on the response belongs in all three cases, but be precise about what it does: it stops a cache storing the *response*. It does not touch the request line already in a log, the browser's history entry, or the header the page sends on its next asset request. ## The copies that already exist If a credential has been in a URL, treat it as disclosed. Not "probably fine because the link expired" — disclosed. Replace it, and then find out how long your log retention keeps the evidence of the exposure, because that retention window is the real duration of the incident, not the token's lifetime. A short expiry shrinks the window in which the leaked copy is useful; it does nothing about the copy existing. ## What to say when you are asked List the channels, say which of them you created and which the platform created for you, and then name the one you cannot recall: the log. Finish by quoting the specification at its actual strength and saying what you would do for the case that forced the URL in the first place.

  • RFC 6750 allows query-string carriage at all — under what condition?
    It defines a URI query-parameter method using the name `access_token`, then says it SHOULD NOT be used unless it is impossible to carry the token in the `Authorization` header or in the request body. That is a strong discouragement with a narrow escape hatch, not a prohibition, and the hatch exists for clients that genuinely cannot set a header.
  • Does moving the credential to the request body instead of the query string remove the leakage?
    It removes the URL channel: bodies are not in the request line a proxy logs, not part of a cache key, not sent in the referring-page header and not bookmarked. It does not remove logging you configured yourself, and it forces the request to stop being a `GET`, which is sometimes the reason the URL was chosen in the first place.
  • The technician's link is single-use and expires in sixty seconds. Is URL carriage fine now?
    Better, not fine. The window in which a leaked copy is useful shrinks to a minute, which is a genuine improvement. The copies still exist afterwards as a durable record that a credential was exposed, and anyone with log access during that minute can use it. Prefer a short handle the landing page exchanges once over the real access token.

saying these in an interview costs you the question

  • Saying HTTPS makes a token in the query string safe
  • Claiming a short expiry makes URL carriage acceptable
  • Thinking a no-store response header keeps the URL out of logs
  • Assuming the specification forbids query-string carriage outright
  • Redacting one log line and calling the leak closed
  • Believing the referring-page channel is fully closed by modern browsers