skip to content

How do you stop urllib.request.urlopen from following a redirect?

level: middleimportance: should knowfreq 35%

answer

  1. urlopen is a facade, not the client
  2. A chain of handlers does the work
  3. Override one method to decline
  4. build_opener plus a handler subclass
  5. redirect_request returning None raises HTTPError

basics

~10 s

Subclass urllib.request.HTTPRedirectHandler so its redirect_request returns None, pass the subclass to build_opener, and call open on that opener. With no handler willing to redirect, the 3xx surfaces as an HTTPError carrying the Location header.

solid answer

~50 s

`urlopen` runs on a default opener whose handler chain includes `urllib.request.HTTPRedirectHandler`, which follows 301, 302, 303, 307 and 308 transparently — up to ten redirects in a chain, after which it raises `HTTPError` rather than looping. To stop it, subclass that handler with a `redirect_request` that returns `None`, hand the subclass to `urllib.request.build_opener`, and use `open` on the resulting opener: with nothing willing to redirect, the 3xx falls through to the default error handler and is raised as an `HTTPError` whose `Location` header you can read. Two behaviours are worth knowing while you are there. A POST redirected by 301, 302 or 303 is re-issued as a GET with the body and its content headers dropped, and urllib refuses to replay a POST body on 307 or 308 — it raises `HTTPError` instead of following.

code

python · 26 lines
python
import urllib.error
import urllib.request

handler = urllib.request.HTTPRedirectHandler()
posted = urllib.request.Request(
    "http://metrics.invalid/start",
    data=b"payload=1",
    headers={"X-Scrape-Id": "42", "Content-Type": "text/plain"},
)

followed = handler.redirect_request(posted, None, 302, "Found", {}, "http://metrics.invalid/final")
print(followed.get_method(), followed.data, followed.header_items())

try:
    handler.redirect_request(posted, None, 307, "Temporary Redirect", {}, "http://metrics.invalid/final")
except urllib.error.HTTPError as exc:
    print("307 refused rather than replayed:", exc.code)


class NoRedirect(urllib.request.HTTPRedirectHandler):
    def redirect_request(self, req, fp, code, msg, headers, newurl):
        return None


opener = urllib.request.build_opener(NoRedirect)
print([type(h).__name__ for h in opener.handlers if "Redirect" in type(h).__name__])

go deeper

for a junior

Know that urlopen follows redirects for you, so the object you get back is the final response and its status is the final status, not the 3xx that started the chain.

for a middle

Explain the opener and handler chain: what build_opener assembles, and how overriding redirect_request on a HTTPRedirectHandler subclass turns a redirect into a raised HTTPError you can inspect.

for a senior

Show judgement about when following must be refused — a body that would be dropped, a header that would travel to another host, a hop you need recorded — and how you re-issue the request deliberately instead.

for a principal

Own where redirect policy lives: one wrapped client with an agreed policy and an audit trail, versus every team subclassing the handler locally and drifting apart, and never a process-wide installed opener hiding the change from readers.

## urlopen is a facade over a handler chain `urllib.request.urlopen` is not the client; it is a convenience wrapper around a default `OpenerDirector` assembled once and cached. That opener holds an ordered chain of handlers — one that knows how to open `http` URLs, one that knows `https`, one that raises for error statuses, one that manages redirects. Anything you want to change about client policy, you change by building your own opener with `urllib.request.build_opener`, which starts from the default chain and lets your handler classes replace the built-in ones of the same type. That is the mental model the question is really testing. There is no `allow_redirects=False` keyword anywhere in `urllib.request`, and there never has been. Policy lives in handlers. ## What the redirect handler does by default `urllib.request.HTTPRedirectHandler` intercepts 301, 302, 303, 307 and 308 responses. For each, it builds a new `Request` for the `Location` target and re-enters the opener, transparently, so the caller sees only the final response. Its class attributes cap the process: `max_redirections` stops a chain at ten hops, and `max_repeats` stops it after four visits to the same URL. Exceed either and it raises `HTTPError` complaining about a redirect loop. Two rewrite rules matter in practice: * **A POST redirected by 301, 302 or 303 becomes a GET.** The handler builds the follow-up request with the method rewritten and the body dropped, and it strips the `Content-Type` and `Content-Length` headers on the way. Your carefully constructed POST arrives at the destination as a bodyless GET. * **A POST is never replayed on 307 or 308.** Rather than re-sending the body, urllib raises `HTTPError` with that status. This surprises people who expect the client to honour the 'repeat the request as-is' intent; the standard library simply declines and hands the decision back to you. Headers you attached with `add_header` — including everything in the `headers` dict you passed to `Request` — *are* copied onto the follow-up request, even when the redirect points at a different host. Only the two content headers are dropped. Headers attached with `add_unredirected_header` are not copied; that method exists precisely so a header is not replayed, and urllib uses it for the headers it fills in itself. If a header must not travel to an unexpected destination, refusing the redirect and re-issuing the request deliberately is the mechanism the library gives you. ## Refusing to follow `redirect_request` is the extension point. It is documented as returning a `Request` to follow, raising `HTTPError` to stop everything, or returning `None` to mean 'I decline, let another handler try'. Since no other handler in the chain wants a 3xx, returning `None` means the response falls through to `HTTPDefaultErrorHandler`, which raises `HTTPError`. The exception then carries the status and the response headers, so `Location` is readable in the handler. ```python import urllib.request class NoRedirect(urllib.request.HTTPRedirectHandler): def redirect_request(self, req, fp, code, msg, headers, newurl): return None opener = urllib.request.build_opener(NoRedirect) ``` The same seam supports the other things people want: logging every hop, capping the chain lower than ten, refusing a redirect that leaves the original host, or refusing one that would drop your body. Override `redirect_request`, decide, and either return a `Request` or refuse. Use `open` on your opener rather than `urlopen`, which keeps the policy local to the call site. `urllib.request.install_opener` makes your opener the process-wide default that `urlopen` uses — convenient in a small script, and a hidden global in anything larger, because a later reader has no way to tell from the `urlopen` call that redirect policy has changed. ## Knowing where you landed When a redirect *is* followed, the object `urlopen` returns has its `url` attribute set to the URL of the last request urllib actually made, so it is the final destination rather than the URL you asked for. Its `status` is the final response's status: a followed chain looks like a plain 200, and nothing on the returned object records the hops. If the chain matters — for an audit trail, or because a redirect is itself the signal — you must refuse the redirect and walk it yourself. ## What an interviewer is listening for The shape of the answer is: urlopen is a facade, redirects are handled by a handler in the opener's chain, you change behaviour by subclassing the handler and building your own opener, and refusing means returning `None` and catching the resulting `HTTPError`. The bonus marks come from knowing that a POST silently becomes a GET on 301, 302 and 303, that 307 and 308 raise instead, and that the final URL is on the response object rather than anywhere in your own code.

  • How many redirects does urllib follow before it gives up?
    The handler stops at ten redirections in one chain and at four visits to the same URL, governed by its `max_redirections` and `max_repeats` class attributes, then raises `HTTPError` about a redirect loop. Raising the ceiling means subclassing the handler and changing those attributes. There is no callback reporting the hops — if you need the chain, refuse the redirect and follow it yourself.
  • What happens to headers you set when urllib follows a redirect to a different host?
    Headers attached with `add_header`, including everything in the `headers` dict passed to `Request`, are copied onto the new request even when the destination host differs; only `Content-Type` and `Content-Length` are dropped. Headers attached with `add_unredirected_header` are not copied — that is exactly what the method is for. For a header that must not travel, either attach it as unredirected or refuse the redirect and re-issue the request deliberately.
  • After a followed redirect, how do you tell where the request actually landed?
    Read the `url` attribute on the object `urlopen` returns; urllib sets it to the full URL of the last request it made, so after a chain it is the final destination rather than what you asked for. The `status` attribute is the final response's status, so a followed redirect simply looks like a 200 and the intermediate hops leave no trace on the returned object.

saying these in an interview costs you the question

  • Assuming urlopen returns the 3xx response itself
  • Expecting a redirects or allow_redirects keyword on urlopen
  • Thinking urllib replays the POST body on every redirect
  • Believing redirect chains loop forever without a limit
  • Reporting the requested URL as the final URL
  • Editing the module-level default opener inside library code

context