skip to content

urllib and http.client

Fetching a URL with nothing installed means urlopen, a Request carrying headers and an explicit timeout, and telling an HTTP status error apart from a connection failure. Interviewers probe redirects.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What is the difference between urllib.error.HTTPError and urllib.error.URLError?

level: middleimportance: must knowfreq 55%

answer

  1. One is a subclass of the other
  2. Answered badly versus never answered
  3. The error object is also a response
  4. Order except clauses most specific first
  5. code, reason, headers and read()

basics

~20 s

HTTPError means the server answered with a status urllib treats as an error; URLError means no HTTP answer arrived at all, such as a name-resolution or connection failure. HTTPError is a subclass of URLError, so catch it first.

solid answer

~40 s

`urllib.error.URLError` subclasses `OSError` and covers everything that stopped the exchange before a status line came back — name resolution failure, refused connection, a TLS handshake error — with the underlying exception stored in its `reason` attribute. `urllib.error.HTTPError` subclasses `URLError` and means the opposite: the server did answer, with a status urllib treats as an error. Because it also inherits from urllib's response wrapper, it doubles as the response — `code`, `reason`, `headers` and `read()` all work on it, which is how you get the error body. Since it is a `URLError`, an `except urllib.error.URLError` clause placed first swallows it, so order handlers most specific first. One trap: a timeout while waiting for the response escapes `urlopen` as a bare `TimeoutError`, not wrapped in `URLError`, so a catch-all handler wants `OSError`.

code

python · 6 lines
python
import urllib.error

print(issubclass(urllib.error.HTTPError, urllib.error.URLError))  # True
print(issubclass(urllib.error.URLError, OSError))                 # True
print(issubclass(TimeoutError, OSError))                          # True
print(issubclass(TimeoutError, urllib.error.URLError))            # False

go deeper

for a junior

Know that urlopen raises rather than returning a 404 response, and that the exception to catch for a bad status is urllib.error.HTTPError while a dead host gives urllib.error.URLError.

for a middle

Explain the inheritance chain — HTTPError under URLError under OSError — and why the order of your except clauses decides which one you actually catch.

for a senior

Demonstrate pulling the error body off the HTTPError for diagnostics, separating retryable transport failures from server answers, and knowing that a stalled read arrives as a bare TimeoutError.

for a principal

Own the failure taxonomy: which classes retry with backoff, which page a human, and how you stop urllib exception types leaking across module boundaries into every caller's error handling.

## Two failures that feel the same and are not When `urllib.request.urlopen` does not give you a response, exactly one of two things happened. Either the conversation never got far enough for the server to say anything — DNS did not resolve, the TCP connection was refused or reset, the TLS handshake failed — or the server answered perfectly well and the answer was a status urllib treats as an error. urllib gives these different exception types, and the whole point of the distinction is that they call for different reactions: the first is a transport problem you might retry or route around, the second is an answer you must read. ## The hierarchy ``` OSError └─ urllib.error.URLError └─ urllib.error.HTTPError (also a readable response object) ``` `URLError` subclasses `OSError`, so any handler that already catches `OSError` catches it. `HTTPError` subclasses `URLError`. That single inheritance edge is the source of the most-asked question on this topic: because a `HTTPError` *is a* `URLError`, writing the general clause first silently absorbs the specific one. ```python try: resp = urllib.request.urlopen(url, timeout=5) except urllib.error.HTTPError as exc: # must come first ... except urllib.error.URLError as exc: # transport failure ... ``` Reverse those two clauses and you can no longer tell a 503 from a refused connection, because both land in the same branch and only one of them has a status code. ## HTTPError is also the response `HTTPError` is unusual: besides subclassing `URLError`, it inherits from urllib's response wrapper, which makes the exception object a file-like response. Inside the handler you can use the `code` attribute for the numeric status, `reason` for the status text, `headers` for the response headers, and `read()` for the body bytes. That last one matters more than it sounds: an API that explains *why* it rejected you does so in the body of the 4xx, and the only place that body exists is on the exception instance. Read it inside the `except` block — once the object is discarded, the socket goes with it. The `reason` attribute means different things on the two classes. On `HTTPError` it is the status phrase the server sent. On `URLError` it is usually the underlying exception object — a `socket.gaierror` from a failed lookup, a `ConnectionRefusedError`, an SSL error — so log its `repr` or check its type rather than matching on message text, which varies by platform and resolver. ## Which statuses raise at all urllib raises `HTTPError` for statuses outside the 2xx range, with one large exception: redirect statuses are consumed by the redirect handler and followed rather than raised, so a followed redirect looks to you like the final 200. If the redirect chain hits urllib's limit, or the handler declines to follow it, then it surfaces as an `HTTPError` with the 3xx code. This is why there is no `raise_for_status()` idiom in the standard library. The raising happens automatically, and the interesting work is deciding, in the handler, which codes you treat as fatal, which as retryable, and which as an expected answer that happens to be non-2xx. ## The timeout hole Here is the part that separates someone who has read the docs from someone who has debugged this. urllib wraps the errors raised while it connects and sends the request into a `URLError`. It does not wrap what happens next. A timeout that fires while waiting for the response status line propagates out of `urlopen` as a plain `TimeoutError`, and `TimeoutError` and `URLError` are siblings under `OSError` — neither one catches the other. Code that carefully handles `HTTPError` and `URLError` and nothing else will crash on the single most common production failure mode. Two defensible fixes: catch `OSError` as the outermost transport clause (it covers `URLError`, `TimeoutError` and anything else the socket layer throws), or name `TimeoutError` explicitly alongside `URLError` so the retry path is obvious to a reader. Since Python 3.10, `socket.timeout` is an alias of the builtin `TimeoutError`, so there is only one name to catch. ## What to say in an interview State the inheritance chain, state that `HTTPError` is simultaneously an exception and a response, state the clause ordering rule, then show judgement: a 5xx and a refused connection may both be retryable, but only one of them tells you the service is alive; a 4xx is almost never worth retrying and its body is the diagnostic. Mentioning the bare `TimeoutError` on the read path is the detail that marks real operational experience.

  • How do you read the body of a failed request when urlopen raises?
    The `HTTPError` instance is the response: calling `read()` on it returns the body bytes, its `headers` attribute holds the response headers, and `code` the status. Do it inside the handler, before the object is discarded, because the underlying socket is closed when the object goes away. A `URLError` has no body — there is nothing to read when the server never answered.
  • Which exception does a name-resolution failure raise, and what does its reason attribute hold?
    `urllib.error.URLError`, whose `reason` is the `socket.gaierror` the resolver raised. On `URLError` the `reason` is usually the underlying exception object rather than a string, so log its `repr` or inspect its type instead of matching on message text — the wording differs by platform and resolver.
  • Why is catching OSError around urlopen sometimes the right call?
    `URLError` and `TimeoutError` are both `OSError` subclasses but siblings, so no single urllib exception covers both. urllib wraps failures raised while connecting and sending into `URLError`, but a timeout that fires while waiting for the response propagates as a plain `TimeoutError`. Catching `OSError` covers every transport failure; keep a `HTTPError` clause above it when a server answer deserves different treatment.

URLError is the letter coming back stamped undeliverable; HTTPError is a reply that says no — and you can still open that envelope and read the reason inside.

saying these in an interview costs you the question

  • Catching URLError before HTTPError and losing the status
  • Believing a 404 returns a response instead of raising
  • Thinking HTTPError carries no readable response body
  • Assuming every network failure surfaces as URLError
  • Treating the reason attribute as a plain message string
  • Expecting a raise_for_status style call in the stdlib

context

open as a page

Using urllib.request, how do you send a POST with custom headers and a body?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Build a urllib.request.Request whose data argument is a bytes body and whose headers argument is a dict, then hand that Request to urlopen. Supplying data switches the method from GET to POST; a str body raises TypeError.

open as a page

How do you stop urllib.request.urlopen from following a redirect?

level: middleimportance: should knowfreq 35%

basics

~10 s

Subclass urllib.request.HTTPRedirectHandler so its redirect_request returns None, pass the subclass to build_opener, and call open on that opener. With no handler willing to redirect, the 3xx surfaces as an HTTPError carrying the Location header.

open as a page

A metrics scraper passes timeout=2 to urllib.request.urlopen yet a call still runs for minutes. What does that timeout actually bound?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It is a socket timeout applied to each blocking operation — the connect and every individual read — not a deadline for the whole exchange. A target trickling bytes resets it indefinitely, so a total budget has to be enforced outside the call.

open as a page