skip to content

Using urllib.request, how do you send a POST with custom headers and a body?

level: juniorimportance: should knowfreq 45%

answer

  1. Two objects, then one call
  2. Headers live on the request object
  3. The body picks the verb
  4. Bytes only, never str
  5. data= means POST unless overridden

basics

~20 s

Build a urllib.request.Request whose data argument is a bytes body and whose headers argument is a dict, then hand that Request to urlopen. Supplying data switches the method from GET to POST; a str body raises TypeError.

solid answer

~40 s

`urlopen` accepts either a URL string or a `urllib.request.Request`, and only the `Request` can carry headers — there is no `headers` keyword on `urlopen`. So construct `Request(url, data=body, headers={...})`, or attach them one at a time with `add_header`, then call `urlopen(req, timeout=...)`. The body must be a bytes-like object, an iterable of bytes, or a file object; a `str` raises `TypeError` with an explicit message, so encode form data with `urllib.parse.urlencode(...).encode()` and JSON with `json.dumps(...).encode()`. Because `data` is not `None`, urllib sends POST unless you override it with the `method` argument. urllib then fills in `Content-Type: application/x-www-form-urlencoded` when you set none, plus `Content-Length` (or chunked transfer encoding when the length is unknown), `Host`, `Connection: close` and a `User-Agent` of `Python-urllib/3.14`.

code

python · 12 lines
python
import urllib.parse
import urllib.request

body = urllib.parse.urlencode({"job": "scrape", "window": "5m"}).encode("utf-8")
req = urllib.request.Request(
    "https://metrics.invalid/ingest",
    data=body,
    headers={"X-Scrape-Id": "42"},
)
print(req.get_method())     # POST, because data is not None
print(req.data)             # b'job=scrape&window=5m'
print(req.header_items())   # [('X-scrape-id', '42')] - note the casing

go deeper

for a junior

Be ready to fetch a URL with nothing installed: urlopen for a GET, and a Request carrying a bytes body for a POST. Remember that the response you read back is bytes, not text.

for a middle

Explain why supplying a body flips the method, which headers urllib fills in on your behalf, and how urlencode plus encode produces a form body while JSON needs its content type set by hand.

for a senior

Show that you always pass an explicit timeout, use the response as a context manager, and know what the stdlib client does not give you: no pooling, no retries, no session. Say when that is enough.

for a principal

Own the call on whether services take an HTTP client dependency at all. The stdlib path costs nothing to install but hands you no retry policy, connection reuse or instrumentation seam, and that cost lands on every team that copies the snippet.

## Two objects, one call `urllib.request` splits an HTTP call into a **request object** and a **call that performs it**. `urlopen` is the performer; it takes either a plain URL string or a `urllib.request.Request` instance. Everything about the request other than the URL, the body and the timeout lives on the `Request`: the headers, the method override, the proxy flags. This is the first thing to internalise, because the most common beginner mistake is reaching for a `headers` keyword on `urlopen` that does not exist and never has. ```python import json import urllib.request payload = json.dumps({"scrape": "ok"}).encode("utf-8") req = urllib.request.Request( "https://example.invalid/ingest", data=payload, headers={"Content-Type": "application/json"}, ) # urlopen(req, timeout=5) would send it ``` ## The body decides the verb `Request` takes `data=None` by default. When `data` is not `None`, urllib's HTTP handler treats the request as a POST: the method returned by the request object is `POST` rather than `GET`. There is no separate switch. If you want a different verb you pass `method="PUT"` (or `"DELETE"`, `"PATCH"`, `"HEAD"`) to the constructor, and that explicit method wins over the data-implies-POST rule. `urlopen` itself has no method parameter at all, so a verb other than GET always means you built a `Request`. ## The body must be bytes HTTP bodies are octets, and urllib refuses to guess an encoding for you. Passing a `str` raises `TypeError` with the message that POST data should be bytes, an iterable of bytes, or a file object. Note *when* it is raised: the constructor happily stores the string, and the error surfaces later, when the handler prepares the request — so the traceback points at `urlopen`, not at the line where you built the `Request`. The two everyday encodings are: * form-encoded — `urllib.parse.urlencode({"k": "v"}).encode("utf-8")`, producing `b"k=v"`; * JSON — `json.dumps(obj).encode("utf-8")`, and you must set `Content-Type` yourself, because urllib will not infer it from the payload. A file object or any iterable of bytes is also accepted. When urllib cannot determine a length for the body it switches to chunked transfer encoding instead of sending `Content-Length`. ## What urllib adds for you Just before the request goes out, the HTTP handler tops up the headers. If you supplied a body and no content type, it adds `Content-Type: application/x-www-form-urlencoded` — which is right for form data and wrong for everything else, so set it explicitly for JSON. It adds `Content-Length` when the body has a length, `Host` derived from the URL, `Connection: close` (urllib does not keep connections alive; every call is a fresh connection), and a default `User-Agent` of `Python-urllib/` plus the interpreter's major and minor version. Many public endpoints reject that user agent, which is why a request that works from a browser can fail from a script until you override the header. These fill-ins are attached as *unredirected* headers, meaning they are deliberately not copied onto a follow-up request if the response is a redirect. ## The header-casing surprise Headers you pass in the `headers` dict all go through `add_header`, which stores each key as `key.capitalize()`. `X-Scrape-Id` becomes `X-scrape-id`. On the wire this is harmless, because HTTP header names are case-insensitive, but it matters twice: when you inspect the request object in a test, and when you call `has_header`, `get_header` or `remove_header`, all of which look up the capitalized spelling. ## Reading the response For `http` and `https` URLs, `urlopen` returns an `http.client.HTTPResponse`. It is a context manager and a binary file-like object, so the idiomatic shape is: ```python import urllib.request with urllib.request.urlopen("https://example.invalid/", timeout=5) as resp: status = resp.status body = resp.read().decode("utf-8") ``` `read()` returns `bytes`; decoding is your job, using the charset from the response headers if you care about correctness. The `status` attribute holds the numeric status, `headers` the response headers, and `url` the URL that was actually fetched. ## What the standard library will not do The stdlib client is deliberately thin. There is no connection pooling, no retry policy, no session object holding cookies by default, no automatic JSON encoding or decoding, and no timeout unless you pass one. That is the honest trade: zero dependencies, and every policy decision left to you. Knowing where the line falls is what an interviewer is checking for — the mechanics above are enough for a health check, a metrics scrape or a build script, and reaching for something heavier is a judgement call rather than a reflex.

  • If you pass a body but set no Content-Type header, what content type does urllib send?
    `application/x-www-form-urlencoded`, added by the HTTP handler just before the request leaves, together with `Content-Length` — or chunked transfer encoding when the body has no knowable length. Both are attached as unredirected headers, so they are dropped if the request is redirected. If you are posting JSON, set the header yourself; urllib does not inspect the payload to guess.
  • Why does the request object show X-scrape-id after you passed X-Scrape-Id?
    `add_header` — which the `headers` dict also flows through — stores keys as `key.capitalize()`, so only the first letter survives as upper case. HTTP header names are case-insensitive on the wire, so no server notices, but your own `has_header`, `get_header` and `remove_header` lookups must use the capitalized spelling, and a test asserting on the request's header keys will fail if it assumes the original casing.
  • How do you send a PUT or a DELETE with urllib.request?
    Pass the `method` argument to the `Request` constructor, for example `Request(url, data=body, method="PUT")`. That explicit method overrides the data-implies-POST rule, and a request with `method="DELETE"` and no data simply sends no body. `urlopen` has no method parameter of its own, so the `Request` object is the only place a verb can be set.

The Request is the addressed, stamped envelope; urlopen is the act of posting it. You cannot write on the envelope after you have dropped it in the box.

saying these in an interview costs you the question

  • Passing a str body and expecting urllib to encode it
  • Believing urlopen takes a headers keyword argument
  • Thinking urlopen(url, data) still sends a GET
  • Assuming urllib serialises a dict body automatically
  • Forgetting the response is bytes and needs decoding

context