How does the HTTP Accept-Language header work, and why is the Accept-Charset header effectively obsolete?
answer
- BCP 47 tags: lang-Script-REGION
- filtering = prefix on subtag boundary; lookup = truncate right
- answer with Content-Language + Vary
- stock the base language or fall back
- Accept-Charset deprecated — UTF-8 won
basics
~20 sAccept-Language lists language tags such as en-GB or fr that a client prefers for the response; the server picks a translation and reports it in Content-Language. Accept-Charset is obsolete because everything is UTF-8 now, so browsers stopped sending it.
solid answer
~50 s`Accept-Language` carries **language tags** (BCP 47): `en`, `en-GB`, `pt-BR`, `zh-Hant`. The server matches them against the translations it actually has, picks one, and reports the result in **`Content-Language`** — plus `Vary: Accept-Language` so caches key on it. Matching is not string equality. Two standard algorithms exist: *basic filtering*, where `en` matches `en-GB` by prefix on whole subtags, and *lookup*, which progressively truncates the requested tag (`en-GB-oxendict` → `en-GB` → `en`) until something is found, then falls back to a configured default. Note the asymmetry: `en` requests match `en-GB` content, but a request for `en-GB` should not silently get `en-US` without truncation logic. `Accept-Charset` was deprecated in RFC 7231 and browsers no longer send it: UTF-8 is universal, and sending the header mostly leaked fingerprinting entropy. Character encoding today travels as a `charset` parameter on `Content-Type`, and servers should just emit UTF-8.
code
http · 8 linesGET /docs/intro HTTP/1.1
Host: example.com
Accept-Language: en-GB, en, fr
HTTP/1.1 200 OK
Content-Type: text/html;charset=utf-8
Content-Language: en-GB
Vary: Accept-Languagego deeper
Know that Accept-Language lists preferred languages, Content-Language reports what was served, and Accept-Charset is dead because everything is UTF-8.
Explain BCP 47 tag structure, filtering versus lookup matching, and the need for Vary: Accept-Language.
Cover the operational traps: stocking base languages, normalising the header at the edge to protect cache hit rate, and persisting an explicit user choice over the header hint.
Decide the localisation URL strategy (per-locale paths versus one negotiated URL) against SEO, caching, and link-sharing needs, and set the fallback chain policy across the estate.
## Accept-Language Natural language is a negotiable axis just like media type. The client sends the languages it can read: ``` Accept-Language: en-GB, en, fr ``` Each entry is a **language tag** as defined by BCP 47 — a hyphen-separated sequence of subtags: primary language (`en`), optional script (`Hant`), optional region (`GB`), optional variants (`oxendict`). Tags are case-insensitive, though the convention is lowercase language, Titlecase script, UPPERCASE region: `zh-Hant-TW`. The wildcard `*` is allowed and means "any language", which is what a client says when it will take whatever you have. ## Matching: filtering vs lookup RFC 4647 defines two ways to match a requested tag against available tags: **Basic filtering** — a range matches a tag if the tag equals the range or begins with the range followed by a hyphen. So the range `en` matches `en`, `en-GB`, `en-US`, but *not* `eng` (subtag boundaries are respected, it is not raw prefix matching). **Lookup** — used when you must return exactly one result. Truncate the requested tag from the right, one subtag at a time, until an available tag matches: `en-GB-oxendict` → `en-GB-oxendict`? → `en-GB`? → `en`? → default. Note that single-character subtags are skipped during truncation, and that lookup never widens sideways — it will not turn `en-GB` into `en-US` unless you also stock `en`. The practical consequence: **stock the base language.** If your catalogue has `en-US` and `pt-BR` only, a client asking for `en-GB` or `pt-PT` gets nothing by lookup and falls to your default. Publishing generic `en` and `pt` entries, or explicitly mapping regional variants, avoids the surprise of a UK user seeing your fallback language. ## Reporting the choice The response should say which language it actually is: ``` Content-Language: en-GB Vary: Accept-Language ``` `Content-Language` describes the intended audience of the representation and may list several tags for genuinely multilingual documents. `Vary: Accept-Language` is required for any shared cache to be correct — without it, the first visitor's language is served to everyone. Because Accept-Language values are long and highly variable, edge caches normally **normalise** the header down to the small set of languages actually supported before it enters the cache key. ## Where the language preference should really come from A seasoned answer notes that `Accept-Language` is a *default*, not a decision. It reflects OS/browser configuration, which is often wrong (a shared machine, a corporate image, a traveller). Production systems therefore treat it as the first-visit hint and then persist an explicit user choice — in a cookie, a profile setting, or a URL/locale path segment — and keep the URL distinct per language where SEO matters. Serving different languages from one URL is legal but makes indexing and link sharing harder. It is also a fingerprinting surface: the full ordered list of languages is unusually identifying, which is why privacy-hardened browsers trim it. ## Accept-Charset: why it is gone `Accept-Charset` was the request-side control for character encoding — e.g. `Accept-Charset: utf-8, iso-8859-1`. It was **deprecated in RFC 7231** and is not sent by any current browser. The reasons: - **UTF-8 won.** Essentially all modern content is UTF-8; negotiating an encoding solves a problem nobody has any more. - **Poor value, real cost.** It added per-request bytes and fingerprinting entropy for near-zero benefit. - **Wrong layer.** Encoding is a property of the representation and travels perfectly well as a media-type parameter on the response: `Content-Type: text/html;charset=utf-8`. For `application/json`, the encoding is fixed as UTF-8 by the JSON media type itself, so a `charset` parameter there is redundant (and harmless but noisy). A server that still branches on `Accept-Charset` is coding for a client that does not exist. The correct modern behaviour is: always emit UTF-8, always label text media types with `charset=utf-8`, and ignore `Accept-Charset` entirely. Note the name collision to avoid in interviews: `Accept-Charset` is about character encoding, `Accept-Encoding` is about compression — unrelated concerns despite the similar names.
- A user's browser sends `Accept-Language: en-GB` and your catalogue has only `en-US` and `de`. What should happen?By lookup, `en-GB` truncates to `en`, which you do not stock, so the request falls through to your configured default. The fix is to publish a generic `en` entry (or map regional variants explicitly) rather than relying on clients to ask for exactly the tags you happen to have.
- Should a production site rely on Accept-Language alone to choose the UI language?No. It reflects browser or OS configuration, which is frequently not what the person wants, and it is a fingerprinting surface. Use it as the first-visit hint, then persist an explicit user choice in a cookie or profile — and prefer distinct URLs or locale path segments when discoverability and link sharing matter.
- How is Accept-Charset different from Accept-Encoding?Accept-Charset negotiated the character encoding of text (UTF-8 vs ISO-8859-1) and is deprecated. Accept-Encoding negotiates content compression (gzip, br, zstd) and is very much alive. The similar names are a common source of confusion; charset now travels as a parameter on Content-Type.
Lookup is like asking for a book in 'British English, Oxford spelling', then settling for British English, then just English, then whatever's on the shelf — you never sideways-swap to American English unless the plain 'English' shelf exists.
saying these in an interview costs you the question
- Confusing Accept-Charset (character encoding, obsolete) with Accept-Encoding (compression, current).
- Matching language tags with naive case-sensitive string equality.
- Assuming a request for `en` cannot be satisfied by `en-GB` content.
- Serving negotiated languages without `Vary: Accept-Language`.
- Treating Accept-Language as an authoritative user choice instead of a default hint.