skip to content

How does CPython choose sys.stdout's text encoding, and how do you force it to UTF-8?

level: middleimportance: should knowfreq 35%

answer

  1. Something outside the program decides it
  2. Fixed when the interpreter starts up
  3. Locale-derived, and stderr plays by different rules
  4. PYTHONIOENCODING, UTF-8 mode, or reconfigure

basics

~10 s

CPython builds sys.stdout at startup with the encoding derived from the process locale, so a non-UTF-8 locale can make non-ASCII output raise UnicodeEncodeError. Force UTF-8 with PYTHONIOENCODING, UTF-8 mode, or sys.stdout.reconfigure.

solid answer

~40 s

`sys.stdout` is an `io.TextIOWrapper`, and its `encoding` and `errors` are fixed when the interpreter initializes, from the process locale. On a UTF-8 host that is UTF-8 and nothing surprises you; on a host with a legacy 8-bit locale the same `print()` of a non-ASCII string raises `UnicodeEncodeError`, because the default error handler for stdout is strict. `sys.stderr` always uses `backslashreplace`, so tracebacks never die of an encoding error. Three overrides, in widening scope: `PYTHONIOENCODING=utf-8` (or `utf-8:backslashreplace`) changes only the standard streams; UTF-8 mode, via `-X utf8` or `PYTHONUTF8=1`, changes the locale-encoding default across the interpreter and is auto-enabled when the locale is C or POSIX, but is otherwise opt-in on 3.14; and `sys.stdout.reconfigure(encoding="utf-8")` changes it from inside the running program.

code

console · 2 lines
console
PYTHONIOENCODING=ascii python3 -c "print('caf\u00e9')"
PYTHONIOENCODING=ascii:backslashreplace python3 -c "print('caf\u00e9')"

go deeper

for a junior

Recall that printing non-ASCII text can fail on a machine whose environment is not set up for UTF-8, and that the encoding of the output stream comes from outside the program rather than from the code.

for a middle

Explain the mechanics: the stream's codec and error handler are fixed at interpreter startup from the locale, stderr uses backslashreplace, and PYTHONIOENCODING, UTF-8 mode and reconfigure are the three overrides with different scopes.

for a senior

Diagnose it as an environment difference rather than a code defect, read encoding, errors and sys.flags.utf8_mode from the failing host, and choose deliberately between forcing UTF-8 and accepting a lossy error handler.

for a principal

Own the text-encoding contract across the fleet: one declared output encoding, set where every process inherits it, so behaviour does not depend on whichever locale a base image or scheduler happens to provide.

## Where the encoding comes from Writing `str` to a stream requires encoding it to bytes, and something must pick the codec. `sys.stdout` is an `io.TextIOWrapper` created during interpreter startup, and the encoding it is given is derived from the process's locale — the same value `locale.getencoding()` reports. This is decided once. Nothing re-reads the locale later, so a program cannot change it by calling into the locale module halfway through. The **error handler** matters just as much as the codec, and the two standard output streams are given different ones on purpose: * `sys.stdout` uses a strict handler in the ordinary case, so a character the codec cannot represent raises `UnicodeEncodeError` rather than silently mangling data. * `sys.stderr` uses `backslashreplace`. Unrepresentable characters become escape sequences instead of an exception — deliberately, so that a traceback containing an exotic character can always be printed. A diagnostic that cannot be displayed is worse than a lossy one. You can read both at runtime: `sys.stdout.encoding` and `sys.stdout.errors`. ## Why this shows up as an environment bug The classic report is "it works on my machine and crashes on the host". The developer's shell has a UTF-8 locale; the host runs the process under a minimal or legacy environment whose locale encoding is an 8-bit codec. The first non-ASCII character — a name, a currency symbol, a check mark in a status line — raises `UnicodeEncodeError` on a `print()` that has never failed anywhere else. The code is identical; the interpreter's startup configuration is not. One mitigation is already built in: **UTF-8 mode is enabled automatically when the locale is `C` or `POSIX`**, so the most common bare-environment case does end up with UTF-8. It is the genuinely non-UTF-8 locale, not the empty one, that still bites. ## The three overrides **`PYTHONIOENCODING=encoding[:errors]`** — the narrowest. It changes only the standard streams, and it takes an optional error handler after a colon: `PYTHONIOENCODING=utf-8:backslashreplace`. Use it when the process must emit UTF-8 regardless of where it runs and you do not want to touch anything else about text handling. **UTF-8 mode**, enabled with `-X utf8` on the command line or `PYTHONUTF8=1` in the environment — the broadest. It makes the interpreter ignore the locale for text encoding purposes and use UTF-8 as the locale-encoding default throughout, with `surrogateescape` on stdin and stdout so undecodable input round-trips rather than raising. `sys.flags.utf8_mode` reports whether it is on. On CPython 3.14 it is **opt-in** except for the automatic C/POSIX case above — do not write code that assumes it is on. **`sys.stdout.reconfigure(encoding="utf-8", errors="backslashreplace")`** — from inside the program, available on `io.TextIOWrapper` since Python 3.7. It flushes and re-creates the text layer over the same buffered stream. The catch is ordering: it only governs writes that happen after it runs, so it belongs as close to the top of the entrypoint as you can put it, and it cannot help output produced by an import that ran earlier. As a rule, prefer configuring the environment for a service you deploy, and reach for `reconfigure` when you are shipping a program whose environment you do not control and cannot instruct. ## Two things not to confuse it with `sys.getdefaultencoding()` always returns `'utf-8'` on Python 3 and is about the implicit `str`/`bytes` conversions inside the interpreter — it says nothing about your streams. And the encoding of `sys.stdout` is independent of its **buffering**: `PYTHONIOENCODING` does not make output appear sooner, and `PYTHONUNBUFFERED` does not make an unrepresentable character printable. They are two separate startup decisions about the same object. ## Verifying instead of assuming Print `sys.stdout.encoding`, `sys.stdout.errors` and `sys.flags.utf8_mode` from inside the deployed environment. Those three values explain essentially every mystery in this area, and they take one line to obtain. If the answer is a legacy codec, decide deliberately: force UTF-8 and keep the data intact, or keep the locale's codec and choose a lossy error handler so the process cannot be brought down by a character.

  • Why does sys.stderr not raise UnicodeEncodeError on the same character that breaks sys.stdout?
    Because `sys.stderr` is created with the `backslashreplace` error handler while `sys.stdout` gets a strict one in the ordinary case. Unrepresentable characters on stderr become escape sequences rather than an exception, which is deliberate: a traceback must always be printable, even when it contains text the stream's codec cannot represent. You can confirm it by reading `sys.stderr.errors`.
  • What is the difference between PYTHONIOENCODING and enabling UTF-8 mode?
    `PYTHONIOENCODING` retargets only the standard streams, and accepts an error handler after a colon. UTF-8 mode, via `-X utf8` or `PYTHONUTF8=1`, tells the whole interpreter to use UTF-8 as the locale-encoding default rather than consulting the locale, and uses `surrogateescape` on stdin and stdout. `sys.flags.utf8_mode` reports it. On 3.14 it is opt-in, apart from being auto-enabled under a C or POSIX locale.
  • Does sys.getdefaultencoding() tell you what sys.stdout will use?
    No. `sys.getdefaultencoding()` always returns `'utf-8'` on Python 3 and describes the interpreter's internal default for implicit str/bytes conversions. The stream's codec is a separate, locale-derived startup decision, readable as `sys.stdout.encoding`. Quoting the former as evidence about the latter is a common way to misdiagnose an encoding failure that only happens on one host.

saying these in an interview costs you the question

  • Says sys.getdefaultencoding() reports the stream's codec
  • Assumes UTF-8 mode is on by default in 3.14
  • Thinks PYTHONIOENCODING also changes stream buffering
  • Believes the locale is re-read after interpreter startup
  • Cannot explain why stderr survives the same character
  • Encodes to bytes and writes to sys.stdout to dodge the error

context