skip to content

Which encoding do sys.stdout and sys.stdin use, and what does PYTHONIOENCODING change?

level: middleimportance: should knowfreq 30%

answer

  1. The streams are text wrappers as well
  2. One of the three cannot afford to raise
  3. A variable with two colon-separated halves
  4. A method reconfigures it in place
  5. Piping changes buffering, not the codec

basics

~10 s

The standard streams are text wrappers that use the locale encoding, or UTF-8 under UTF-8 mode. PYTHONIOENCODING overrides that as encodingname:errorhandler, with either half optional; sys.stdout.reconfigure() changes it from inside a running program.

solid answer

~40 s

`sys.stdin`, `sys.stdout` and `sys.stderr` are `io.TextIOWrapper` objects over the process's binary streams, and their encoding is the locale encoding — UTF-8 when UTF-8 mode is on. Their error handlers differ: stdout and stdin default to `strict`, so printing a character the encoding cannot represent raises `UnicodeEncodeError`, while `sys.stderr` uses `backslashreplace` so that reporting an error can never itself fail. `PYTHONIOENCODING` overrides both parts in the form `encodingname:errorhandler`, and either half may be omitted, so `PYTHONIOENCODING=:replace` keeps the encoding and only relaxes the handler. Inside a running program, `sys.stdout.reconfigure(encoding="utf-8", errors="backslashreplace")` does the same job (3.7+). Two gotchas: redirecting to a pipe does not change the encoding in Python — only the buffering — and on Windows the variable is ignored for the interactive console unless `PYTHONLEGACYWINDOWSSTDIO` is also set.

code

console · 4 lines
console
$ PYTHONIOENCODING=ascii:backslashreplace python3 -c "print('caf\u00e9')"
caf\xe9
$ PYTHONIOENCODING=utf-8 python3 -c "import sys; print(sys.stdout.encoding, sys.stdout.errors)"
utf-8 strict

go deeper

for a junior

Know that print() can raise UnicodeEncodeError when the stream's encoding cannot represent a character, and that sys.stdout.encoding tells you which encoding is in play.

for a middle

Explain that the streams are text wrappers using the locale encoding, that stderr alone defaults to backslashreplace, and give PYTHONIOENCODING's encoding:errorhandler syntax plus the reconfigure() alternative.

for a senior

An interviewer expects a diagnosis under pressure: distinguish an encoding failure from a buffering one, choose strict versus backslashreplace based on whether the output is machine-parsed or human-read, and never silence it with errors='ignore'.

for a principal

Own the contract for machine-readable output: which encoding your services emit, whether mis-encoding should be loud or lossy, and how that is enforced in runtime configuration rather than left to each host's locale.

### The streams are files too `sys.stdout` is not a magic object: it is an `io.TextIOWrapper` around a buffered binary stream, exactly like the object `open()` returns in text mode, and it carries the same two properties — `sys.stdout.encoding` and `sys.stdout.errors`. Everything that is true about a text file's encoding is true here, with a few deliberate differences. ### The defaults The encoding is the locale encoding, or UTF-8 when UTF-8 mode is enabled (`-X utf8` / `PYTHONUTF8=1`, or automatically under the `C`/`POSIX` locale). The error handlers are chosen per stream: * `sys.stdout` and `sys.stdin` use `strict`. Encoding a character the codec cannot represent raises `UnicodeEncodeError` — the classic `print()` traceback on a host whose locale is ASCII or a single-byte code page. * `sys.stderr` uses `backslashreplace`. Unrepresentable characters become escape sequences rather than an exception, on the principle that the channel you use to report failures must not be able to fail. Under UTF-8 mode, stdin and stdout switch to `surrogateescape`, so bytes that are not valid UTF-8 survive a decode/encode round trip instead of raising. ### `PYTHONIOENCODING` The environment variable overrides the streams' encoding and error handler at startup. Its syntax is `encodingname:errorhandler`, and either half is optional: * `PYTHONIOENCODING=utf-8` — force the encoding, keep the default handlers. * `PYTHONIOENCODING=:backslashreplace` — keep the encoding, never raise on output. * `PYTHONIOENCODING=ascii:xmlcharrefreplace` — both. It is the right tool when you do not own the program you are running — a third-party entry point whose output goes into a log pipeline, for example. When you do own the code, `sys.stdout.reconfigure(encoding="utf-8", errors="backslashreplace")` (added in 3.7) is better, because it does not depend on how the process was launched. `reconfigure()` on a stream that already has buffered content flushes it first, and it is only available on text streams. ### The gotchas interviews probe **Piping does not change the encoding.** Unlike some runtimes, CPython does not switch to a different codec when stdout is not a terminal. What redirection does change is *buffering*: interactive streams are line-buffered, redirected ones are block-buffered, which is why output can appear to vanish until the process exits. `-u` or `PYTHONUNBUFFERED=1` turns that off. Confusing the two — blaming an encoding for missing output — is a common wrong diagnosis. **The console on Windows is special.** Since 3.6 the interactive Windows console is written through the wide-character Win32 API rather than an encoded byte stream, and `PYTHONIOENCODING` is ignored for it unless `PYTHONLEGACYWINDOWSSTDIO` is also set. Redirected output on Windows is unaffected by that carve-out and honours the variable normally. **Streams and files are separate decisions.** `PYTHONIOENCODING` says nothing about `open()`; UTF-8 mode changes both. A program can therefore print fine while reading its data files wrongly, or the reverse, and a diagnosis that does not distinguish the two chases the wrong knob. **Detached and replaced streams.** `sys.stdout` is a normal attribute: a test harness or a logging shim may have replaced it with an object that has no `encoding` at all, so defensive code should not assume the attribute exists. And `sys.stdout.buffer` gives you the underlying binary stream when you genuinely need to write bytes — the correct way to emit pre-encoded output rather than fighting the wrapper. ### Diagnosing the classic failure A batch job prints a name containing an accented letter and dies with `UnicodeEncodeError` on one machine only. The sequence is: read `sys.stdout.encoding` on both machines; if it is not UTF-8, look at `LANG`/`LC_ALL` in the unit or job definition and at whether the image ships any locale at all. The immediate mitigation is `PYTHONUTF8=1` or `PYTHONIOENCODING=utf-8` on the process. The durable fix depends on what the output is for: if it feeds a log pipeline that expects UTF-8, set the encoding and keep `strict` so mis-encoding is loud; if it is human-facing diagnostics, `backslashreplace` is the humane handler. Reaching for `errors="ignore"` is the wrong answer in both cases — it silently deletes the characters that identified the problem.

  • Why does sys.stderr use a different error handler from sys.stdout?
    Because the channel that reports failures must not be able to fail. `sys.stderr` uses `backslashreplace`, so an unrepresentable character becomes an escape sequence instead of raising `UnicodeEncodeError` in the middle of a traceback or a log line. `sys.stdout` keeps `strict` because it usually carries data that another program will parse, where silent substitution would be worse than an error.
  • Does redirecting stdout to a file or pipe change its encoding?
    No. CPython uses the same codec whether or not the stream is a terminal; what changes is buffering, from line-buffered on a tty to block-buffered when redirected, which is why output can seem to disappear until the process exits. `-u` or `PYTHONUNBUFFERED=1` restores immediate writes. Blaming an encoding for delayed output is a common misdiagnosis.
  • When would you write to sys.stdout.buffer instead of using print()?
    When you already hold bytes and must not have them re-encoded — emitting binary output, forwarding a payload verbatim, or writing content whose encoding was decided elsewhere. `sys.stdout.buffer` is the underlying binary stream beneath the text wrapper. Mixing the two needs care: flush the text layer before writing bytes, or the interleaving will not match the order of your calls.

saying these in an interview costs you the question

  • Thinks Python switches codec when output is piped
  • Believes PYTHONIOENCODING also changes open()'s default
  • Reaches for errors='ignore' to stop the traceback
  • Assumes all three streams share one error handler
  • Expects the variable to affect the Windows console
  • Confuses missing output with an encoding problem

context