Which encoding do sys.stdout and sys.stdin use, and what does PYTHONIOENCODING change?
answer
- The streams are text wrappers as well
- One of the three cannot afford to raise
- A variable with two colon-separated halves
- A method reconfigures it in place
- Piping changes buffering, not the codec
basics
~10 sThe standard streams are text wrappers that use the locale encoding, or UTF-8 under UTF-8 mode. PYTHONIOENCODING overrides that as encodingname:errorhandler, with either half optional; sys.stdout.reconfigure() changes it from inside a running program.
solid answer
~40 s`sys.stdin`, `sys.stdout` and `sys.stderr` are `io.TextIOWrapper` objects over the process's binary streams, and their encoding is the locale encoding — UTF-8 when UTF-8 mode is on. Their error handlers differ: stdout and stdin default to `strict`, so printing a character the encoding cannot represent raises `UnicodeEncodeError`, while `sys.stderr` uses `backslashreplace` so that reporting an error can never itself fail. `PYTHONIOENCODING` overrides both parts in the form `encodingname:errorhandler`, and either half may be omitted, so `PYTHONIOENCODING=:replace` keeps the encoding and only relaxes the handler. Inside a running program, `sys.stdout.reconfigure(encoding="utf-8", errors="backslashreplace")` does the same job (3.7+). Two gotchas: redirecting to a pipe does not change the encoding in Python — only the buffering — and on Windows the variable is ignored for the interactive console unless `PYTHONLEGACYWINDOWSSTDIO` is also set.
code
console · 4 lines$ PYTHONIOENCODING=ascii:backslashreplace python3 -c "print('caf\u00e9')"
caf\xe9
$ PYTHONIOENCODING=utf-8 python3 -c "import sys; print(sys.stdout.encoding, sys.stdout.errors)"
utf-8 strictgo deeper
Know that print() can raise UnicodeEncodeError when the stream's encoding cannot represent a character, and that sys.stdout.encoding tells you which encoding is in play.
Explain that the streams are text wrappers using the locale encoding, that stderr alone defaults to backslashreplace, and give PYTHONIOENCODING's encoding:errorhandler syntax plus the reconfigure() alternative.
An interviewer expects a diagnosis under pressure: distinguish an encoding failure from a buffering one, choose strict versus backslashreplace based on whether the output is machine-parsed or human-read, and never silence it with errors='ignore'.
Own the contract for machine-readable output: which encoding your services emit, whether mis-encoding should be loud or lossy, and how that is enforced in runtime configuration rather than left to each host's locale.
### The streams are files too `sys.stdout` is not a magic object: it is an `io.TextIOWrapper` around a buffered binary stream, exactly like the object `open()` returns in text mode, and it carries the same two properties — `sys.stdout.encoding` and `sys.stdout.errors`. Everything that is true about a text file's encoding is true here, with a few deliberate differences. ### The defaults The encoding is the locale encoding, or UTF-8 when UTF-8 mode is enabled (`-X utf8` / `PYTHONUTF8=1`, or automatically under the `C`/`POSIX` locale). The error handlers are chosen per stream: * `sys.stdout` and `sys.stdin` use `strict`. Encoding a character the codec cannot represent raises `UnicodeEncodeError` — the classic `print()` traceback on a host whose locale is ASCII or a single-byte code page. * `sys.stderr` uses `backslashreplace`. Unrepresentable characters become escape sequences rather than an exception, on the principle that the channel you use to report failures must not be able to fail. Under UTF-8 mode, stdin and stdout switch to `surrogateescape`, so bytes that are not valid UTF-8 survive a decode/encode round trip instead of raising. ### `PYTHONIOENCODING` The environment variable overrides the streams' encoding and error handler at startup. Its syntax is `encodingname:errorhandler`, and either half is optional: * `PYTHONIOENCODING=utf-8` — force the encoding, keep the default handlers. * `PYTHONIOENCODING=:backslashreplace` — keep the encoding, never raise on output. * `PYTHONIOENCODING=ascii:xmlcharrefreplace` — both. It is the right tool when you do not own the program you are running — a third-party entry point whose output goes into a log pipeline, for example. When you do own the code, `sys.stdout.reconfigure(encoding="utf-8", errors="backslashreplace")` (added in 3.7) is better, because it does not depend on how the process was launched. `reconfigure()` on a stream that already has buffered content flushes it first, and it is only available on text streams. ### The gotchas interviews probe **Piping does not change the encoding.** Unlike some runtimes, CPython does not switch to a different codec when stdout is not a terminal. What redirection does change is *buffering*: interactive streams are line-buffered, redirected ones are block-buffered, which is why output can appear to vanish until the process exits. `-u` or `PYTHONUNBUFFERED=1` turns that off. Confusing the two — blaming an encoding for missing output — is a common wrong diagnosis. **The console on Windows is special.** Since 3.6 the interactive Windows console is written through the wide-character Win32 API rather than an encoded byte stream, and `PYTHONIOENCODING` is ignored for it unless `PYTHONLEGACYWINDOWSSTDIO` is also set. Redirected output on Windows is unaffected by that carve-out and honours the variable normally. **Streams and files are separate decisions.** `PYTHONIOENCODING` says nothing about `open()`; UTF-8 mode changes both. A program can therefore print fine while reading its data files wrongly, or the reverse, and a diagnosis that does not distinguish the two chases the wrong knob. **Detached and replaced streams.** `sys.stdout` is a normal attribute: a test harness or a logging shim may have replaced it with an object that has no `encoding` at all, so defensive code should not assume the attribute exists. And `sys.stdout.buffer` gives you the underlying binary stream when you genuinely need to write bytes — the correct way to emit pre-encoded output rather than fighting the wrapper. ### Diagnosing the classic failure A batch job prints a name containing an accented letter and dies with `UnicodeEncodeError` on one machine only. The sequence is: read `sys.stdout.encoding` on both machines; if it is not UTF-8, look at `LANG`/`LC_ALL` in the unit or job definition and at whether the image ships any locale at all. The immediate mitigation is `PYTHONUTF8=1` or `PYTHONIOENCODING=utf-8` on the process. The durable fix depends on what the output is for: if it feeds a log pipeline that expects UTF-8, set the encoding and keep `strict` so mis-encoding is loud; if it is human-facing diagnostics, `backslashreplace` is the humane handler. Reaching for `errors="ignore"` is the wrong answer in both cases — it silently deletes the characters that identified the problem.
- Why does sys.stderr use a different error handler from sys.stdout?Because the channel that reports failures must not be able to fail. `sys.stderr` uses `backslashreplace`, so an unrepresentable character becomes an escape sequence instead of raising `UnicodeEncodeError` in the middle of a traceback or a log line. `sys.stdout` keeps `strict` because it usually carries data that another program will parse, where silent substitution would be worse than an error.
- Does redirecting stdout to a file or pipe change its encoding?No. CPython uses the same codec whether or not the stream is a terminal; what changes is buffering, from line-buffered on a tty to block-buffered when redirected, which is why output can seem to disappear until the process exits. `-u` or `PYTHONUNBUFFERED=1` restores immediate writes. Blaming an encoding for delayed output is a common misdiagnosis.
- When would you write to sys.stdout.buffer instead of using print()?When you already hold bytes and must not have them re-encoded — emitting binary output, forwarding a payload verbatim, or writing content whose encoding was decided elsewhere. `sys.stdout.buffer` is the underlying binary stream beneath the text wrapper. Mixing the two needs care: flush the text layer before writing bytes, or the interleaving will not match the order of your calls.
saying these in an interview costs you the question
- Thinks Python switches codec when output is piped
- Believes PYTHONIOENCODING also changes open()'s default
- Reaches for errors='ignore' to stop the traceback
- Assumes all three streams share one error handler
- Expects the variable to affect the Windows console
- Confuses missing output with an encoding problem