What encoding does Python's open() use when you do not pass encoding=?
answer
- The call names no encoding at all
- The environment decides, not the file
- Same code, different host, different characters
- One locale call reports the value
- Windows ANSI code page versus UTF-8
basics
~20 sA text-mode open() with no encoding= uses the locale encoding that locale.getencoding() reports: usually UTF-8 on Linux and macOS, but the ANSI code page on Windows. It is not guaranteed to be UTF-8, so pass encoding= explicitly.
solid answer
~40 sIn text mode, `open(path)` wraps the binary file in a `io.TextIOWrapper` and decodes with the **locale encoding** — what `locale.getencoding()` returns (added in 3.11; before that, `locale.getpreferredencoding(False)`). On a typical Linux or macOS box that is UTF-8, on Windows it is the ANSI code page such as `cp1252`, and under a `C`/`POSIX` locale the interpreter falls back to UTF-8 via locale coercion or UTF-8 mode. So the same source file reads different characters on different hosts, which is how mojibake and `UnicodeDecodeError` appear only in CI or in a container. The fix is to say what you mean: `encoding="utf-8"` for data you own, or `encoding="locale"` (3.10+) when reading the user's environment on purpose. Note that `sys.getdefaultencoding()` is a different thing — it is always `'utf-8'` and governs `str`/`bytes` conversions, not files.
code
python · 17 linesimport locale, os, tempfile
print("locale encoding:", locale.getencoding())
print("str/bytes default:", __import__("sys").getdefaultencoding())
path = os.path.join(tempfile.gettempdir(), "note.txt")
with open(path, "w", encoding="utf-8") as f:
f.write("caf\u00e9")
with open(path) as f: # decodes with the locale encoding
print("implicit:", f.read(), f.encoding)
with open(path, encoding="locale") as f: # 3.10+: the same thing, written down
print("explicit locale:", f.read())
with open(path, encoding="utf-8") as f: # portable and correct
print("explicit utf-8:", f.read())go deeper
Be ready to say in one sentence that a text-mode open() with no encoding= uses the locale encoding, not UTF-8, and that the habit is to pass encoding="utf-8" for files your program owns.
Explain where the locale encoding comes from on Unix versus Windows, name locale.getencoding() as the way to read it, and separate it cleanly from sys.getdefaultencoding() and sys.getfilesystemencoding().
An interviewer expects you to connect the default to real incidents: mojibake or UnicodeDecodeError that appears only on a build agent or in a container, and a remediation plan that makes the encoding explicit rather than tweaking the host's locale.
Own the policy question. Decide whether the codebase declares encodings at every call site, standardises on process-wide UTF-8 mode, or both, and how that decision is enforced for new code and for dependencies you do not control.
### The call that says nothing `open(path)` in text mode returns an `io.TextIOWrapper` layered over a buffered binary file. That wrapper needs an encoding, and if the call does not supply one, CPython asks the *environment* rather than the file. The value it uses is the **locale encoding**: on Unix, whatever the C library reports for the `LC_CTYPE` category at the moment the stream is created; on Windows, the process's ANSI code page (`cp1252` on a typical US/Western European install, `cp932` on a Japanese one). Since 3.11 you can read that value directly with `locale.getencoding()`; on older versions the idiom was `locale.getpreferredencoding(False)`. The consequence is that the encoding is a property of the *machine*, not of the code and not of the file. A module that reads a data file with `open(path)` decodes it as UTF-8 on a developer's laptop, as `cp1252` on a Windows build agent, and as UTF-8 again inside a minimal container — or as ASCII on an old image whose locale is unset. Nothing in the source line changes; only the ambient locale does. ### What the locale encoding actually is On Unix, CPython reads the codeset of the current `LC_CTYPE` locale, which is driven by `LC_ALL`, `LC_CTYPE` and `LANG` in that precedence order. If the locale is `C` or `POSIX` — common in containers, cron jobs and CI runners — the codeset is ASCII, which would make Python unable to read its own UTF-8 source data. Since 3.7 the interpreter defends against this twice: it first tries **locale coercion** (PEP 538), re-setting the locale to a UTF-8 variant such as `C.UTF-8`, and if that is impossible it enables **UTF-8 mode** (PEP 540), which forces UTF-8 regardless of the locale. On macOS the interpreter always uses UTF-8 for the filesystem encoding, and Windows separates the console (which speaks UTF-16 through the Win32 API) from files (which still get the ANSI code page unless UTF-8 mode is on). ### The look-alikes Three functions get confused with each other, and interviewers probe the confusion: * `locale.getencoding()` — the locale encoding; **this** is the default for `open()`. * `sys.getdefaultencoding()` — always `'utf-8'` on Python 3. It is the implicit encoding for `str`/`bytes` conversions, and it has nothing to do with files. Answering with this is a classic wrong answer. * `sys.getfilesystemencoding()` — how *path names* are encoded and decoded when crossing the OS boundary, not how file *contents* are decoded. ### It is not only `open()` Every stdlib helper that opens a text file for you inherits the same default: `pathlib.Path.read_text()` and `write_text()`, `configparser.ConfigParser.read()`, and `subprocess.run(..., text=True)` for a child process's pipes. Anything that takes an already-open text object — `json.load(fp)`, `csv.reader(fp)` — simply inherits whatever decoding that object was built with, so the bug is upstream of them. Binary mode has no such default at all: `open(path, "rb")` hands you `bytes` and never decodes, and passing `encoding=` alongside `"rb"` raises `ValueError`. That is the right mode when the exact bytes matter and you intend to decode yourself. ### Saying what you mean There are three honest positions, and the point is to pick one explicitly: 1. **The data has a known encoding.** Pass `encoding="utf-8"`. This is correct for almost every file your own program wrote, every config file in your repository, and every network-sourced payload with a declared charset. 2. **The data belongs to the user's environment.** Pass `encoding="locale"` (accepted since 3.10). It is exactly the old behaviour, but written down — and it suppresses `EncodingWarning` when implicit-encoding auditing is turned on. 3. **You do not want text at all.** Open in binary and decode deliberately. Because the default is invisible, teams usually make it visible with tooling instead of vigilance: run the test suite with the interpreter's implicit-encoding warning enabled, or set UTF-8 mode process-wide so that at least the behaviour is the same everywhere. Either way, the interview answer that matters is the first sentence: `open()` with no `encoding=` reads your environment, not your file.
- Is sys.getdefaultencoding() the value that open() uses?No. `sys.getdefaultencoding()` returns `'utf-8'` on every Python 3 build and describes the implicit `str`/`bytes` conversion encoding; it never varies with the locale and never controls file decoding. The value `open()` uses with no `encoding=` is the locale encoding, which `locale.getencoding()` reports and which differs by platform and environment.
- What does encoding="locale" mean, and why would you ever write it?Accepted by `open()` since 3.10, it explicitly requests the current locale encoding — identical runtime behaviour to omitting the argument, but it records that the choice was deliberate. It is the right value when you are genuinely reading text produced by the user's environment, such as a console-generated file, and it silences `EncodingWarning` when implicit-encoding auditing is enabled.
- Does sys.getfilesystemencoding() have anything to do with this default?No. It reports how path names are encoded when crossing into the operating system, not how a file's contents are decoded. The two can differ: on Windows the filesystem encoding is UTF-8 while the default for `open()` is the ANSI code page. Mixing them up leads people to "fix" a decoding bug by changing something that only affects filenames.
It is like reading a letter in whatever language is spoken where you happen to be standing, instead of the language the letter was written in.
saying these in an interview costs you the question
- Says open() always defaults to UTF-8
- Answers with sys.getdefaultencoding(), which is always utf-8
- Thinks the file's own bytes determine the decoding
- Believes encoding= only matters on Windows
- Claims binary mode still decodes with a default encoding
- Blames the data file when the same code works locally