skip to content

A fraud-scoring service reads a cached rules file with open() and no encoding=, and one host decodes it as mojibake. How would you find every such call site before the next release?

level: seniorimportance: nice to knowfreq 18%

answer

  1. Do not grep for it by hand
  2. The interpreter can report the call sites
  3. A warning that is off by default
  4. Escalate it to an error in CI
  5. Only executed code is covered

basics

~10 s

Run the test suite under -X warn_default_encoding (or PYTHONWARNDEFAULTENCODING=1). Every open() or io.TextIOWrapper built without encoding= then raises an EncodingWarning at its own call site; escalate it to an error with -W error::EncodingWarning.

solid answer

~50 s

The interpreter has an auditing switch for exactly this: `-X warn_default_encoding`, or `PYTHONWARNDEFAULTENCODING=1`, added in 3.10 with PEP 597. With it on, every `open()` and `io.TextIOWrapper` constructed without an `encoding=` argument emits an `EncodingWarning` pointing at the caller, so a run of the test suite prints the inventory instead of you grepping for `open(`. Escalate with `-W error::EncodingWarning` in CI so new call sites cannot land, and check `sys.flags.warn_default_encoding` if code needs to know. Then fix each hit deliberately: `encoding="utf-8"` for files the service itself wrote — the cached rules blob is one — and `encoding="locale"` for the handful that genuinely read environment-produced text, which also silences the warning. Coverage is the catch: the warning only fires on code that actually executes, so pair it with a static sweep, and consider setting `PYTHONUTF8=1` in the deployment so untouched dependency code lands on UTF-8 too.

code

console · 4 lines
console
$ python3 -X warn_default_encoding -c "import io; io.TextIOWrapper(io.BytesIO(b'hi')).read()"
<string>:1: EncodingWarning: 'encoding' argument not specified
$ python3 -X warn_default_encoding -W error::EncodingWarning -m unittest discover
$ PYTHONWARNDEFAULTENCODING=1 python3 -m pytest -q

go deeper

for a junior

Know that Python can warn you about text files opened without an encoding argument, and that the fix is to pass encoding="utf-8" rather than to change the machine.

for a middle

Be able to name the interpreter option and environment variable that emit EncodingWarning, explain that they are off by default, and show how -W turns the warning into a CI failure.

for a senior

An interviewer expects the whole loop: confirm the mis-decode from the raw bytes, inventory the call sites at runtime, decide utf-8 versus locale per site, invalidate the poisoned cache, and gate regressions in CI.

for a principal

Own the standard rather than the incident: which repositories run the audit, whether deployments set UTF-8 mode, how the rule reaches dependencies you do not control, and how it is sequenced across a fixed release cadence.

### Why the bug shows up only on one host A text-mode `open()` with no `encoding=` decodes with the locale encoding, which is a property of the machine. The fraud-scoring service writes a cached rules blob on a host whose locale is UTF-8 and reads it back on a replica whose locale is not — or reads a cache file that a *previous* release wrote under a different locale. Either way the bytes are fine and the decoding is wrong, so the failure is silent: the file loads, the strings are subtly corrupted, and the stale cached value keeps scoring traffic until someone notices the accented merchant names. A crash would have been kinder. With a three-week release train, the expensive move is to fix the one call site you found, ship, and discover the next one three weeks later. The goal is an inventory in a single cycle. ### The auditing switch PEP 597 (3.10) added `EncodingWarning` and the interpreter option that emits it. Start the process with `-X warn_default_encoding`, or set `PYTHONWARNDEFAULTENCODING=1`, and every construction of a text stream without an explicit encoding warns — `open()`, `io.TextIOWrapper`, and stdlib call paths that build one for you. Crucially the warning is reported at the *caller's* line, not inside the io module, so the output is a list of files and line numbers you can work through. `sys.flags.warn_default_encoding` exposes the switch to code that wants to behave differently under audit. The warning is inert by default: without the option, `EncodingWarning` is never emitted, so there is no cost to leaving it off in production. In CI, escalate it with `-W error::EncodingWarning`, which turns each hit into a test failure and prevents regressions once you have cleaned up. ### Working the list Each hit deserves a decision, not a blanket edit: * **Data the service owns** — its cache files, its config files, its fixtures, anything it wrote itself. Pass `encoding="utf-8"`. This is the majority. * **Text produced by the user's environment** — output captured from a console tool, a file the operator hand-edited on a Windows box. Pass `encoding="locale"`, which is the same runtime behaviour as before, states the intent, and suppresses the warning. * **Not text at all** — anything where the exact bytes matter. Open in binary and decode explicitly where the knowledge lives. A cached artefact deserves one more step: a file written under a mis-decode is already poisoned, so plan invalidation. Versioning the cache key, or writing a small header the reader validates, converts "a stale cached value silently scores wrong" into "the cache misses and rebuilds". ### The limits of the technique The warning is a *runtime* signal. It fires only for code that executes, so its coverage is your test suite's coverage; a rarely taken error path stays invisible. Two complements are worth the effort: 1. **A static sweep.** A grep or a lint rule for text-mode opens with no encoding argument catches unexecuted code. Several third-party linters ship such a rule; the point is that the two techniques have complementary blind spots. 2. **Process-wide UTF-8 mode.** `PYTHONUTF8=1` in the deployment makes every remaining implicit open — including inside dependencies you will not be editing — default to UTF-8. It does not fix explicit-but-wrong encodings, and it turns previously silent mis-decodes into `UnicodeDecodeError`, which is why you want the audit *first*: you want to know what will start failing before it starts failing in production. ### Confirming the diagnosis first Before any of it, prove the theory rather than assuming it. Read the artefact in binary and inspect the bytes: a UTF-8 encoded `é` is the two bytes `0xC3 0xA9`, and text that decodes to a two-character sequence beginning with `Ã` is the signature of UTF-8 bytes read through a single-byte codec. Compare `locale.getencoding()` and `sys.stdout.encoding` between a healthy host and the sick one; the difference is usually a missing `LANG` in a service unit or a slimmer base image. Fixing only the host's locale is tempting and wrong — it leaves the same landmine for the next environment. Make the code state its encoding, and use the environment change only as immediate mitigation. ### What good looks like afterwards One CI job runs the suite with the warning escalated to an error, the deployment sets UTF-8 mode, every remaining implicit open is a conscious `encoding="locale"`, and the cache carries a version so a bad generation cannot outlive the release that produced it.

  • Why not just set the locale correctly on the affected host and move on?
    Because it fixes one machine, not the code. The next image, runner or operating system leaves the same landmine, and a service whose correctness depends on an environment variable has an undeclared dependency. Set the locale or UTF-8 mode as immediate mitigation if you must, then make the call sites explicit so the behaviour no longer varies by host.
  • What does the warning miss?
    Everything that does not run. It is emitted when a text stream is constructed, so its coverage equals your test coverage — rare error paths, optional features and lazily imported modules stay silent. Pair it with a static check for text-mode opens lacking an encoding argument, and treat process-wide UTF-8 mode as the safety net for dependency code you will not be editing.
  • How do you confirm that mojibake is a decoding problem rather than corrupted data?
    Read the artefact in binary and look at the bytes. Well-formed UTF-8 for an accented letter is a two-byte sequence; if those bytes are intact but the text shows a two-character sequence beginning with a capital A-tilde, the file is fine and the reader used a single-byte codec. Comparing `locale.getencoding()` between a healthy host and the failing one usually names the culprit.

saying these in an interview costs you the question

  • Greps for open( instead of using the interpreter's warning
  • Fixes only the host locale and calls it done
  • Thinks EncodingWarning is emitted by default
  • Adds errors='ignore' to make the traceback disappear
  • Assumes the warning covers code the tests never run
  • Leaves the poisoned cache entry in place

context