How would you set a Python-wide policy for codec error handlers at a system's boundaries?
answer
- One decision per boundary, not per call
- Who should learn that data is bad
- OS values are handles, not text
- One handler should never be allowed
- A registry hook makes loss countable
basics
~20 sAssign a handler per boundary by audience: 'strict' wherever you control the producer, 'surrogateescape' only for OS-supplied values that must round-trip, lossy handlers only on human-facing output, and 'ignore' nowhere. Register a custom handler when you need lossiness counted.
solid answer
~50 sTreat the handler as a boundary contract rather than a per-call convenience. Ingress you control gets `'strict'`, so bad input is reported to whoever produced it instead of being absorbed. OS-facing values — paths, `sys.argv`, environment — get `'surrogateescape'`, but only as a short-lived carrier that is converted before it enters storage or a message. Human-facing egress, meaning logs, error text and display, gets `'replace'` or `'backslashreplace'`. `'ignore'` is banned outright: it deletes bytes with no signal, shifts offsets, and can strip exactly the characters a later validator was meant to catch. Where you need lossiness to be measurable, `codecs.register_error` lets you install a named handler that increments a counter and then substitutes, so silent degradation becomes a metric. The organisational half matters as much: one boundary helper module, a review rule that a bare `errors='ignore'` is a defect, and tests that feed deliberately invalid bytes.
code
python · 14 linesimport codecs
replaced = 0
def counting_replace(exc):
global replaced
if isinstance(exc, UnicodeDecodeError):
replaced += 1
return ("�", exc.end)
raise exc
codecs.register_error("counting_replace", counting_replace)
text = b"a\xffb\xfec".decode("utf-8", errors="counting_replace")
print(ascii(text), replaced)go deeper
Recall that the handler choice depends on where in the program the text is crossing, and that the default is to raise. You are not expected to design the policy, only to follow the boundary helper the codebase provides.
Explain why lossy handling on stored or re-parsed text is worse than an exception, and how a custom handler is registered by name so a boundary can substitute and count at the same time.
Show that you would place handlers by boundary and audience, quarantine bad records instead of absorbing them, and instrument every lossy path so degradation appears in metrics rather than in a customer report.
Own the accountability question: which boundary is answerable for bad bytes, what is centralised so the policy is reviewable, and how you prevent a convenience default from quietly turning failures into corrupted content across services.
At a small scale, choosing an error handler looks like a local decision. At system scale it is a **contract**, because the handler decides who learns about bad data and when. Every lossy handler converts a loud, local failure into a quiet change of content that some other component discovers later — often in a different process, a different team's service, or a customer-visible field. A policy is therefore about placing failure deliberately rather than about picking a favourite handler. ## Ingress you control gets `'strict'` If another service, an internal producer or your own upload path sends bytes that do not decode, that is a defect in the producer, and absorbing it destroys the only signal that would get it fixed. Failing loudly at the edge also keeps the blast radius one record wide if you quarantine rather than abort: 1. reject the message, 2. record the offending bytes, 3. continue. The rejected payload is evidence; a `'replace'`-ed one is not. ## OS-supplied values get `'surrogateescape'`, with a containment rule Filenames, `sys.argv` and environment variables on Unix are bytes with no declared encoding, and Python already decodes them this way. That is correct — the round trip through `os.fsencode` is what lets you reopen a badly named file. What must be written down is that such a string is **a handle, not text**: - it may be passed back to OS APIs, - and it must be converted before it enters a queue, a database column, a serialised message or an outbound response. Exporting a lone surrogate into a message body means every downstream consumer inherits a decoding problem they did not create and cannot diagnose. ## Human-facing egress gets a lossy handler, and only there Logs, exception text, terminal output and UI strings must never fail; a logging path that raises during error handling is how an incident loses its own evidence. - `'backslashreplace'` is the strongest choice for diagnostics because byte values survive as ASCII; - `'replace'` is fine for display where a marker is enough. The discipline is that this applies to output that will be read by a person and never re-parsed by a machine. Text that is stored, forwarded or re-parsed keeps `'strict'`. ## `'ignore'` is banned It is the only handler with **no trace whatsoever**: no exception, no marker, no length preservation. Beyond losing data it has a security dimension, because deleting bytes before validation changes what a validator sees relative to what a downstream component will act on. Any code review that finds `errors='ignore'` should treat it as a defect requiring justification, not a style preference. ## Make lossiness measurable `codecs.register_error(name, handler)` installs a handler under your own name; a handler receives the exception instance and returns a `(replacement, resume_index)` pair, where the replacement is a `str` when decoding, and the resume index is normally `exc.end`. That hook is the right place to: - increment a counter, - sample the offending bytes into a log, - or apply a domain-specific substitution. The registry is **process-global** and the name must be registered before any call uses it, which in practice means registering it during application start-up, in the same module that owns the boundary helpers. A handler that raises — or re-raises the exception it was given — restores strict behaviour for cases it does not want to absorb, which is how you write a handler that only handles decoding failures and refuses encoding ones. ## Centralise the decision In a codebase of any size, the practical failure mode is not a wrong policy but a hundred call sites that each made their own. One module owning “decode inbound”, “render for logs” and “decode OS value” gives you a single place to change the policy, a single place to instrument it, and something concrete for a review checklist to point at. Pair it with tests that feed deliberately invalid bytes through each boundary and assert the expected behaviour, because these paths are otherwise exercised only in production: - a truncated multi-byte sequence, - a byte no codec accepts, - a name carrying an escaped byte. The judgement an interviewer is listening for is that the handler choice is not about avoiding exceptions. It is about deciding, per boundary, **who is accountable for bad bytes and how visible that is** — and about noticing that the cheapest way to make a system quietly wrong is to standardise on a handler that never complains.
- What does a function registered with codecs.register_error have to return?A two-tuple of a replacement and the index at which decoding or encoding resumes — normally `exc.end`, so the offending slice is consumed. When decoding, the replacement is a `str`. The handler receives the exception instance, so it can inspect `object`, `start`, `end` and `reason`, and re-raising it restores strict behaviour for cases it chooses not to absorb.
- Why is errors='ignore' worth an outright ban rather than case-by-case judgement?It is the only handler that leaves no evidence: no exception, no marker, no preserved length. Case-by-case judgement fails in practice because the harm appears far from the call site, and because deleting bytes before validation changes what a validator sees relative to what a downstream component acts on. A blanket rule is cheaper to review than an argument per occurrence.
- How do you keep a policy like this from decaying across a large codebase?Put the decisions behind named boundary helpers — decode inbound, render for logs, decode OS value — so the policy exists as code rather than as documentation. Add a review rule for bare handler arguments, register any custom handler at start-up in that same module, and test each boundary with deliberately invalid bytes so the paths are exercised outside production.
It is a returns policy rather than a shop-floor judgement call: the interesting question is not what to do with one damaged item but which counter is allowed to accept damage at all.
saying these in an interview costs you the question
- Standardising on one handler for the whole codebase
- Treating handler choice as a way to avoid exceptions
- Letting surrogate-bearing strings into storage or messages
- Applying lossy handling to text that is later re-parsed
- Adding lossy handling with no counter or log
- Assuming a custom handler works without registering it first