An email-digest sender intermittently raises UnicodeEncodeError on some subject lines — how do you diagnose and fix it?
answer
- Intermittent points at data, not load
- The exception object names the culprit
- One field splits the diagnosis in two
- Escaped bytes mean the bug is upstream
- Lossy handling belongs on log output only
basics
~20 sCatch the exception and read encoding, reason, start and object[start:end]: that names the exact character and the target codec. Intermittent means data-dependent, not load-dependent, so reproduce with those characters rather than retrying at higher volume.
solid answer
~50 sStart from the exception object, not the stack trace. `UnicodeEncodeError` carries `encoding`, `object`, `start`, `end` and `reason`, and `exc.object[exc.start:exc.end]` is the offending character; print it with `ascii()` so the terminal cannot mangle the evidence. The `reason` splits the diagnosis in two. `ordinal not in range(128)` and its kin mean the target codec is too narrow for legitimate text, so the fix is the encoding at that boundary, not a handler. `surrogates not allowed` means the string already carries escaped bytes from an earlier decode, so the real defect is upstream and no handler at the write site is correct. “Intermittent” is a clue in itself: encoding failures are deterministic in the data, so the variable is which records a given run touched. Only after the boundary is right do you add a lossy handler — `'backslashreplace'` for logs, never `'ignore'` on the payload itself.
code
python · 10 linessubject = "digest for \udcff user"
try:
subject.encode("utf-8")
except UnicodeEncodeError as exc:
print(exc.encoding, ascii(exc.object[exc.start:exc.end]), exc.start, exc.reason)
def log_safe(text):
return text.encode("utf-8", errors="backslashreplace").decode("ascii")
print(log_safe(subject))go deeper
Know that this exception means the target codec cannot represent a character in the string, and that the exception object itself says which character. Do not guess from the traceback alone.
Be able to catch it and print encoding, reason and object[start:end], and to say why widening the codec at that boundary is a different fix from adding an error handler.
Demonstrate the full triage: rule out load as a cause, split on reason, trace escaped bytes back to the decode that planted them, and place lossy handling on diagnostics only, with a regression test built from the captured string.
Own the boundary contract across services: state where text is validated, where lossy rendering is permitted, and how degradation is counted, so a single mis-encoded field cannot quietly reach subscribers.
## The failure is data-dependent The first thing to establish is that a `UnicodeEncodeError` is **never random**. Encoding a given string with a given codec either works or fails, every time, on every machine. When a service fails “intermittently” the variable is the *input*, and the investigation is therefore about which records reached the failing line rather than about timing or load. In a digest sender that is easy to demonstrate: if most subject lines are served from a cache of previously rendered text — say an 83% cache-hit rate — then only the freshly rendered minority ever traverses the encode path with new user data in it, and the failure rate tracks the miss rate rather than traffic. That reframing alone rules out the tempting but wrong theories: - it is not an intermittent timeout to the mail relay, - and it is not a race between workers. ## Interrogate the exception object Step two is to make the exception talk. `UnicodeEncodeError` is a `UnicodeError` and therefore a `ValueError`, and the instance carries the same five attributes as its decoding counterpart: - `encoding` names the target codec, - `object` is the whole `str` being encoded, - `start` and `end` bound the failing slice, - and `reason` explains it. Wrap the failing boundary in a handler that logs `exc.encoding`, `exc.reason`, `exc.start` and `ascii(exc.object[exc.start:exc.end])`. Using `ascii()` matters: printing the raw character to a log stream that itself cannot encode it produces a second `UnicodeEncodeError` inside the error handler, which is how these incidents turn into empty log lines. ## Two kinds of failure Step three is to read `reason`, because it partitions the problem cleanly. ### Real text, too-narrow codec If the reason is of the `ordinal not in range(128)` family, the string is legitimate text and the codec at that boundary is too narrow — an em dash or an accented name headed into an ASCII or legacy encoder. The correct fix is at the boundary: encode with a codec that covers the text. Reaching for `errors='replace'` here downgrades a subscriber's name to question marks in an outbound message, which is a product defect, not a fix. If some part of the pipeline genuinely cannot carry the characters — a protocol field with a restricted charset — then the transformation should be explicit and appropriate to that field's rules, applied deliberately rather than as a codec error handler bolted onto a generic write. ### Escaped bytes, not text If the reason is `surrogates not allowed`, the diagnosis is completely different. The string contains lone surrogates in U+DC80–U+DCFF, which means some earlier decode used `errors='surrogateescape'` — very often implicitly, by reading a filename, an environment variable or a command-line argument. No handler at the write site is the right answer, because the string is not text: it is a carrier for raw bytes. Walk back to the boundary that produced it and decide there whether to keep the bytes, reject the record, or convert it to a printable form. ## Where a lossy handler belongs Step four is placement of any lossy handler you do add. Split the boundaries by audience. - **For diagnostics**, a helper such as `text.encode("utf-8", errors="backslashreplace").decode("ascii")` guarantees a log line can always be written and keeps the byte values visible, so the next occurrence is self-describing. - **For the outbound payload**, keep `'strict'` and let the failure be loud, because a silently mangled subject line reaches a person. - `'ignore'` deserves particular suspicion: it removes characters with no signal at all, so a run that “succeeded” may have shipped a truncated name that nobody notices for months. ## Prevention Step five is to keep the finding from recurring. 1. Add a **regression test** that feeds the exact offending string — the one you recovered from `exc.object` — through the same boundary; unlike the production incident, that test is fully deterministic. 2. Emit a **counter** whenever a lossy handler actually fires, so degradation is measurable rather than invisible. 3. And record the **encode/decode contract** for each boundary in the code that owns it, since these bugs are almost always a boundary whose codec was never chosen consciously. ## The shape of the answer The answer that lands in an interview is the shape of the reasoning: - exceptions of this class are data-dependent, - the exception object identifies the exact character and codec, - `reason` distinguishes “wrong codec for real text” from “string is carrying escaped bytes”, - and error handlers belong on human-facing output, never on the payload.
- Why does 'intermittent' point at the data rather than at load?Encoding is deterministic: the same string and codec always give the same result. So variability comes from which records a run touched. Caching, sampling or partitioning can make the bad path rare — if most subjects are served from cache, only the freshly rendered minority reaches the encode, and the failure rate tracks misses rather than request volume.
- What tells you the failure belongs upstream rather than at the failing write?The `reason` on the exception. `surrogates not allowed` means the string carries lone surrogates planted by an earlier decode with `'surrogateescape'`, typically from a filename, environment variable or argv. No handler at the write site is correct; the boundary that produced the string must decide whether to keep the bytes, reject the record or convert it.
- How do you keep the error handler in your logging path from failing too?Render defensively before logging: `text.encode("utf-8", errors="backslashreplace").decode("ascii")` always produces pure ASCII, so the log write cannot itself raise. Printing the raw offending character into a stream that cannot encode it raises a second UnicodeEncodeError inside the handler, which is how incidents end up with empty or truncated log lines.
saying these in an interview costs you the question
- Treating an intermittent encode failure as a load or timing issue
- Adding errors='ignore' to the outbound payload to stop the crash
- Logging the raw offending character and crashing again
- Ignoring exc.reason, which distinguishes the two root causes
- Assuming the writing code is at fault when surrogates are involved