Stripping a byte-order mark with text[1:] truncated real data in a 6,800-row geocoding batch — how should a Python ingest remove a BOM safely?
answer
- Blind slicing eats real characters
- Remove only when it is actually there
- Decode once, at the boundary
- removeprefix, not [1:]
- Fixture with and without the mark
basics
~20 sDo not slice blindly. Decode with the utf-8-sig codec, which removes a leading mark only when one is there, or guard the removal with str.removeprefix('\ufeff'). Handle it once, at the decode boundary, and keep a no-mark fixture in the tests.
solid answer
~40 sThe bug is a fix that assumes its own precondition: `text[1:]` removes the first character whether or not it is U+FEFF, so a partner file without a mark loses a real character from its first field — a silent truncation that no exception announces. The right removal is conditional. Best is to never materialise the mark: decode the bytes with `utf-8-sig`, which strips one leading signature when present and behaves exactly like `utf-8` when absent. If you already hold a `str`, use `text.removeprefix("\ufeff")`, or guard with `str.startswith`. For mixed producers, sniff the first bytes against `codecs.BOM_UTF32_LE`/`_BE`, then `codecs.BOM_UTF16_LE`/`_BE`, then `codecs.BOM_UTF8`, and pick the codec from that. Do the whole thing once, at ingest, and hand clean `str` downstream.
code
python · 7 linesmarked = "\ufeffid,lat"
clean = "id,lat"
print(repr(marked.removeprefix("\ufeff"))) # 'id,lat'
print(repr(clean.removeprefix("\ufeff"))) # 'id,lat' (nothing to remove)
print(repr(clean[1:])) # 'd,lat' (real character lost)
print(repr(clean.strip())) # 'id,lat' (strip never touches U+FEFF)go deeper
Take away one rule: never remove a character you have not checked for. str.removeprefix or a str.startswith guard, and better still let the utf-8-sig codec do it during decode.
Explain why the slice is unsafe and why utf-8-sig is not — it strips only when the signature is present and is otherwise a plain UTF-8 decode. Know that str.strip() will not remove U+FEFF.
An interviewer wants the diagnosis and the containment: a silent truncation with no exception, one decode boundary rather than scattered strips, byte sniffing for mixed producers, and fixtures covering both marked and unmarked input.
Own the ingest contract across producers — which encodings are accepted, who normalizes, and how a producer changing its exporter is detected rather than absorbed. Recording per-source signature behaviour turns a mystery truncation into an attributable change.
## What actually went wrong The first fix for a stray byte-order mark is nearly always `text[1:]` or `line[1:]`. It works on the file that prompted it and is wrong in general, because slicing removes *a* character, not *the mark*. Feed it a 6,800-row batch from a partner whose exporter does not write a signature and the first field of the first row quietly loses its leading character — `id` becomes `d`, an address `12 Bridge St` becomes `2 Bridge St`. Nothing raises. The batch runs green, the geocoder resolves fewer rows or resolves them to the wrong place, and the damage surfaces days later as a truncation nobody can reproduce because the file that reproduces it is the *clean* one. The general lesson: a removal must be conditional on the thing being present. Everything below is a way of making it conditional. ## Fix it at the decode boundary The cleanest answer is to never let the mark become a character. Decoding with `utf-8-sig` removes exactly one leading signature if it is there and decodes byte-for-byte identically to `utf-8` if it is not. It is idempotent in the sense that matters: safe on both kinds of file, so the ingest path needs no branch and no flag. That is why signature handling belongs at the edge where bytes become text, not sprinkled through the consumers. One function turns bytes into `str`; everything downstream receives text that is already clean; no parser, no key lookup, no comparison has to know that byte-order marks exist. ## If you already hold a str Sometimes the decode happened somewhere you do not control. Then the safe removals are: - `text.removeprefix("\ufeff")` — removes it if present, returns the string unchanged otherwise. Available since Python 3.9. - an explicit guard: `text[1:] if text.startswith("\ufeff") else text`. And the unsafe ones worth naming so you can reject them in review: - `text[1:]` — the truncation bug above. - `text.strip()` or `text.lstrip()` — U+FEFF is a format character in category `Cf`, not whitespace, so these do nothing to it while happily eating leading spaces you may have wanted. - `text.replace("\ufeff", "")` — removes marks *anywhere*, including a U+FEFF that is legitimate data in the middle of a document. - stripping per line rather than per file — a signature can only be at the very start of the stream; a per-line strip is a licence to corrupt every line. ## Mixed producers: sniff the bytes When files arrive from several sources, one partner eventually sends UTF-16. Read the first four bytes and dispatch on them: test `codecs.BOM_UTF32_LE` and `codecs.BOM_UTF32_BE` first, then `codecs.BOM_UTF16_LE` and `codecs.BOM_UTF16_BE`, then `codecs.BOM_UTF8`, falling back to `utf-8-sig`. The order matters: the UTF-32 little-endian mark begins with the UTF-16 little-endian mark, so testing two bytes first mislabels every UTF-32 file. Note also that a mark is *evidence*, not a guarantee — most UTF-8 files carry none, so the fallback has to be a real decision rather than a shrug. ## The write side The same batch usually emits files too. Default to plain `utf-8`; choose `utf-8-sig` only for an output whose named consumer needs the signature to pick an encoding. Two encode-side traps travel with that choice: `open()` builds a fresh incremental encoder per stream, so appending with `encoding="utf-8-sig"` plants a second signature mid-file, and calling `str.encode("utf-8-sig")` per chunk and concatenating gives one signature per chunk. Encode the document once. ## Making it stay fixed This class of bug returns because the test corpus contains only files that reproduce the original symptom. Pin both cases: - two fixtures, one with a signature and one without, pushed through the *same* ingest function; - assertions on exact equality of the parsed first key (`assert header[0] == "id"`), never on how the output looks — the character is zero-width, so a visual check passes on a corrupt file; - an assertion that a field which never had a mark comes back with its full length, which is the specific regression that slicing introduced; - when a partner file misbehaves in production, log `raw[:4]` in hex rather than the decoded text, since the decoded text is exactly what hides the problem. A useful operational habit is to record, per source, whether that producer emits a signature. It costs nothing, it turns "the batch truncated" into "this producer changed its exporter", and it makes the mixed-producer sniffing above auditable rather than magic.
- Where in the pipeline should signature handling live?At the single point where bytes become text. Decode with `utf-8-sig` (or a sniffing dispatcher for mixed producers) in one ingest function, and hand clean `str` to everything downstream. Spreading the concern into parsers and consumers guarantees an inconsistent one: some strip, some do not, and the ones that do usually strip unconditionally, which is how the truncation got in.
- How do you stop this regression coming back?Two fixture files through the same ingest path — one with a signature, one without — and assert exact equality of the parsed first key rather than eyeballing output, because a zero-width character passes any visual check. Add an assertion that a field which never carried a mark keeps its full length; that is precisely what the slicing fix broke, and nothing else in the suite would notice.
- A partner starts sending UTF-16 instead. What changes?The read path has to sniff rather than assume. Read the first four bytes and test `codecs.BOM_UTF32_LE`/`_BE`, then `codecs.BOM_UTF16_LE`/`_BE`, then `codecs.BOM_UTF8`, falling back to `utf-8-sig`. Test the four-byte marks first: the UTF-32 little-endian mark begins with the UTF-16 little-endian one, so a two-byte-first sniffer mislabels every UTF-32 file.
saying these in an interview costs you the question
- Slices text[1:] without checking for a mark
- Uses strip() and expects U+FEFF to go
- Replaces every U+FEFF in the whole document
- Strips per line instead of once per file
- Encodes each chunk with utf-8-sig and concatenates
- Ships with no fixture file lacking a mark