A Ruby job importing a legacy Latin-1 customer file raises ArgumentError: invalid byte sequence in UTF-8 on split; how do you diagnose and fix it?
answer
- the label lies, the bytes are fine
- Encoding.default_external labels what you read
- valid_encoding? returns false
- encode("UTF-8", "ISO-8859-1") at the boundary
- scrub is lossy; Windows-1252 lurks
basics
~20 sThe Latin-1 bytes were labelled with the default external encoding, UTF-8, so the string is invalid and split or any Regexp match raises. Confirm with encoding, valid_encoding? and bytes, then transcode at the boundary with encode("UTF-8", "ISO-8859-1").
solid answer
~40 sReading the file labels its bytes with `Encoding.default_external`, usually UTF-8, so a Latin-1 `é` (byte `0xE9`) sits in a string labelled UTF-8. Reading does not validate; `valid_encoding?` is `false`, and the first method that must decode characters, such as `split` or a Regexp match, raises `ArgumentError: invalid byte sequence in UTF-8`. Diagnose on one failing row with `encoding`, `valid_encoding?` and `bytes`. Fix it once, at the input boundary: `line.encode("UTF-8", "ISO-8859-1")`, or give the encoding when opening the file. ISO-8859-1 maps all 256 byte values, so that conversion never raises. `scrub` would only hide the bug by turning every `é` into U+FFFD. Check a sample, too: if the producer really wrote Windows-1252, its curly quotes and euro sign land on C1 control characters under a Latin-1 label.
code
ruby · 9 linesdef to_utf8(row)
return row if row.valid_encoding?
row.encode("UTF-8", "ISO-8859-1")
end
legacy = "Jos\xE9;Garc\xEDa".dup # Latin-1 bytes labelled UTF-8
to_utf8(legacy).split(";") # => ["José", "García"]
to_utf8("Zoë;Ng").split(";") # => ["Zoë", "Ng"], left alonego deeper
Recall that Ruby labels file data with a default encoding and that valid_encoding? tells you when the bytes do not fit that label.
Explain why the crash appears at split rather than at read, and show the encode call with a source-encoding argument that repairs the row.
Diagnose on one row, fix at the input boundary, verify output, and rule out Windows-1252 and double conversion before shipping the importer.
Push the contract upstream: agree an encoding with the producer, reject or quarantine rows that fail validation, and make encoding part of the import's documented interface.
## What actually went wrong A Ruby `String` is bytes plus an `Encoding` label. When an importer reads a file without saying what encoding it is in, Ruby labels the result with `Encoding.default_external`, which on a typical server is UTF-8. The legacy system wrote **ISO-8859-1** (Latin-1), where `é` is the single byte `0xE9`. In UTF-8, `0xE9` must start a three-byte sequence, so a lone `0xE9` makes the string **invalid**, but nothing checks that at read time. The failure surfaces later, in whichever method first needs to decode characters: - `String#split` and every Regexp match (`=~`, `match?`, `gsub` with a Regexp) raise `ArgumentError: invalid byte sequence in UTF-8`. - Other methods may quietly succeed, so the crash appears far from the read, often on the one customer whose name has an accent. ## Diagnosing on one failing row 1. **Reproduce with the exact row** that fails, not the whole file. 2. **Ask Ruby what it believes**: `row.encoding` shows the label, `row.valid_encoding?` returns `false`. 3. **Look at the bytes**: `row.bytes.find { it > 127 }` returning `233` (0xE9) next to ASCII letters is the signature of Latin-1 `é`. 4. **Confirm with the producer** which encoding the export uses. Ruby does not guess encodings; `valid_encoding?` only says the bytes do not fit the label they carry. ```ruby line = File.foreach("customers_legacy.txt").first line.encoding # => #<Encoding:UTF-8> (default_external) line.valid_encoding? # => false line.bytes.find { it > 127 } # => 233, a lone 0xE9 utf8 = line.encode("UTF-8", "ISO-8859-1") utf8.valid_encoding? # => true utf8.split(";") # no ArgumentError ``` ## Fixing it at the boundary - **Transcode, do not relabel to UTF-8.** `encode("UTF-8", "ISO-8859-1")` reads the bytes as Latin-1 and writes proper UTF-8. `force_encoding("UTF-8")` changes nothing, because the label already says UTF-8. - **Do it once, where data enters.** Either pass the external encoding when opening the file (the IO layer can transcode as it reads) or convert each record immediately after reading. Everything downstream then sees valid UTF-8 only. - **Expect no conversion errors from Latin-1.** Every one of the 256 byte values has a character in ISO-8859-1, so `encode` from it to UTF-8 does not raise `Encoding::UndefinedConversionError` or `Encoding::InvalidByteSequenceError`. - **Verify after converting.** Assert `valid_encoding?` on output and count converted rows in the job's log. ## Traps a senior engineer checks for | Symptom after the fix | Likely cause | Response | |---|---|---| | `“` and `€` show up as invisible control characters | producer wrote **Windows-1252**, not Latin-1 | use `encode("UTF-8", "Windows-1252")` | | `é` becomes `é` | a row was already UTF-8 and got converted again | convert only rows whose `valid_encoding?` is false | | accents replaced by `�` | someone added `scrub` | remove it; the bytes were never corrupt | Windows-1252 and ISO-8859-1 agree on `0xE9` but differ in `0x80`–`0x9F`: Windows-1252 puts curly quotes, dashes and the euro sign there, while ISO-8859-1 has C1 control characters. Characters U+0080–U+009F in converted output are the tell. For a file that mixes old Latin-1 rows with newer UTF-8 rows, a per-row rule works: keep rows that are already `valid_encoding?`, transcode the rest. It is a heuristic, since some Latin-1 byte pairs happen to form valid UTF-8, so prefer fixing the producer or splitting the file by origin. ## Preventing a recurrence 1. **Write the source encoding into the importer's configuration**, next to the file path, rather than relying on whatever `default_external` the server happens to have. 2. **Add a fixture row with accented names** (`José`, `Müller`, `Zoë`) to the importer's tests; ASCII-only fixtures are why this bug reached production. 3. **Validate at the boundary**: reject or quarantine rows whose converted text still fails `valid_encoding?`, and count them in the job's output. 4. **Log the row identifier and the offending byte**, never the whole customer record, so the next failure is diagnosable without leaking personal data. ## Why scrub is the wrong default `scrub` replaces each invalid sequence with a replacement string, U+FFFD for a Unicode encoding. It is the right tool for input that is genuinely corrupt, such as a truncated upload, where you accept losing the damaged bytes. Here the bytes were valid Latin-1 all along; scrubbing turns `José` into `Jos�` and destroys customer data silently. Use it only after transcoding, on rows you have decided to accept with losses, and log each one.
- Why is String#scrub the wrong default fix for a Ruby import of Latin-1 data?`scrub` replaces invalid sequences with U+FFFD for a Unicode encoding, so `José` becomes `Jos�`. The bytes were valid Latin-1; only the label was wrong. Transcoding with `encode("UTF-8", "ISO-8859-1")` keeps every character. Keep `scrub` for genuinely corrupt input you accept to lose, and log those rows.
- In a Ruby import, how do you tell whether the legacy file is Latin-1 or Windows-1252?Both map `0xE9` to `é`; they differ in `0x80`–`0x9F`. Windows-1252 puts curly quotes, dashes and `€` there, ISO-8859-1 has C1 controls. Convert a sample with `encode("UTF-8", "ISO-8859-1")` and look for characters U+0080–U+009F; if they appear where punctuation belongs, re-run with `"Windows-1252"` as the source.
saying these in an interview costs you the question
- The file is corrupt, so scrub every line and move on
- Calling force_encoding('UTF-8') on each line fixes the invalid bytes
- The ArgumentError means File.read failed to decode the file
- A UTF-8 magic comment in the importer script changes how file data is labelled
- Converting ISO-8859-1 to UTF-8 can raise UndefinedConversionError for some bytes