In Ruby, when does String#encode raise Encoding::UndefinedConversionError instead of Encoding::InvalidByteSequenceError, and which options suppress each?
answer
- valid character vs broken bytes
- both inherit from EncodingError
- undef: :replace and invalid: :replace
- replace: defaults to ? or U+FFFD
- error_char vs error_bytes
basics
~20 sUndefinedConversionError means a valid source character has no counterpart in the target, such as € into ISO-8859-1. InvalidByteSequenceError means the source bytes are broken for their own encoding. undef: :replace and invalid: :replace substitute a replacement string instead.
solid answer
~40 s`encode` raises by default on both problems, and both classes inherit from `EncodingError`, a `StandardError`. **Undefined**: the source is fine but the destination lacks the character; exporting `"€5"` to a legacy Latin-1 system raises `U+20AC from UTF-8 to ISO-8859-1`, and the exception's `error_char` returns `"€"`. **Invalid**: the source bytes are broken for their own label, like a lone `0xE9` in a UTF-8 string, and `error_bytes` shows them. `undef: :replace` and `invalid: :replace` swap in the `replace:` string, which defaults to `"?"` for a non-Unicode target and U+FFFD for a Unicode one. `fallback:` takes a Hash or Proc for per-character mappings such as `"€" => "EUR"`. One surprise: encoding a binary string with bytes above 0x7F to UTF-8 raises the undefined error, not the invalid one.
code
ruby · 10 linesLEGACY_MAP = { "€" => "EUR" }
def to_legacy(text, record_id)
text.encode("ISO-8859-1")
rescue Encoding::UndefinedConversionError => e
warn "record=#{record_id} char=#{e.error_char.dump} #{e.source_encoding} -> #{e.destination_encoding}"
text.encode("ISO-8859-1", fallback: ->(ch) { LEGACY_MAP.fetch(ch, "?") })
end
to_legacy("Zoë paid €5", 42) # logs char="\u20AC", returns "Zo\xEB paid EUR5"go deeper
Recall which error means broken input and which means a character the target cannot hold, and that both are EncodingError subclasses.
Explain the invalid:, undef:, replace: and fallback: options, the default replacement for Unicode and non-Unicode targets, and the binary-to-UTF-8 surprise.
Choose an export policy per field: fail loudly for names and money, map with fallback: where meaning survives, and log error_char with the record id.
Weigh lossy export against rejecting records, and decide who owns widening a legacy system to UTF-8 versus maintaining conversion rules forever.
## Two different ways a conversion fails `String#encode` reads characters in the source encoding and writes them in the destination encoding. Each half of that can fail: - **Reading fails**: the source bytes are not a valid sequence in the source encoding. Ruby raises `Encoding::InvalidByteSequenceError`. - **Writing fails**: the source character is valid, but the destination encoding has no way to represent it. Ruby raises `Encoding::UndefinedConversionError`. Both classes live under `Encoding`, inherit from `EncodingError`, and through it from `StandardError`, so a plain `rescue` catches them. A third sibling, `Encoding::ConverterNotFoundError`, means Ruby has no converter between the two encodings at all. | | `UndefinedConversionError` | `InvalidByteSequenceError` | |---|---|---| | Problem is in | the destination | the source | | Example | `"€".encode("ISO-8859-1")` | `"caf\xE9".encode("ISO-8859-1")` | | Extra reader | `error_char` | `error_bytes`, `incomplete_input?` | | Option that suppresses it | `undef: :replace` | `invalid: :replace` | Both also expose `source_encoding` and `destination_encoding`, which make log lines precise. ## The options `encode` accepts keyword options: - **`invalid: :replace`** replaces each invalid source sequence with the replacement string. - **`undef: :replace`** replaces each character that the destination cannot hold. - **`replace: "…"`** sets the replacement string. Without it, Ruby uses `"\uFFFD"` (�) for a Unicode destination and `"?"` otherwise. Given on its own, it enables replacement of undefined characters but not of invalid bytes. - **`fallback:`** takes a Hash, Proc or Method that maps an undefined character to its substitute, so `"€"` can become `"EUR"` instead of `"?"`. It is consulted only when `undef: :replace` is absent: with both given, the replacement string wins. A Hash that has no entry (and no default) for a character leaves the error raised. - **`xml: :text`** and **`xml: :attr`** escape `&`, `<` and `>` and write undefined characters as numeric character references such as `€`. ```ruby export = "Zoë paid €5" export.encode("ISO-8859-1") # Encoding::UndefinedConversionError: U+20AC from UTF-8 to ISO-8859-1 export.encode("ISO-8859-1", undef: :replace) # "Zo\xEB paid ?5" export.encode("ISO-8859-1", fallback: { "€" => "EUR" }) # "Zo\xEB paid EUR5" "caf\xE9".encode("ISO-8859-1", invalid: :replace) # "caf?" ``` `ë` survives because ISO-8859-1 contains it; only `€` is undefined there. ## The binary surprise A string labelled `ASCII-8BIT` (alias `BINARY`) has no invalid byte sequences: every byte is valid. Bytes above `0x7F` simply have **no character meaning**, so converting them to UTF-8 is an undefined conversion: ```ruby "\x80".b.encode("UTF-8") # Encoding::UndefinedConversionError: "\x80" from ASCII-8BIT to UTF-8 ``` Candidates who expect `InvalidByteSequenceError` here are thinking of the bytes as broken text rather than as bytes with no text meaning. The fix is to state what the bytes really are, for example `encode("UTF-8", "ISO-8859-1")`, not to add `invalid: :replace`, which does not apply. ## Reading the exception The exception objects carry enough detail to log a precise line instead of a stack trace: - `UndefinedConversionError#error_char` returns the character that could not be converted, such as `"€"`. - `InvalidByteSequenceError#error_bytes` returns the broken bytes, and `readagain_bytes` returns bytes that were read ahead and will be retried. - `InvalidByteSequenceError#incomplete_input?` returns `true` when the input simply ended in the middle of a character, as in a truncated upload; the message then reads like `incomplete "\xE9" on UTF-8`. - Both classes provide `source_encoding_name` and `destination_encoding_name`. Messages follow a fixed pattern: `U+20AC from UTF-8 to ISO-8859-1` for an undefined character, and a form such as `"\xE9" followed by "x" on UTF-8` for an invalid sequence in the middle of the text. ## Choosing a policy when exporting 1. **Decide whether loss is acceptable.** Invoices and legal names usually are not; prefer failing loudly and fixing the data. 2. **If loss is acceptable, prefer `fallback:` over `?`.** A mapping like `"€" => "EUR"` keeps meaning; a question mark does not. 3. **Rescue narrowly and log `error_char` with the record id**, plus `source_encoding` and `destination_encoding`, so the offending character is visible without dumping personal data. 4. **Never reach for `invalid: :replace` to silence an undefined-character error**: it covers the other failure and leaves the exception in place. ## Common misreadings - "Both errors mean the input is corrupt." Only the invalid one does; the undefined one means the target is too small. - "The replacement is always �." It is `?` whenever the destination is not a Unicode encoding. - "`invalid: :replace` handles everything." It only covers broken source bytes.
- In Ruby, what does String#encode's replace: option do when given without invalid: or undef:?It sets the replacement string and switches on replacement for undefined characters only. `"€5".encode("ISO-8859-1", replace: "?")` returns `"?5"`, but a string with invalid source bytes still raises `Encoding::InvalidByteSequenceError` unless `invalid: :replace` is also passed.
- In Ruby, what does "€ & <b>".encode("US-ASCII", xml: :text) return?`"€ & <b>"`. The `xml: :text` option escapes `&`, `<` and `>` and writes each character the destination cannot hold as an upper-case hexadecimal character reference, so nothing raises and nothing is lost for an XML consumer.
saying these in an interview costs you the question
- Both encoding errors mean the input file is corrupt
- UndefinedConversionError means the source bytes are malformed
- invalid: :replace also covers characters the target encoding lacks
- The replacement character is always U+FFFD, whatever the target encoding
- Converting BINARY bytes above 0x7F to UTF-8 raises InvalidByteSequenceError