In Ruby, what is the difference between String#encode and String#force_encoding, and when is each the right call?
answer
- one changes bytes, one changes a label
- encode returns a new String
- force_encoding mutates, returns self
- relabel with the truth, then transcode
- encode(dst, src) does both at once
basics
~20 sString#encode returns a new string whose bytes are converted so the same characters exist in the target encoding. String#force_encoding keeps the bytes and only relabels the receiver. Relabel when the label is wrong; encode when bytes must change.
solid answer
~30 s`encode("ISO-8859-1")` **transcodes**: `"é"` goes from bytes `[195, 169]` in UTF-8 to `[233]` in Latin-1, and the receiver is untouched (`encode!` is the in-place form). By default it raises `Encoding::UndefinedConversionError` or `Encoding::InvalidByteSequenceError` when a character cannot be converted. `force_encoding` **relabels**: it never converts or validates, it rewrites the receiver's encoding tag and returns `self`, so on a frozen string it raises `FrozenError`. The classic pairing is bytes written as Latin-1 that arrived labelled UTF-8: `force_encoding("ISO-8859-1")` states the truth, then `encode("UTF-8")` converts, and `str.encode("UTF-8", "ISO-8859-1")` does both without mutating. Calling `force_encoding("UTF-8")` on those bytes changes nothing useful.
code
ruby · 9 liness = "é"
latin = s.encode("ISO-8859-1")
latin.bytes # => [233]
s.bytes # => [195, 169] receiver untouched
label_only = s.dup.force_encoding("ISO-8859-1")
label_only.bytes # => [195, 169] same bytes
label_only.length # => 2 read as two Latin-1 characters
label_only.equal?(label_only.force_encoding("ISO-8859-1")) # => true, returns selfgo deeper
Remember the one-line rule: encode changes bytes and returns a new string, force_encoding only changes the label on the same bytes.
Walk through the Latin-1 repair step by step, say that force_encoding returns self and mutates, and that same-encoding encode is a copy that validates nothing.
Show where the repair belongs, at the input boundary, and how to avoid double transcoding by converting only strings you know are mislabelled.
Argue for one internal encoding across the codebase and explicit conversion at every external edge, so encoding bugs surface at the edge instead of deep in business logic.
## Bytes and the label on them A Ruby `String` is a byte sequence plus an `Encoding`. Two different repairs exist because two different things can be wrong: - the **label** is wrong: the bytes are fine Latin-1 text, but Ruby believes they are UTF-8; - the **bytes** are wrong for where they are going: the text is correct UTF-8, but a legacy system needs Latin-1 bytes. `force_encoding` fixes the first problem and `encode` fixes the second. ## String#encode transcodes `encode(dst)` reads the receiver's characters using its current encoding and writes the same characters as bytes of `dst`, returning a **new** String. The receiver keeps its bytes and label; `encode!` is the mutating variant. - `"é".encode("ISO-8859-1").bytes` returns `[233]`, while the original `"é".bytes` stays `[195, 169]`. - `encode(dst, src)` interprets the receiver as `src` regardless of its label, then transcodes to `dst`. It is the non-mutating way to say "these bytes are really Latin-1". - A character missing from the target raises `Encoding::UndefinedConversionError`; broken source bytes raise `Encoding::InvalidByteSequenceError`. The `invalid:`, `undef:`, `replace:` and `fallback:` options change that. - When source and destination encodings are the same, `encode` is a plain copy: it does not validate or repair anything unless you pass `invalid: :replace`. ## String#force_encoding relabels `force_encoding(enc)` changes only the encoding tag of the **receiver** and returns `self`. The bytes are identical before and after. - It never raises for "bad" bytes. Relabelling `"\xE9"` as UTF-8 succeeds and simply yields a string whose `valid_encoding?` is `false`. - Because it modifies the receiver, it raises `FrozenError` on a frozen string. Call it on a copy (`dup`) or, for binary work, use `String#b`, which returns a relabelled copy. ## Side by side | | `encode(dst)` | `force_encoding(enc)` | |---|---|---| | Bytes | converted | unchanged | | Receiver | untouched, new String returned | modified, `self` returned | | Validates input | yes, raises by default | never | | Typical use | produce bytes a consumer needs | correct a wrong label | | In-place variant | `encode!` | is already in place | ## The relabel-then-transcode idiom Data read from a legacy Latin-1 source is usually labelled with `Encoding.default_external`, commonly UTF-8. The repair takes two steps: 1. **Tell the truth about the bytes**: `force_encoding("ISO-8859-1")`, or pass `"ISO-8859-1"` as the second argument to `encode`. 2. **Convert to what the application uses**: `encode("UTF-8")`. 3. **Check** that `valid_encoding?` is `true` on the result before letting it travel further. ```ruby raw = "caf\xE9".dup # Latin-1 bytes labelled UTF-8 raw.valid_encoding? # => false raw.force_encoding("ISO-8859-1") raw.encode("UTF-8") # => "café" "caf\xE9".encode("UTF-8", "ISO-8859-1") # => "café", receiver untouched ``` Doing only step 2 on the mislabelled string does nothing useful: `raw.encode("UTF-8")` while `raw` is still labelled UTF-8 is a same-encoding copy. Doing only step 1 with the wrong target, `force_encoding("UTF-8")`, just re-asserts the lie. ## Choosing in practice - **An HTTP body or file whose bytes are right but whose label is wrong**: relabel with `force_encoding` on a copy, or use `encode(dst, src)` to convert in one step. - **A consumer that needs different bytes**, such as a legacy system expecting ISO-8859-1 or a Windows tool expecting UTF-16LE: `encode(dst)`, with an explicit policy for characters the target lacks. - **Bytes that are not text at all**, such as a digest or a compressed payload: neither; take a binary copy with `b` and never transcode. - **Unknown input you merely want to make safe to display**: `scrub`, accepting that damaged sequences become U+FFFD. A useful interview habit is to say, before choosing, whether the bytes or the label is wrong. That one sentence decides the method. ## Traps - Believing `force_encoding` converts: it produces mojibake or invalid strings when the bytes do not match the new label. - Believing `encode` fixes invalid bytes in place: it returns a new String, and with identical encodings it copies without checking. - Relabelling a frozen string, or a string literal in a file whose magic comment freezes literals: `FrozenError`. - Transcoding twice: running `encode("UTF-8", "ISO-8859-1")` on text that is already correct UTF-8 turns `é` into `é`, the classic double-encoding mojibake.
- In Ruby, why does calling encode("UTF-8") on a UTF-8-labelled string with invalid bytes not repair it?Source and destination are the same encoding, so `encode` just copies the string without checking it. Pass `invalid: :replace` to make it replace the broken sequences, or call `scrub`. If the bytes are really another encoding, neither is right: use `encode("UTF-8", "ISO-8859-1")` so the characters survive.
- In Ruby, what does force_encoding do when the receiver is frozen, and what do you use instead?It raises `FrozenError`, because relabelling modifies the receiver. Use `str.dup.force_encoding(enc)` to relabel a copy, `str.encode(dst, src)` to reinterpret and convert without mutation, or `str.b` when you want a binary copy of the same bytes.
- In Ruby, how does double transcoding produce "é" instead of "é"?The text was already valid UTF-8, bytes `[195, 169]`, but code treated it as Latin-1 and ran `encode("UTF-8", "ISO-8859-1")`. Each byte became its own Latin-1 character, `Ã` and `©`, and was converted again. Convert only strings whose `valid_encoding?` is false or whose true source you know.
force_encoding is swapping the language sticker on an envelope: the letter inside is untouched, so a French letter labelled German is still French and now misread. encode is having the letter translated and rewritten, producing a new letter that says the same thing in the other language.
saying these in an interview costs you the question
- force_encoding converts the bytes into the new encoding
- encode only changes the label, so it can never raise
- force_encoding raises when the bytes are invalid for the new encoding
- force_encoding('UTF-8') repairs Latin-1 text that arrived labelled UTF-8
- encode mutates the receiver in place and returns it