skip to content

Encodings & Bytes

Every Ruby string carries an Encoding: encode transcodes while force_encoding only relabels bytes, and length counts characters, not bytes. Interviewers probe mojibake and invalid input.

on this pageshow

explore

questions

5

In Ruby, what is the difference between String#encode and String#force_encoding, and when is each the right call?

level: middleimportance: must knowfreq 55%

answer

  1. one changes bytes, one changes a label
  2. encode returns a new String
  3. force_encoding mutates, returns self
  4. relabel with the truth, then transcode
  5. encode(dst, src) does both at once

basics

~20 s

String#encode returns a new string whose bytes are converted so the same characters exist in the target encoding. String#force_encoding keeps the bytes and only relabels the receiver. Relabel when the label is wrong; encode when bytes must change.

solid answer

~30 s

`encode("ISO-8859-1")` **transcodes**: `"é"` goes from bytes `[195, 169]` in UTF-8 to `[233]` in Latin-1, and the receiver is untouched (`encode!` is the in-place form). By default it raises `Encoding::UndefinedConversionError` or `Encoding::InvalidByteSequenceError` when a character cannot be converted. `force_encoding` **relabels**: it never converts or validates, it rewrites the receiver's encoding tag and returns `self`, so on a frozen string it raises `FrozenError`. The classic pairing is bytes written as Latin-1 that arrived labelled UTF-8: `force_encoding("ISO-8859-1")` states the truth, then `encode("UTF-8")` converts, and `str.encode("UTF-8", "ISO-8859-1")` does both without mutating. Calling `force_encoding("UTF-8")` on those bytes changes nothing useful.

code

ruby · 9 lines
ruby
s = "é"
latin = s.encode("ISO-8859-1")
latin.bytes                  # => [233]
s.bytes                      # => [195, 169]  receiver untouched

label_only = s.dup.force_encoding("ISO-8859-1")
label_only.bytes             # => [195, 169]  same bytes
label_only.length            # => 2  read as two Latin-1 characters
label_only.equal?(label_only.force_encoding("ISO-8859-1"))  # => true, returns self

go deeper

for a junior

Remember the one-line rule: encode changes bytes and returns a new string, force_encoding only changes the label on the same bytes.

for a middle

Walk through the Latin-1 repair step by step, say that force_encoding returns self and mutates, and that same-encoding encode is a copy that validates nothing.

for a senior

Show where the repair belongs, at the input boundary, and how to avoid double transcoding by converting only strings you know are mislabelled.

for a principal

Argue for one internal encoding across the codebase and explicit conversion at every external edge, so encoding bugs surface at the edge instead of deep in business logic.

## Bytes and the label on them A Ruby `String` is a byte sequence plus an `Encoding`. Two different repairs exist because two different things can be wrong: - the **label** is wrong: the bytes are fine Latin-1 text, but Ruby believes they are UTF-8; - the **bytes** are wrong for where they are going: the text is correct UTF-8, but a legacy system needs Latin-1 bytes. `force_encoding` fixes the first problem and `encode` fixes the second. ## String#encode transcodes `encode(dst)` reads the receiver's characters using its current encoding and writes the same characters as bytes of `dst`, returning a **new** String. The receiver keeps its bytes and label; `encode!` is the mutating variant. - `"é".encode("ISO-8859-1").bytes` returns `[233]`, while the original `"é".bytes` stays `[195, 169]`. - `encode(dst, src)` interprets the receiver as `src` regardless of its label, then transcodes to `dst`. It is the non-mutating way to say "these bytes are really Latin-1". - A character missing from the target raises `Encoding::UndefinedConversionError`; broken source bytes raise `Encoding::InvalidByteSequenceError`. The `invalid:`, `undef:`, `replace:` and `fallback:` options change that. - When source and destination encodings are the same, `encode` is a plain copy: it does not validate or repair anything unless you pass `invalid: :replace`. ## String#force_encoding relabels `force_encoding(enc)` changes only the encoding tag of the **receiver** and returns `self`. The bytes are identical before and after. - It never raises for "bad" bytes. Relabelling `"\xE9"` as UTF-8 succeeds and simply yields a string whose `valid_encoding?` is `false`. - Because it modifies the receiver, it raises `FrozenError` on a frozen string. Call it on a copy (`dup`) or, for binary work, use `String#b`, which returns a relabelled copy. ## Side by side | | `encode(dst)` | `force_encoding(enc)` | |---|---|---| | Bytes | converted | unchanged | | Receiver | untouched, new String returned | modified, `self` returned | | Validates input | yes, raises by default | never | | Typical use | produce bytes a consumer needs | correct a wrong label | | In-place variant | `encode!` | is already in place | ## The relabel-then-transcode idiom Data read from a legacy Latin-1 source is usually labelled with `Encoding.default_external`, commonly UTF-8. The repair takes two steps: 1. **Tell the truth about the bytes**: `force_encoding("ISO-8859-1")`, or pass `"ISO-8859-1"` as the second argument to `encode`. 2. **Convert to what the application uses**: `encode("UTF-8")`. 3. **Check** that `valid_encoding?` is `true` on the result before letting it travel further. ```ruby raw = "caf\xE9".dup # Latin-1 bytes labelled UTF-8 raw.valid_encoding? # => false raw.force_encoding("ISO-8859-1") raw.encode("UTF-8") # => "café" "caf\xE9".encode("UTF-8", "ISO-8859-1") # => "café", receiver untouched ``` Doing only step 2 on the mislabelled string does nothing useful: `raw.encode("UTF-8")` while `raw` is still labelled UTF-8 is a same-encoding copy. Doing only step 1 with the wrong target, `force_encoding("UTF-8")`, just re-asserts the lie. ## Choosing in practice - **An HTTP body or file whose bytes are right but whose label is wrong**: relabel with `force_encoding` on a copy, or use `encode(dst, src)` to convert in one step. - **A consumer that needs different bytes**, such as a legacy system expecting ISO-8859-1 or a Windows tool expecting UTF-16LE: `encode(dst)`, with an explicit policy for characters the target lacks. - **Bytes that are not text at all**, such as a digest or a compressed payload: neither; take a binary copy with `b` and never transcode. - **Unknown input you merely want to make safe to display**: `scrub`, accepting that damaged sequences become U+FFFD. A useful interview habit is to say, before choosing, whether the bytes or the label is wrong. That one sentence decides the method. ## Traps - Believing `force_encoding` converts: it produces mojibake or invalid strings when the bytes do not match the new label. - Believing `encode` fixes invalid bytes in place: it returns a new String, and with identical encodings it copies without checking. - Relabelling a frozen string, or a string literal in a file whose magic comment freezes literals: `FrozenError`. - Transcoding twice: running `encode("UTF-8", "ISO-8859-1")` on text that is already correct UTF-8 turns `é` into `é`, the classic double-encoding mojibake.

  • In Ruby, why does calling encode("UTF-8") on a UTF-8-labelled string with invalid bytes not repair it?
    Source and destination are the same encoding, so `encode` just copies the string without checking it. Pass `invalid: :replace` to make it replace the broken sequences, or call `scrub`. If the bytes are really another encoding, neither is right: use `encode("UTF-8", "ISO-8859-1")` so the characters survive.
  • In Ruby, what does force_encoding do when the receiver is frozen, and what do you use instead?
    It raises `FrozenError`, because relabelling modifies the receiver. Use `str.dup.force_encoding(enc)` to relabel a copy, `str.encode(dst, src)` to reinterpret and convert without mutation, or `str.b` when you want a binary copy of the same bytes.
  • In Ruby, how does double transcoding produce "é" instead of "é"?
    The text was already valid UTF-8, bytes `[195, 169]`, but code treated it as Latin-1 and ran `encode("UTF-8", "ISO-8859-1")`. Each byte became its own Latin-1 character, `Ã` and `©`, and was converted again. Convert only strings whose `valid_encoding?` is false or whose true source you know.

force_encoding is swapping the language sticker on an envelope: the letter inside is untouched, so a French letter labelled German is still French and now misread. encode is having the letter translated and rewritten, producing a new letter that says the same thing in the other language.

saying these in an interview costs you the question

  • force_encoding converts the bytes into the new encoding
  • encode only changes the label, so it can never raise
  • force_encoding raises when the bytes are invalid for the new encoding
  • force_encoding('UTF-8') repairs Latin-1 text that arrived labelled UTF-8
  • encode mutates the receiver in place and returns it
open as a page

In Ruby, what do String#length, String#bytesize and String#grapheme_clusters each count for a non-ASCII string?

level: juniorimportance: should knowfreq 50%

basics

~10 s

String#length (alias size) counts characters as the string's encoding defines them, String#bytesize counts raw bytes, and String#grapheme_clusters splits text into user-perceived characters. Outside ASCII the three numbers can all differ.

open as a page

In Ruby, what is the ASCII-8BIT (BINARY) encoding, what does String#b return, and why can Encoding::CompatibilityError follow?

level: middleimportance: should knowfreq 28%

basics

~20 s

ASCII-8BIT, aliased BINARY, labels a string as raw bytes with no characters above 0x7F. String#b returns a copy of the same bytes labelled ASCII-8BIT. Joining a binary string holding high bytes with non-ASCII UTF-8 raises Encoding::CompatibilityError.

open as a page

In Ruby, when does String#encode raise Encoding::UndefinedConversionError instead of Encoding::InvalidByteSequenceError, and which options suppress each?

level: middleimportance: should knowfreq 30%

basics

~20 s

UndefinedConversionError means a valid source character has no counterpart in the target, such as € into ISO-8859-1. InvalidByteSequenceError means the source bytes are broken for their own encoding. undef: :replace and invalid: :replace substitute a replacement string instead.

open as a page

A Ruby job importing a legacy Latin-1 customer file raises ArgumentError: invalid byte sequence in UTF-8 on split; how do you diagnose and fix it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The Latin-1 bytes were labelled with the default external encoding, UTF-8, so the string is invalid and split or any Regexp match raises. Confirm with encoding, valid_encoding? and bytes, then transcode at the boundary with encode("UTF-8", "ISO-8859-1").

open as a page