In Ruby, what is the ASCII-8BIT (BINARY) encoding, what does String#b return, and why can Encoding::CompatibilityError follow?
answer
- bytes, not text
- every byte valid, length equals bytesize
- String.new with no argument
- b returns a relabelled copy
- one side must be ASCII-only
basics
~20 sASCII-8BIT, aliased BINARY, labels a string as raw bytes with no characters above 0x7F. String#b returns a copy of the same bytes labelled ASCII-8BIT. Joining a binary string holding high bytes with non-ASCII UTF-8 raises Encoding::CompatibilityError.
solid answer
~40 s`Encoding::ASCII_8BIT`, aliased `Encoding::BINARY` and shown by Ruby 4.0 as `#<Encoding:BINARY (ASCII-8BIT)>`, is the label for data that is bytes rather than text: digests, packed integers, compressed or encrypted payloads. Every byte sequence is valid in it and `length` equals `bytesize`. `String.new` with no argument returns an empty ASCII-8BIT string. `s.b` returns a **new** string with the same bytes labelled ASCII-8BIT, so `"é".b` is `"\xC3\xA9"` with length 2; unlike `force_encoding`, it leaves the receiver alone and works on frozen strings. Ruby combines two strings of different ASCII-compatible encodings only when at least one side is ASCII-only, so appending UTF-8 `"é"` to a binary buffer holding `0xFF` raises `Encoding::CompatibilityError`. For byte buffers, `append_as_bytes` (Ruby 3.4+) concatenates without validation or conversion.
go deeper
Recall that ASCII-8BIT, also called BINARY, means raw bytes, and that String#b gives you a binary copy without changing the original.
Explain the concatenation rule behind Encoding::CompatibilityError and why ASCII-only test data hides it, and contrast b with force_encoding.
Keep binary buffers binary end to end, use b or append_as_bytes deliberately, and add non-ASCII fixtures so compatibility errors appear in tests, not production.
Draw a clear line in the codebase between byte-oriented and text-oriented layers, with explicit conversion at the seam, so encoding negotiation never happens by accident.
## Bytes with no text meaning Most Ruby strings carry a text encoding such as UTF-8. Some data is not text at all: the output of a digest, a packed binary header, a gzip payload, an image. For those, Ruby has **`Encoding::ASCII_8BIT`**, also reachable as **`Encoding::BINARY`**. In Ruby 4.0 its `inspect` reads `#<Encoding:BINARY (ASCII-8BIT)>`, and its `names` include both `"ASCII-8BIT"` and `"BINARY"`. What the label means in practice: - **Every byte sequence is valid**, so `valid_encoding?` is always `true`. - **One byte is one character**, so `length` equals `bytesize`, and indexing is by byte. - **Bytes 0x00–0x7F read as ASCII**, which is why the encoding is ASCII-compatible; bytes above `0x7F` have no character meaning. ## Where binary strings come from - `String.new` with **no argument** returns an empty ASCII-8BIT string. A string literal, by contrast, gets the script encoding, usually UTF-8. - `String#b` returns a binary copy of any string. - Many byte-oriented APIs, such as `Array#pack`, digests and binary reads, return ASCII-8BIT strings. ## String#b compared with force_encoding | | `s.b` | `s.force_encoding("BINARY")` | |---|---|---| | Result | new String | the receiver, `self` | | Receiver | unchanged | relabelled in place | | Frozen receiver | works | raises `FrozenError` | | Bytes | identical | identical | ```ruby s = "é" t = s.b # => "\xC3\xA9" t.encoding # => #<Encoding:BINARY (ASCII-8BIT)> t.length # => 2 s.encoding # => #<Encoding:UTF-8>, untouched ``` `b` is the tool when you need to compare, hash or slice by bytes without disturbing the original, for example comparing a received signature with a computed one byte for byte. ## Why CompatibilityError appears When you concatenate, interpolate or compare strings, Ruby must pick one encoding for the result. For two ASCII-compatible encodings the rules are: 1. **Same encoding**: fine. 2. **Different encodings, at least one side ASCII-only**: fine; the result takes the other side's encoding. 3. **Different encodings, both sides contain non-ASCII bytes**: Ruby cannot tell how to read the mixture and raises `Encoding::CompatibilityError`, with a message such as `incompatible character encodings: BINARY (ASCII-8BIT) and UTF-8`. ```ruby buf = "\xFF".b buf + "abc" # fine: "abc" is ASCII-only buf + "é" # Encoding::CompatibilityError ``` This is why a bug stays hidden in tests with ASCII fixtures and appears in production with the first accented name. `Encoding.compatible?(a, b)` returns the resulting encoding or `nil`, which helps when diagnosing. ## Building binary buffers Choose one representation per buffer: - keep it **binary**, and convert text into bytes explicitly before appending, for example `buf << name.b`, which is binary plus binary; - or, since **Ruby 3.4**, use **`append_as_bytes`**, which concatenates the given strings or integers as raw bytes with no encoding validation or conversion and returns `self`. Do not mix a binary buffer with UTF-8 fragments and hope that every fragment stays ASCII. ## Checking before you combine A few query methods make encoding bugs visible early: - `str.encoding` shows the label, and `str.ascii_only?` returns `true` when every byte is below `0x80`, which is exactly the condition under which mixing is always safe. - `Encoding.compatible?(a, b)` returns the encoding a combination would get, or `nil` when Ruby would raise, as it does for `"\xFF".b` and `"é"`. - `str.valid_encoding?` catches strings whose bytes do not fit their text label; for a binary string it is always `true`, so it cannot tell you whether the bytes are meaningful. In tests, include at least one fixture with non-ASCII text on every path that combines strings from different sources. The rule in step 2 above means ASCII-only fixtures pass every compatibility check, which is how this class of bug slips through review. ## Common mistakes - Treating `b` as a conversion: it never transcodes, so `"é".b` is two bytes, not Latin-1 `0xE9`. - Expecting `"é".encode("BINARY")` to behave like `b`: transcoding to ASCII-8BIT raises `Encoding::UndefinedConversionError` because `é` has no binary character. - Assuming `String.new` returns a UTF-8 string and then appending text to it.
- In Ruby, what does "é".encode("BINARY") do, compared with "é".b?`encode("BINARY")` tries to transcode the character `é` into ASCII-8BIT, which has no characters above `0x7F`, so it raises `Encoding::UndefinedConversionError`. `"é".b` does not transcode at all; it copies the two UTF-8 bytes and labels the copy ASCII-8BIT.
- In Ruby 3.4 and later, what does String#append_as_bytes do that << does not?`append_as_bytes` concatenates its arguments as raw bytes, with no encoding validation, negotiation or conversion, and returns `self`; an Integer argument appends its low-order byte. `<<` must first find a compatible encoding for the result, so mixing a high-byte binary buffer with non-ASCII UTF-8 raises `Encoding::CompatibilityError`.
saying these in an interview costs you the question
- String#b converts the text to Latin-1 bytes
- String.new with no argument returns a UTF-8 string
- A binary string can hold invalid byte sequences
- Any two Ruby strings can be concatenated whatever their encodings
- String#b relabels the receiver in place like force_encoding