In Ruby, what do String#length, String#bytesize and String#grapheme_clusters each count for a non-ASCII string?
answer
- characters, bytes, what a reader sees
- length and size are aliases
- bytesize for byte-limited storage
- bytes gives Integers, chars gives Strings
- combining accent: two chars, one cluster
basics
~10 sString#length (alias size) counts characters as the string's encoding defines them, String#bytesize counts raw bytes, and String#grapheme_clusters splits text into user-perceived characters. Outside ASCII the three numbers can all differ.
solid answer
~40 s`length` and `size` count characters in the string's `Encoding`, so `"café".length` is 4 while `"café".bytesize` is 5, because UTF-8 stores `é` in two bytes. `bytesize` is the number to check against a byte-limited column, a protocol length field or a buffer. For valid text, `chars` returns one-character Strings and `bytes` returns Integers, so their sizes match `length` and `bytesize`. `grapheme_clusters` keeps a base letter together with its combining marks or emoji modifiers: `"cafe\u0301"` has `length` 5 and `bytesize` 6 but 4 clusters, and a thumbs-up with a skin-tone modifier has `length` 2 but one cluster. Use clusters when a limit is about what the user sees, such as a display-name length.
code
ruby · 5 linesname = "Zoë \u{1F44D}\u{1F3FD}"
name.length # => 6
name.bytesize # => 13
name.grapheme_clusters.size # => 5
name.bytes.first(4) # => [90, 111, 195, 171]go deeper
Recall the three numbers for one example string: length counts characters, bytesize counts bytes, grapheme_clusters counts what a reader sees. Say which one a byte-limited column needs.
Explain why the numbers differ: UTF-8 uses one to four bytes per character, and combining marks or emoji modifiers are separate characters. Show byteslice producing an invalid string.
Tie each count to a real limit: bytes for storage and protocols, characters for Ruby slicing, clusters for anything a user reads. Mention normalization before comparing user input.
Frame length rules as a product decision: which limit the database, the API contract and the UI each enforce, and how to keep those three from disagreeing for non-Latin names.
## A Ruby String is bytes plus an Encoding Every `String` in Ruby holds a sequence of bytes and an `Encoding` object that says how to read those bytes as characters. A literal in a source file gets the script encoding, which is **UTF-8** by default, so `"café".encoding` returns `#<Encoding:UTF-8>`. Because the encoding travels with each string, the question "how long is this string?" has more than one honest answer, and Ruby gives each answer its own method. ## Three counts, three methods | Method | Unit counted | `"café"` | `"cafe\u0301"` | thumbs-up + skin tone | |---|---|---|---|---| | `length` / `size` | characters (code points in UTF-8) | 4 | 5 | 2 | | `bytesize` | bytes | 5 | 6 | 8 | | `grapheme_clusters.size` | user-perceived characters | 4 | 4 | 1 | - **`length`** (alias **`size`**) walks the bytes using the string's encoding and counts characters. In UTF-8 a character takes one to four bytes. - **`bytesize`** returns the raw byte count without decoding anything. - **`grapheme_clusters`** returns an Array of Strings, each one a unit a reader sees as a single symbol: a base letter plus combining accents, or an emoji plus its modifiers. `each_grapheme_cluster` yields them one at a time. The companion methods return arrays whose sizes line up with the counts for valid text: - `bytes` returns an Array of **Integers**, one per byte. - `chars` returns an Array of one-character **Strings**, one per character. - `codepoints` returns an Array of Integers, one per character. ## Seeing it in irb ```ruby "café".length # => 4 "café".bytesize # => 5 "café".bytes # => [99, 97, 102, 195, 169] "cafe\u0301".chars.size # => 5, the accent is its own character "cafe\u0301".grapheme_clusters # => ["c", "a", "f", "é"] "\u{1F44D}\u{1F3FD}".length # => 2 "\u{1F44D}\u{1F3FD}".grapheme_clusters.size # => 1 ``` `"cafe\u0301"` spells the accent as a separate combining character (U+0301), so `chars` splits it from the `e`, while `grapheme_clusters` keeps them together. `"café"` typed directly usually uses the precomposed `é` (U+00E9), which is one character. The two strings look identical on screen but differ in `length`, `bytesize` and `==`. ## Which count to enforce where 1. **Storage and wire limits are bytes.** A database column declared in bytes, a fixed-size header, or a payload cap should be checked with `bytesize`. 2. **Validation of "at most N characters"** in the Ruby sense uses `length`, which is what `str[0, n]` also counts. 3. **Limits a human reads**, such as a 20-symbol display name or a tweet-style counter, use `grapheme_clusters.size`, so that an accented letter or a flag emoji counts once. ## Slicing by characters or by bytes `str[0, n]` slices by characters, so it never cuts a UTF-8 sequence in half. `byteslice(0, n)` slices by bytes, which is fast and exact for binary data but can split a multi-byte character: `"héllo".byteslice(0, 2)` returns `"h\xC3"`, a string still labelled UTF-8 whose `valid_encoding?` is `false`. Later operations such as `split` or a Regexp match on it raise `ArgumentError: invalid byte sequence in UTF-8`. If you must cut on a byte budget, check `valid_encoding?` on the result or trim the broken tail with `scrub("")`. ## Where a string's encoding comes from The counts above depend on the label, so it helps to know how a string gets one: - A **string literal** takes the script encoding of its source file. That default is **UTF-8**, and `__ENCODING__` returns it inside the file. - Data **read from a file, socket or pipe** is labelled with `Encoding.default_external` unless the reader says otherwise, which is usually UTF-8 on a modern system. - `String.new` with no argument is labelled **ASCII-8BIT**, where every byte counts as one character, so `length` equals `bytesize`. The same bytes therefore report different lengths under different labels. `"é".b.length` is 2, because `String#b` relabels a copy as binary and each byte becomes its own character. ## Iterating without building arrays `bytes`, `chars` and `grapheme_clusters` allocate a whole Array. For long strings, the iterator forms `each_byte`, `each_char` and `each_grapheme_cluster` yield one element at a time and return an Enumerator when called without a block, which keeps memory flat when you only need to count or stop early. ## Common mistakes - Treating `length` as a byte count, as C's `strlen` would be. - Assuming `chars` groups an accent with its letter; that is what `grapheme_clusters` does. - Enforcing a byte-limited column with `length`, which passes 255 multi-byte characters that need far more than 255 bytes. - Comparing a precomposed and a decomposed spelling with `==` and expecting `true`; `unicode_normalize` brings both to one form first.
- In Ruby, what does "héllo".byteslice(0, 2) return, and why is it risky?It returns `"h\xC3"`: the first two bytes, which cut the two-byte `é` in half. The result is still labelled UTF-8 but `valid_encoding?` is `false`, so a later `split` or Regexp match raises `ArgumentError`. Slice by characters with `str[0, n]`, or check `valid_encoding?` and trim the broken tail with `scrub("")` when a byte budget is unavoidable.
- In Ruby, how do you truncate a display name to 20 visible symbols without splitting an emoji or accent?Work on clusters: `name.grapheme_clusters.first(20).join`. `name[0, 20]` counts characters, so it can keep a base emoji and drop its skin-tone modifier, or keep an `e` and drop its combining accent. For a byte-limited store, check `bytesize` of the joined result as a separate rule.
saying these in an interview costs you the question
- String#length returns the number of bytes, like C's strlen
- bytesize and length are always equal for Ruby strings
- String#chars keeps an accented letter and its combining mark together
- Ruby strings carry no encoding; they are plain byte arrays
- Checking length is enough to fit UTF-8 text into a byte-limited column