skip to content

String Handling

Ruby strings are mutable byte sequences tagged with an encoding, built by interpolation or heredocs and searched with Regexp. Interviewers probe frozen literals, encodings and anchors.

on this pageshow

explore

questions

21

In Ruby, how do single-quoted and double-quoted string literals differ, and where do %q and %Q fit?

level: juniorimportance: must knowfreq 76%

answer

  1. one interpolates, one does not
  2. single quotes: only \\ and \' escape
  3. '\n' is backslash plus n
  4. %q is single, %Q and % are double
  5. #@var shorthand in double quotes

basics

~10 s

Double-quoted strings process escape sequences such as \n and \t and interpolate #{...}; single-quoted strings do neither, except for ' and \. %q(...) behaves like single quotes, %Q(...) and %(...) like double quotes.

solid answer

~40 s

A double-quoted literal is **processed**: escape sequences such as `\n`, `\t` and `\u00e9` become the characters they name, and `#{expr}` is evaluated and inserted. A single-quoted literal is taken almost **verbatim**: it recognises only `\'` and `\\`, so `'\n'` is two characters, a backslash and an `n`, and `'#{total}'` stays literal text. The percent forms choose the same behaviour with a delimiter you pick: `%q(...)` acts like single quotes, `%Q(...)` and bare `%(...)` like double quotes, which removes the need to escape embedded quotes: `%q(She said "it's hot")`. In double quotes, the shorthands `#@total`, `#@@count` and `#$mode` also interpolate instance, class and global variables without braces, a trap when a literal `#@` is meant.

code

ruby · 8 lines
ruby
total = 7.5
"Total:\t#{total}"  # => "Total:\t7.5" (a real tab)
'Total:\t#{total}'  # => "Total:\\t\#{total}"
'\n'.length          # => 2
"\n".length          # => 1
'it\'s'              # => "it's"
%q(say "hi")          # => "say \"hi\""
%(#{total} due)      # => "7.5 due"

go deeper

for a junior

Recall that double quotes interpolate and process escapes, single quotes do not, and that %q and %Q mirror them. Be able to predict '\n'.length.

for a middle

Explain the exact escape rules on both sides, including unknown escapes in double quotes and the #@ivar shorthand, and when a percent literal beats escaping.

for a senior

Spot the production versions of these bugs: a style cleanup that turns \n into literal text, or #@ in user-facing copy that silently swallows text.

for a principal

Decide what a team's linter should enforce for quote style and why the rule needs exceptions for text containing #{, backslashes or quotes.

## Two kinds of literal Ruby has two basic string literal forms, and they differ in how much the parser **processes** the text between the quotes. | Feature | `"double"` | `'single'` | |---|---|---| | `#{expr}` interpolation | yes | no, kept as text | | `\n`, `\t`, `\u00e9`, `\e` escapes | yes | no | | `\'` | not needed | yes, a single quote | | `\\` | a backslash | a backslash | | Unknown escape like `\d` | the character itself: `"d"` | kept: backslash and `d` | | `#@ivar`, `#@@cvar`, `#$global` shorthand | yes | no | - **Double quotes** are the everyday choice when text contains values or control characters. - **Single quotes** are the choice for text that must stay exactly as typed: regular-expression sources, Windows paths, templates that contain `#{` on purpose. Both produce ordinary `String` objects; once parsed, a literal without interpolation is the same kind of string whichever quotes wrote it, so pick by meaning. ## Escapes in detail In a double-quoted literal the backslash starts an **escape sequence**: - `"\n"` is one character, a newline; `"\t"` is a tab. - `"\u00e9"` is the character é; `"\u{1F600}"` takes a code point in braces. - `"\""` embeds a double quote. - `"\d"` is simply `"d"`: an unknown escape yields the character itself. In a single-quoted literal only two sequences exist: 1. `'\''` is a single quote. 2. `'\\'` is one backslash. Everything else keeps its backslash: `'\n'.length` is `2`, and `'C:\temp'` keeps both the backslash and the `t`. ## Percent literals When the text contains quotes, a **percent literal** lets you pick a delimiter instead of escaping: ```ruby %q(She said "it's too hot") # like single quotes %Q(Order ##{order_id} is ready) # like double quotes %(Order ##{order_id} is ready) # bare % is %Q ``` - `%q` behaves like single quotes: no interpolation, and only backslashes and the chosen delimiter can be escaped. - `%Q` and bare `%` behave like double quotes. - Delimiters can be `()`, `[]`, `{}`, `<>` or a pair of other punctuation such as `|` or `:`. Bracket pairs may nest: `%q(a (b) c)` is valid. ## The receipt example Printing a coffee-shop receipt header shows each form in its natural place: ```ruby shop = "Bean There" order = 1042 puts "#{shop}\tOrder ##{order}" # tab and interpolation puts 'Prices include VAT\n' # prints the backslash and n puts %q(Ask for our "barista's pick") ``` The second line prints `Prices include VAT\n` literally, a common surprise when someone switches quote styles during a style-guide cleanup. ## Adjacent literals Ruby joins **adjacent string literals** at parse time, which is handy for long text split across lines: ```ruby footer = "Thank you for visiting " \ 'Bean There' # => "Thank you for visiting Bean There" ``` - Any mix of single-quoted, double-quoted and percent literals can be joined, as long as a percent literal is not the last one; `"a" %q{c}` is parsed as a method call instead. - The trailing backslash continues the line; without it, the second literal would be a separate statement. - Each piece keeps its own rules: the double-quoted part interpolates, the single-quoted part does not. ## Interpolation shorthands In double-quoted strings `#` followed by a sigil interpolates without braces: - `"#@total"` inserts `@total`. - `"#@@count"` inserts the class variable `@@count`. - `"#$mode"` inserts the global variable `$mode`. This is why `"Email: support#@example"`-style text can silently lose part of itself: `#@example` reads an instance variable, usually `nil`, and inserts an empty string. Escape the hash (`"\#@"`) or use single quotes. ## Choosing in practice - Use **double quotes** when you interpolate or need an escape. - Use **single quotes** or `%q` for literal backslashes and `#{` sequences. - Use `%Q` or `%(...)` when interpolated text also contains double quotes. - Many teams let a linter enforce one default quote style; the behaviour differences above are what make a mechanical switch risky.

  • What does "support#@example" produce inside a method where @example is unset?
    `#@example` is interpolation shorthand for the instance variable `@example`. Unset, it reads as `nil`, which interpolates as an empty string, so the result is `"support"`. Escape the hash as `\#` or use single quotes.
  • Which characters can be escaped inside %q(...)?
    Only the backslash and the delimiters you chose. `%q` otherwise behaves like a single-quoted string: no interpolation and no `\n`-style escapes. Nested bracket pairs such as `%q(a (b) c)` need no escaping.
  • What does "\d" evaluate to in a double-quoted Ruby literal?
    `"d"`. In a double-quoted string an unrecognised escape yields the character itself, so the backslash disappears. That is why regular-expression text is usually written as a Regexp literal or in single quotes.

saying these in an interview costs you the question

  • Single-quoted strings interpolate #{} but skip escape sequences
  • '\n' is a newline in single quotes too
  • %q is the interpolating percent literal and %Q is the literal one
  • Single quotes are much faster at run time, so always prefer them
  • Only #{} interpolates; #@ivar in double quotes is plain text
open as a page

In Ruby, why does calling name.strip! inside a method change the caller's string, while name = name.strip does not?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Ruby hands the method a reference to the caller's String object. strip! edits that object in place, so the caller sees it; name = name.strip builds a new string and rebinds only the method's local parameter.

open as a page

In Ruby, how do the results of String#=~, String#match and String#match? differ, and which of them set $~ and $1?

level: juniorimportance: must knowfreq 62%

basics

~20 s

=~ returns the index where the match starts or nil, match returns a MatchData or nil, and match? returns true or false. =~ and match set $~, $1 and the other match variables; match? sets none of them.

open as a page

In Ruby, what is the difference between String#encode and String#force_encoding, and when is each the right call?

level: middleimportance: must knowfreq 55%

basics

~20 s

String#encode returns a new string whose bytes are converted so the same characters exist in the target encoding. String#force_encoding keeps the bytes and only relabels the receiver. Relabel when the label is wrong; encode when bytes must change.

open as a page

In Ruby, how do <<~, <<- and plain << heredocs differ, and what does quoting the identifier change?

level: middleimportance: must knowfreq 52%

basics

~20 s

Plain <<EOS needs its terminator at column 0 and keeps text as written; <<-EOS lets the terminator be indented; <<~EOS also removes the smallest common indentation from the body. Quoting the identifier as 'EOS' turns off interpolation and escapes.

open as a page

In Ruby, how do <<, + and += differ when building a string that another variable also references?

level: middleimportance: must knowfreq 66%

basics

~20 s

String#<< appends in place and returns the same object, so every variable holding it sees the change. + returns a new string, and s += x is shorthand for s = s + x, which rebinds only s.

open as a page

In Ruby, what does the # frozen_string_literal: true magic comment change, and which strings in that file stay mutable?

level: middleimportance: must knowfreq 56%

basics

~10 s

The comment makes every plain string literal in that one file frozen and deduplicated, so mutating one raises FrozenError. Interpolated literals, strings returned by methods, +"" buffers and strings from other files stay mutable.

open as a page

In Ruby, why does /^[A-Z]{2}-\d{3}$/ accept "AB-123\n<script>" as a licence plate, and how do \A and \z fix it?

level: middleimportance: must knowfreq 58%

basics

~20 s

Ruby's ^ and $ match at the start and end of every line, so a value whose first line is a valid plate passes. \A and \z match only the whole string's start and end, so /\A[A-Z]{2}-\d{3}\z/ rejects extra lines.

open as a page

In Ruby, what do String#length, String#bytesize and String#grapheme_clusters each count for a non-ASCII string?

level: juniorimportance: should knowfreq 50%

basics

~10 s

String#length (alias size) counts characters as the string's encoding defines them, String#bytesize counts raw bytes, and String#grapheme_clusters splits text into user-perceived characters. Outside ASCII the three numbers can all differ.

open as a page

In Ruby, how do String#chomp, chop and strip differ when cleaning a line read from input, and when would each surprise you?

level: juniorimportance: should knowfreq 54%

basics

~20 s

chomp removes one trailing line ending (\n, \r or \r\n) and nothing else; chop removes the last character whatever it is; strip removes leading and trailing ASCII whitespace and NUL. All three return new strings.

open as a page

In Ruby, how do String#sub and String#gsub behave when given a replacement string, a Hash, or a block?

level: juniorimportance: should knowfreq 55%

basics

~20 s

sub replaces the first match and gsub every match, each returning a new string. A replacement string may use back-references such as \1 or \k<name>, a Hash maps matched text to replacements, and a block computes each replacement.

open as a page

In Ruby, what is the ASCII-8BIT (BINARY) encoding, what does String#b return, and why can Encoding::CompatibilityError follow?

level: middleimportance: should knowfreq 28%

basics

~20 s

ASCII-8BIT, aliased BINARY, labels a string as raw bytes with no characters above 0x7F. String#b returns a copy of the same bytes labelled ASCII-8BIT. Joining a binary string holding high bytes with non-ASCII UTF-8 raises Encoding::CompatibilityError.

open as a page

In Ruby, when does String#encode raise Encoding::UndefinedConversionError instead of Encoding::InvalidByteSequenceError, and which options suppress each?

level: middleimportance: should knowfreq 30%

basics

~20 s

UndefinedConversionError means a valid source character has no counterpart in the target, such as € into ISO-8859-1. InvalidByteSequenceError means the source bytes are broken for their own encoding. undef: :replace and invalid: :replace substitute a replacement string instead.

open as a page

In Ruby, how do format, sprintf and String#% build a fixed-width receipt line, and what happens on mismatched arguments?

level: middleimportance: should knowfreq 46%

basics

~20 s

format and sprintf are the same Kernel method, and fmt % args calls it with the arguments; %-18s left-justifies in 18 characters and %7.2f prints two decimals. Too few arguments raise ArgumentError, a missing named key raises KeyError.

open as a page

In Ruby, how does String#split treat runs of whitespace, a single-space separator and trailing empty fields?

level: middleimportance: should knowfreq 40%

basics

~20 s

split with no argument or with " " splits on runs of whitespace and ignores leading whitespace. Any other separator splits on each occurrence, keeping empty fields between separators but dropping trailing ones unless a negative limit is passed.

open as a page

In Ruby, how do you extract the region code and serial number from a licence plate using named captures and MatchData?

level: middleimportance: should knowfreq 45%

basics

~10 s

Name each group with (?<region>…) and (?<number>…), call match, and read md[:region] and md[:number] from the MatchData, after checking it is not nil. named_captures returns all of them as a Hash.

open as a page

A Ruby job importing a legacy Latin-1 customer file raises ArgumentError: invalid byte sequence in UTF-8 on split; how do you diagnose and fix it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The Latin-1 bytes were labelled with the default external encoding, UTF-8, so the string is invalid and split or any Regexp match raises. Confirm with encoding, valid_encoding? and bytes, then transcode at the boundary with encode("UTF-8", "ISO-8859-1").

open as a page

In Ruby 4.0, what is a chilled string literal, and why does mutating one usually run with no visible warning?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A chilled string is a literal from a file with no frozen_string_literal comment. It reports frozen? as false and mutates normally; the first mutation emits a deprecation warning, which is invisible unless deprecation warnings are enabled.

open as a page

In Ruby, what do unary +"..." and -"..." (String#+@ and String#-@) return, and when would you use each?

level: seniorimportance: should knowfreq 30%

basics

~20 s

+str returns str itself when it is already mutable without warning, otherwise an unfrozen copy. -str returns a frozen, deduplicated string equal to str, so equal strings share one object. Use + for buffers, - for repeated values.

open as a page

In Ruby 4.0, how do Regexp.timeout, Regexp.new's timeout: keyword and Regexp.linear_time? help protect a service that matches untrusted input?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Regexp.timeout= sets a process-wide limit in seconds, Regexp.new(src, timeout:) sets one per pattern, and an overrun raises Regexp::TimeoutError. Regexp.linear_time? reports whether this interpreter can match a pattern in linear time, for checking your own patterns.

open as a page

In Ruby, why do ljust, center and format("%-18s") misalign receipt columns when item names contain accented or CJK characters?

level: seniorimportance: nice to knowfreq 16%

basics

~10 s

ljust, rjust, center and format widths count characters, not terminal columns. A decomposed accent adds an invisible character and a CJK character fills two columns, so padding by character count leaves rows misaligned.

open as a page