skip to content

In Ruby, how do String#sub and String#gsub behave when given a replacement string, a Hash, or a block?

level: juniorimportance: should knowfreq 55%

answer

  1. first match vs every match
  2. a String pattern is literal text
  3. '\1' and \k<name> in replacements
  4. Hash keys are the matched text
  5. the block sees $~ and $1

basics

~20 s

sub replaces the first match and gsub every match, each returning a new string. A replacement string may use back-references such as \1 or \k<name>, a Hash maps matched text to replacements, and a block computes each replacement.

solid answer

~40 s

`sub(pattern, replacement)` changes the first match and `gsub` every match; both return a new String. A String pattern is literal text, so `"AB.123".sub(".", "-")` replaces the dot only. In a **replacement string**, `\1`, `\k<name>` and `\0` refer to the match, but a double-quoted `"\1"` is a control character, so write `'\1'` or `"\\1"`; `$1` inside the replacement string is plain text. With a **Hash**, the matched text is looked up as a String key and the value is used; a missing key removes the match, and Symbol keys never match. With a **block**, Ruby yields the matched text, sets `$~` and `$1` for that match, and uses the block's value. `gsub` with neither a replacement nor a block returns an Enumerator of matches.

code

ruby · 8 lines
ruby
def normalise_plate(raw)
  compact = raw.upcase.gsub(/[\s-]+/, "")
  compact.sub(/\A([A-Z]{2})(\d{3})\z/, '\1-\2')
end

normalise_plate("ab 123")   # => "AB-123"
normalise_plate("AB-123")   # => "AB-123"
normalise_plate("a1")       # => "A1", unchanged: no match, so validate next

go deeper

for a junior

Recall sub versus gsub, that both return a new string, and the three replacement forms: string, Hash, block.

for a middle

Explain back-reference escaping in single versus double quotes, why interpolating $1 is stale, and the Hash rules for missing and Symbol keys.

for a senior

Choose the block form whenever logic is involved, keep normalisation steps small and tested, and validate after normalising rather than trusting the rewrite.

for a principal

Set where input normalisation happens so every entry point shares it, and when a canonical-format rule belongs in a value object instead of scattered gsub calls.

## First match or every match `String#sub` replaces the **first** match of a pattern and `String#gsub` replaces **every** match. Both return a new String and leave the receiver unchanged. Both accept either a Regexp or a String as the pattern: - a **Regexp** pattern is matched as a regular expression; - a **String** pattern is treated as literal text, so `"."` means a dot, not "any character". ```ruby "AB.123".sub(".", "-") # => "AB-123", literal dot "AB-123".gsub(/\d/, "#") # => "AB-###" ``` The replacement comes from one of three sources, and each has its own rules. ## A replacement string A String replacement may contain **back-references**: - `\1`, `\2`, … insert numbered captures, and `\0` or `\&` the whole match; - `\k<name>` inserts a named capture; - `` \` `` and `\'` insert the text before and after the match. The trap is escaping. Ruby's string literal consumes one level of backslashes first: | You write | The method receives | Result for `"AB-123".sub(/(\w+)-(\d+)/, …)` | |---|---|---| | `'\2-\1'` | `\2-\1` | `"123-AB"` | | `"\\2-\\1"` | `\2-\1` | `"123-AB"` | | `"\2-\1"` | control characters | `"\u0002-\u0001"` | | `"#{$2}-#{$1}"` | text built **before** the call | uses stale `$1`/`$2` | `$1` written inside the replacement string is plain text. Interpolating `#{$1}` evaluates it before `sub` runs, so it reads whatever the previous match left. When the replacement needs Ruby code, use the block form. ## A Hash With a Hash, each matched text is looked up as a **key**, and the value, converted with `to_s`, replaces it: - a missing key replaces the match with an empty string, deleting it; - Symbol keys never match, because the key is the matched **String**; - a Hash with a default value or default proc supplies replacements for unlisted matches. For licence plates typed by hand, a Hash fixes look-alike letters in the digit part: ```ruby FIX = { "O" => "0", "I" => "1" } "AB-1O3".sub(/(?<=-)\w{3}\z/) { |serial| serial.gsub(/[OI]/, FIX) } # => "AB-103" ``` ## A block With a block, Ruby calls the block for each match, passing the matched text, and uses the block's return value. Inside the block `$~`, `$1` and the named captures of **that** match are set, which the string form cannot offer: ```ruby "ab 123".gsub(/[a-z]+/) { |letters| letters.upcase } # => "AB 123" "AB-123".gsub(/(?<d>\d)/) { ($~[:d].to_i + 1).to_s } # => "AB-234" ``` ## Choosing between the three forms | Need | Form | Example | |---|---|---| | fixed text, maybe rearranging captures | replacement string | `sub(/(\w+)-(\d+)/, '\2-\1')` | | a fixed mapping from matched text to output | Hash | `gsub(/[OI]/, "O" => "0", "I" => "1")` | | logic, arithmetic, lookups or method calls | block | `gsub(/\d/) { \|d\| (d.to_i + 1).to_s }` | The string form is the most compact but the easiest to break with escaping; the block form is the most flexible and the easiest to read when anything more than copying captures is involved. ## Normalising user input A typical pipeline for plates typed as `"ab 123"`, `"AB123"` or `"ab-123"`: 1. `upcase` the input. 2. `gsub(/[\s-]+/, "")` to drop separators. 3. `sub(/\A([A-Z]{2})(\d{3})\z/, '\1-\2')` to insert the canonical hyphen. 4. Validate the result with an anchored `match?`. ## Common mistakes - Writing `"\\1"` in some places and `"\1"` in others; only the first is a back-reference in a double-quoted literal. - Passing a Hash with Symbol keys, which silently deletes every match. - Using a String pattern and expecting regexp syntax, or a Regexp built from user text without `Regexp.escape`. - Calling `sub` when every occurrence should change; the first match only is replaced, and tests with one occurrence will not notice. ## Other behaviours worth knowing - `gsub(pattern)` with neither a replacement nor a block returns an **Enumerator** over the matches, for example `"AB-123".gsub(/[A-Z]/).to_a` is `["A", "B"]`. - `sub` and `gsub` also exist as `sub!` and `gsub!`, which modify the receiver. - For removing a fixed substring, `delete_prefix`, `delete_suffix` or `tr` may be clearer than a regexp.

  • In Ruby, why does "AB-123".sub(/(\w+)-(\d+)/, "#{$2}-#{$1}") give a wrong result?
    The interpolation runs before `sub` is called, so `$1` and `$2` come from whatever match ran earlier in the method, often `nil`. Use back-references the method expands itself, `'\2-\1'`, or a block, `{ "#{$2}-#{$1}" }`, where the variables belong to the current match.
  • In Ruby, what does "AB-123".gsub(/[A-Z]/, A: "4") return, and why?
    `"-123"`. With a Hash replacement the matched text, a String, is the lookup key, so the Symbol key `:A` never matches. Missing keys give an empty replacement, so both letters are deleted. Use String keys: `{ "A" => "4" }`.

saying these in an interview costs you the question

  • A String pattern in gsub is compiled as a regular expression
  • "\1" in double quotes inserts the first capture
  • With a Hash replacement, matches without a key are left unchanged
  • $1 is not available inside a gsub block
  • sub replaces every match, gsub only the first