skip to content

In Ruby, how do String#chomp, chop and strip differ when cleaning a line read from input, and when would each surprise you?

level: juniorimportance: should knowfreq 54%

answer

  1. one line ending vs last character
  2. chomp handles \n, \r and \r\n
  3. chop always eats a character
  4. strip: both ends, ASCII whitespace
  5. 4.0: strip(*selectors)

basics

~20 s

chomp removes one trailing line ending (\n, \r or \r\n) and nothing else; chop removes the last character whatever it is; strip removes leading and trailing ASCII whitespace and NUL. All three return new strings.

solid answer

~40 s

`gets.chomp` is the idiom because `chomp` removes **one trailing line separator**, `"\n"`, `"\r"` or `"\r\n"`, and leaves the string alone if there is none. `chop` removes the **last character** unconditionally (treating `"\r\n"` as one), so `chop` on a final line without a newline eats real data. `strip` removes whitespace at **both ends**: space, tab, newline, vertical tab, form feed, carriage return and NUL, but not a Unicode no-break space pasted from a web page. `chomp` also takes an argument: `chomp("")` removes all trailing newlines, and `chomp("cd")` removes that suffix once. Since Ruby 4.0, `strip`, `lstrip` and `rstrip` accept character selectors such as `strip("-+")`. Each has a bang form that edits the string in place.

code

ruby · 7 lines
ruby
"Latte\r\n".chomp    # => "Latte"
"Latte".chomp        # => "Latte"
"Latte".chop         # => "Latt"
"Latte\n\n".chomp("") # => "Latte"
"  oat milk \n".strip # => "oat milk"
"\u00A0Mocha".strip   # => "\u00A0Mocha" (no-break space kept)
"--Mocha--".strip("-") # => "Mocha" (Ruby 4.0)

go deeper

for a junior

Recall gets.chomp and what chomp, chop and strip each remove. Predict "Latte".chop.

for a middle

Explain the exact separator rules of chomp, including chomp("") and a suffix argument, and the fixed whitespace set strip uses.

for a senior

Diagnose production input bugs: data lost to chop on the last line, and no-break spaces surviving strip in imported spreadsheets.

for a principal

Define where input normalisation lives in an import pipeline and how aggressive it may be before it corrupts legitimate data.

## Three cleaners, three targets Reading input line by line leaves a line ending on every string. Ruby offers three methods that look similar and remove different things: | Method | Removes | From | If nothing matches | |---|---|---|---| | `chomp` | one line separator: `"\n"`, `"\r"` or `"\r\n"` | the end | returns an unchanged copy | | `chop` | the last character (`"\r\n"` counts as one) | the end | `""` stays `""` | | `strip` | all leading and trailing whitespace | both ends | returns an unchanged copy | All three return a **new** string. `chomp!`, `chop!` and `strip!` edit the receiver instead. ## chomp: the input idiom `gets.chomp` removes exactly the newline that `gets` kept: - `"Latte\n".chomp` is `"Latte"`. - `"Latte\r\n".chomp` is `"Latte"`, so Windows line endings are handled. - `"Latte".chomp` is `"Latte"`: nothing to remove, nothing removed. - `"Latte\n\n".chomp` is `"Latte\n"`: only one separator goes. `chomp` also takes an argument: 1. `chomp("")` removes **all** trailing `"\n"` and `"\r\n"` sequences, but not a lone `"\r"`. 2. `chomp("cd")` removes that exact suffix once: `"abcdcd".chomp("cd")` is `"abcd"`. ## chop: the character eater `chop` predates `chomp` as an input idiom and is almost never what you want for it: - `"Latte\n".chop` is `"Latte"`, which looks right. - `"Latte".chop` is `"Latt"`: the last line of a file often has no trailing newline, and `chop` removes a real letter. Use `chop` only when you truly mean "drop the final character", such as trimming a trailing comma you appended yourself. ## strip: both ends, defined whitespace `strip` removes leading and trailing **whitespace**, which `String` defines as NUL, tab, line feed, vertical tab, form feed, carriage return and space. It is the right tool for user-typed values such as `" oat milk \n"`. It does **not** remove: - a Unicode no-break space (`"\u00A0"`), common in text copied from web pages or spreadsheets; - zero-width characters; - whitespace in the **middle** of the string. For those cases, use a character selector (below) or a regular-expression replacement. ## Ruby 4.0: selectors for strip Since Ruby 4.0, `strip`, `lstrip` and `rstrip` (and their bang forms) accept `*selectors`, the same character-selector syntax as `delete` and `squeeze`: ```ruby "---Latte+++".strip("-+") # => "Latte" "01234abc56789".strip("0-9") # => "abc" "***TOTAL***".lstrip("*") # => "TOTAL***" ``` With selectors, only the selected characters are removed, not whitespace. On Ruby 3.4 and earlier these calls raise `ArgumentError`, so code that must run on older versions still needs a regular expression. ## Line endings from different systems Files arrive with three conventions, and `chomp` with no argument handles each of them: 1. `"\n"` (Unix and macOS): removed. 2. `"\r\n"` (Windows): removed as a pair. 3. `"\r"` alone (very old Mac files): removed. A reversed pair, `"\n\r"`, is not a separator, so `"abc\n\r".chomp` removes only the `"\r"` and returns `"abc\n"`. `strip` removes all of these characters anyway, which is one reason it is tempting, and one reason it can remove too much from records whose leading or trailing spaces matter. ## Choosing by input - **Lines from a file or socket:** `chomp`, and nothing more if spaces inside the record are significant, as in fixed-width exports. - **Values typed into a form:** `strip`, because users add stray spaces at both ends. - **Known wrapper characters:** on Ruby 4.0, `strip("*")` or `rstrip(",")`; on older Rubies, `delete_prefix` and `delete_suffix` for one exact prefix or suffix, or a regular expression. - **A trailing character you appended yourself:** `chop` is acceptable, though `delete_suffix(",")` states the intent more clearly. ## Checking prefixes while cleaning Input cleaning often pairs with a prefix test. `String#start_with?` takes several candidates and returns `true` if any matches: ```ruby while (line = gets) line = line.chomp next if line.start_with?("#", "//") # skip comment lines record_order(line) end ``` A string argument is matched literally, not as a pattern, so `start_with?(".")` checks for a real dot.

  • Why is gets.chomp preferred over gets.chop for reading lines?
    `chomp` removes a line separator only if one is there, while `chop` always removes the last character. The final line of a file often lacks a newline, and `chop` would cut a real character from it.
  • What does "a\n\n\n".chomp("") return, and how does that differ from chomp with no argument?
    It returns `"a"`: an empty-string argument removes every trailing `"\n"` or `"\r\n"`. With no argument, `chomp` removes only one separator and would return `"a\n\n"`.
  • In Ruby, does start_with?("#", "//") need both prefixes to match?
    No. `String#start_with?` returns `true` if any of its arguments matches the beginning of the string. String arguments are compared literally; a Regexp argument is matched as a pattern anchored at the start.

saying these in an interview costs you the question

  • chop removes the trailing newline only when there is one
  • chomp removes every kind of trailing whitespace
  • strip also removes whitespace inside the string
  • strip removes Unicode no-break spaces along with ASCII spaces
  • strip("-") works in every Ruby 3.x release