skip to content

In Ruby, why does /^[A-Z]{2}-\d{3}$/ accept "AB-123\n<script>" as a licence plate, and how do \A and \z fix it?

level: middleimportance: must knowfreq 58%

answer

  1. line anchors vs string anchors
  2. $ matches before a newline
  3. the m flag is dot-all in Ruby
  4. \Z tolerates one final newline
  5. \A…\z for whole-value validation

basics

~20 s

Ruby's ^ and $ match at the start and end of every line, so a value whose first line is a valid plate passes. \A and \z match only the whole string's start and end, so /\A[A-Z]{2}-\d{3}\z/ rejects extra lines.

solid answer

~40 s

Ruby's `^` and `$` are **always line anchors**: `^` matches at the start of the string or after any newline, `$` at the end or before any newline. `/^[A-Z]{2}-\d{3}$/` therefore finds `AB-123` on the first line of `"AB-123\n<script>"`, and `match?` returns `true` for the whole value. `\A` matches only at the start of the string and `\z` only at its very end, so `/\A[A-Z]{2}-\d{3}\z/` accepts exactly one plate and nothing else. `\Z` also exists; it allows one trailing newline, which is rarely what a validator wants. The `m` flag does not change any of this: in Ruby it only lets `.` match a newline. Use `\A…\z` for every whole-value validation, and keep `^`/`$` for searching line-oriented text.

code

ruby · 6 lines
ruby
log = "AB-123 entered\nCD-456 left\n"
log.scan(/^[A-Z]{2}-\d{3}/)          # => ["AB-123", "CD-456"], ^ per line

"a\nc".match?(/a.c/)                 # => false, . stops at a newline
"a\nc".match?(/a.c/m)                # => true, m makes . match it
/\A[A-Z]{2}-\d{3}\Z/.match?("AB-123\n")  # => true, \Z allows one final newline

go deeper

for a junior

Recall that \A and \z anchor the whole string while ^ and $ anchor lines, and that validation should use \A and \z.

for a middle

Walk through why an embedded newline slips past ^…$, the difference between \z and \Z, and what Ruby's m flag really does.

for a senior

Treat ^…$ in a validator as a security finding, add newline test cases to every validator, and audit patterns copied from other languages.

for a principal

Make whole-value anchoring a reviewed convention, and decide where validation lives so every entry point, not just the web form, runs the strict pattern.

## Two kinds of anchor An **anchor** matches a position, not a character. Ruby has two families: | Anchor | Matches at | `"AB-123\n<script>"` | |---|---|---| | `^` | start of the string, or right after any `\n` | matches at 0 and after the newline | | `$` | end of the string, or right before any `\n` | matches before the newline and at the end | | `\A` | start of the string only | matches at 0 only | | `\z` | end of the string only | matches only after `>` | | `\Z` | end of the string, or before a final single `\n` | matches only at the end here | The key fact is that in Ruby `^` and `$` are **line** anchors with no switch that turns them into string anchors. ## Why the validator passes a hostile value A validation such as `PLATE.match?(params[:plate])` asks "does the pattern match somewhere in this string?". With `/^[A-Z]{2}-\d{3}$/`: 1. `^` matches at position 0. 2. `[A-Z]{2}-\d{3}` consumes `AB-123`. 3. `$` matches right before `\n`, because it is the end of a line. 4. The match succeeds, so the whole value, including `<script>` on the second line, is accepted and stored. The same happens with `"<script>\nAB-123"`, where `^` matches after the newline. Any code that later renders, logs or passes the value to a shell trusts a check that never covered it. ```ruby LOOSE = /^[A-Z]{2}-\d{3}$/ STRICT = /\A[A-Z]{2}-\d{3}\z/ LOOSE.match?("AB-123\n<script>") # => true LOOSE.match?("<script>\nAB-123") # => true STRICT.match?("AB-123\n<script>") # => false STRICT.match?("AB-123") # => true ``` ## Why an extra line is dangerous The accepted value is not only ugly; the lines after the plate travel wherever the plate goes: - into **HTML**, where an unescaped second line can carry markup; - into **log files**, where an embedded newline forges a fake log entry that looks like it came from the application; - into **headers, file names or command arguments**, where a newline can split one value into two. Output escaping still matters in each of those places, but a validator that says "valid plate" should mean exactly one plate. ## `\z` versus `\Z` - **`\z`** is the true end of the string. `"AB-123\n"` fails `/\A[A-Z]{2}-\d{3}\z/`. - **`\Z`** also matches just before one final newline, so `"AB-123\n"` passes `/\A[A-Z]{2}-\d{3}\Z/`, while `"AB-123\n\n"` does not. `\Z` suits a line read with `gets`, which keeps its `"\n"`, but calling `chomp` first and using `\z` is clearer. ## The `m` flag in Ruby In many other regex flavours, `^` and `$` anchor the whole string by default and a "multiline" flag turns them into line anchors, while a separate "dot-all" flag lets `.` cross newlines. Ruby works differently: - `^` and `$` are line anchors all the time; - Ruby's **`m`** flag is the dot-all switch: `/a.c/m` matches `"a\nc"`, `/a.c/` does not; - nothing in the `m` flag affects `^`, `$`, `\A` or `\z`. A pattern copied from another language with `^…$` and no flags is therefore looser in Ruby than its author expected. ## Test inputs every validator should see A validator is only as good as the inputs it has been tried against. For a whole-value plate check, a table-driven test with these cases catches the anchor bug and its relatives: | Input | Expected | What it catches | |---|---|---| | `"AB-123"` | valid | the happy path | | `"AB-123\n"` | invalid | `\Z` or `$` used as the end anchor | | `"AB-123\n<script>"` | invalid | `$` used as the end anchor | | `"<script>\nAB-123"` | invalid | `^` used as the start anchor | | `" AB-123"` | invalid, or valid after `strip` | missing start anchor or missing normalisation | | `"ab-123"` | invalid, or valid after `upcase` | an accidental `i` flag, or missing normalisation | Running each case through `match?` takes a few lines and documents the intended format better than a comment. ## Practical rules - Validate whole values with `\A` and `\z`, every time. - Use `^` and `$` when you really mean lines, for example scanning a multi-line log for plates at the start of each line. - Normalise input before matching (`strip`, `upcase`) rather than widening the pattern. - Add a test with an embedded newline for every validator; it is the cheapest regression guard for this bug.

  • In Ruby, what does the m flag change in /\A.+\z/m?
    Only the dot: with `m`, `.` also matches `\n`, so `.+` can span lines and the pattern accepts any non-empty string. `\A` and `\z` mean the same with or without `m`, and so do `^` and `$`, which are line anchors either way.
  • In Ruby, when is \Z the right end anchor instead of \z?
    When the string legitimately ends with one newline you want to ignore, such as a line from `gets`. `\Z` matches at the end or just before a single final `\n`. For validating a submitted value, `\z` is safer; strip or chomp the input first if trailing whitespace is acceptable.

saying these in an interview costs you the question

  • In Ruby ^ and $ match only at the start and end of the whole string
  • Adding the m flag makes ^ and $ match only at string boundaries
  • \Z and \z are interchangeable end anchors
  • Web form fields never contain newlines, so ^ and $ are safe for validation
  • Anchoring the start with \A is enough; the end anchor is optional