skip to content

In Ruby 4.0, how do Regexp.timeout, Regexp.new's timeout: keyword and Regexp.linear_time? help protect a service that matches untrusted input?

level: seniorimportance: should knowfreq 30%

answer

  1. a clock on every match
  2. process-wide default, nil initially
  3. per-regexp timeout: wins
  4. Regexp::TimeoutError < RegexpError
  5. backreferences defeat linear_time?

basics

~20 s

Regexp.timeout= sets a process-wide limit in seconds, Regexp.new(src, timeout:) sets one per pattern, and an overrun raises Regexp::TimeoutError. Regexp.linear_time? reports whether this interpreter can match a pattern in linear time, for checking your own patterns.

solid answer

~40 s

Since Ruby 3.2, `Regexp.timeout = 1.0` gives every match a default limit in seconds; it starts as `nil`, meaning no limit, and is process-global. `Regexp.new(source, timeout: 0.2)` sets a per-pattern limit that overrides the default, readable with `Regexp#timeout`. A match that runs too long raises `Regexp::TimeoutError`, a `RegexpError` and so a `StandardError`, which the caller should rescue and turn into a rejected input. Ruby 3.2 also added a memoization cache that makes many patterns match in linear time, and `Regexp.linear_time?(re)` tells you whether this interpreter can apply it; backreferences, subexpression calls and conditionals make it return `false`. Use `linear_time?` as a test assertion on your own patterns, a global timeout as a backstop, a tight per-regexp timeout for patterns users supply, and `Regexp.escape` for user text inserted into a pattern.

go deeper

for a junior

Recall that Regexp.timeout= sets a limit in seconds, that it is nil by default, and that an overrun raises Regexp::TimeoutError.

for a middle

Explain how the per-regexp timeout: combines with the global default, and what Regexp.linear_time? reports and which constructs make it false.

for a senior

Lay out the layered defence: global backstop, linearity assertions in tests, tight timeouts for user rules, escaping, input caps and a clean rescue path.

for a principal

Decide whether users may supply patterns at all, what budget a single match may spend, and how the policy is re-verified on each Ruby upgrade.

## The problem these APIs address Ruby's regexp engine backtracks: for some patterns and inputs, matching time grows far faster than the input length, and one crafted request can hold a CPU for seconds or minutes. That risk is highest when the **input** comes from users, and worse when the **pattern** does too. Ruby 3.2 added three tools for it, and Ruby 4.0 keeps all of them. ## `Regexp.timeout`: a process-wide default - `Regexp.timeout = 1.0` sets the default limit, in seconds, for every match in the process. - `Regexp.timeout` returns the current value; it starts as `nil`, meaning **no limit**. - The setting is global to the process, so set it once at boot rather than per request. - Any matching method that runs the regexp, such as `match?`, `=~`, `gsub` or `scan`, is covered. ## `timeout:` on `Regexp.new`: a per-pattern limit - `Regexp.new(source, timeout: 0.2)` attaches a limit to one Regexp; `re.timeout` reads it back. - A regexp literal has no syntax for a timeout, so a per-pattern limit means building the Regexp with `Regexp.new`. How the two combine: | `re.timeout` | `Regexp.timeout` | Result | |---|---|---| | `nil` | `nil` | never times out | | `nil` | a Float | times out after the global value | | a Float | anything | times out after the pattern's own value | When the limit passes, Ruby raises **`Regexp::TimeoutError`** with the message `regexp match timeout`. It inherits from `RegexpError`, which inherits from `StandardError`, so a plain `rescue` in the request handler catches it. ## `Regexp.linear_time?` and the memoization cache Ruby 3.2 introduced a cache-based optimization that lets many, but not all, patterns match in time linear in the input length. Ruby 3.3 extended it to lookarounds and atomic groups, provided they contain no captures and are not nested. `Regexp.linear_time?(re)` (or `Regexp.linear_time?(source, options)`) returns `true` when this interpreter can match `re` in linear time. Constructs that make it `false` include: - **backreferences** such as `\1` or `\k<name>`; - **subexpression calls** such as `\g<name>`; - **conditionals** such as `(?(1)…)`. The documentation stresses that the answer is a property of the interpreter, not of the pattern: another Ruby version or implementation may answer differently, and no compatibility is promised. Treat it as a check for the Ruby you run, re-verified on upgrade. ## A defensive setup for a licence-plate service 1. **Set a global backstop at boot**, for example `Regexp.timeout = 1.0`, so no forgotten pattern can hang a worker indefinitely. 2. **Assert linearity in tests** for every pattern you ship: `assert Regexp.linear_time?(PLATE)`. A pattern that fails the assertion gets rewritten or an explicit tight timeout. 3. **Give user-supplied patterns their own short timeout**: `Regexp.new(rule, timeout: 0.05)`, and reject rules that fail `linear_time?` before storing them. 4. **Escape user text** that is embedded in a pattern: `Regexp.new("\\A#{Regexp.escape(prefix)}")`, so a `.` or `*` in the text stays literal. 5. **Cap input length** before matching; a plate field never needs kilobytes. 6. **Rescue `Regexp::TimeoutError`** at the boundary, reject the input, and log the pattern name without the raw input. ```ruby Regexp.timeout = 1.0 # process-wide backstop PLATE = /\A[A-Z]{2}-\d{3}\z/ Regexp.linear_time?(PLATE) # => true def matches_rule?(rule_source, plate) rule = Regexp.new(rule_source, timeout: 0.05) rule.match?(plate) rescue Regexp::TimeoutError false end ``` ## Choosing the numbers - The **global** value is a safety net for every regexp in the process, including those inside libraries, so it should sit well above the slowest legitimate match; one second is a common starting point. - A **per-pattern** value for untrusted rules can be far smaller, tens of milliseconds, because a plate-sized input matched by a sane rule finishes in microseconds. - Measure before tightening: run the real patterns over the largest legitimate inputs and set the limit a comfortable multiple above that. ## Limits to keep in mind - A timeout bounds damage; it does not make matching cheap. Each timed-out request still spends its full budget of CPU. - `linear_time?` returning `true` means the engine can use the cache, not that it is free: the cache switches on only after a match has failed many times, and then allocates a bitmap that grows with the input length times the pattern's size. - Setting `Regexp.timeout` also affects libraries in the process, so choose a value generous enough for their legitimate patterns.

  • In Ruby, why might Regexp.linear_time?(/(\w+)-\1/) return false?
    The pattern uses a backreference, `\1`, and the memoization cache cannot handle backreferences, so the engine may backtrack without the linear-time guarantee. Rewrite without the backreference if possible, or give the pattern an explicit short `timeout:` and cap the input length.
  • In Ruby, does a per-regexp timeout: of nil switch off the global Regexp.timeout?
    No. A `nil` instance timeout falls through to `Regexp.timeout`; only a Float on the Regexp overrides the global value. A pattern escapes every limit only when both its own timeout and `Regexp.timeout` are `nil`, which is the default before you set anything.
  • In a Ruby service, what should the handler do when Regexp::TimeoutError is raised?
    Rescue it at the boundary, treat the input as invalid, and return a normal validation error rather than a 500. Log the pattern's name and the input length, not the raw input, and count occurrences: repeated timeouts point to an attack or to a pattern that needs rewriting.

saying these in an interview costs you the question

  • Ruby regexps have a built-in timeout by default
  • Regexp.linear_time? returning true is guaranteed on every Ruby version
  • Regexp::TimeoutError is not a StandardError, so a plain rescue misses it
  • A timeout makes a slow pattern cheap, so rewriting it is unnecessary
  • A regexp literal accepts a timeout: option after its flags