skip to content

How does Cucumber-JVM decide whether a step definition string is a Cucumber Expression or a regex?

level: middleimportance: must knowfreq 68%

answer

  1. two pattern languages, one annotation
  2. look at the two ends of the string
  3. anchors and forward slashes
  4. in cucumber-js the argument type decides
  5. capture groups become method arguments

basics

~20 s

Cucumber-JVM reads the annotation string as a regular expression when it starts with a caret, ends with a dollar sign, or is wrapped in forward slashes; otherwise it parses it as a Cucumber Expression. In cucumber-js the argument type decides.

solid answer

~40 s

Two pattern languages share one annotation. **Cucumber-JVM** infers which: the string is compiled as a **regular expression** when it starts with `^`, ends with `$`, or is wrapped in `/.../`, and is parsed as a **Cucumber Expression** otherwise. **cucumber-js** does not infer — `Given(/^.../ , fn)` passes a `RegExp` and is a regex, `Given('...', fn)` passes a string and is an expression. What each buys: an expression reads like the step line and hands the method typed arguments through `{int}`, `{string}`, `{word}` and custom types; a regex buys character classes, quantifiers and backreferences, at the cost of readability and with one method argument per **capture group** — so a group added to pluralise a word changes the method's signature.

code

java · 9 lines
java
@Given("turbine {word} is scheduled for maintenance in {int} days")
public void turbineScheduled(String turbineId, int days) {
    plan.schedule(turbineId, days);
}

@Given("^turbine ([A-Z]-[0-9]{3}) is scheduled for maintenance in ([0-9]+) days$")
public void turbineScheduledRegex(String turbineId, int days) {
    plan.schedule(turbineId, days);
}

go deeper

for a junior

Be ready to say that the text in a step-definition annotation can be written in two ways, and that the readable one with curly-brace placeholders is called a Cucumber Expression.

for a middle

Explain the actual inference rule in Cucumber-JVM — anchors or slash delimiters mean regex, anything else is an expression — and that in cucumber-js the argument type settles it instead.

for a senior

Show the consequence you have hit in production: a capture group changing a method's arity, or a loose regex claiming a sentence someone else wrote as an expression and turning the step ambiguous.

for a principal

Own the convention: which syntax is the default across teams, when regex is allowed, and how review catches a pattern loose enough to poach another team's steps.

## Two pattern languages behind one annotation A step definition binds a line of a Gherkin scenario to code through a **pattern string** — the text inside `@Given`/`@When`/`@Then` in Cucumber-JVM, or the first argument of `Given(...)` in cucumber-js. Cucumber accepts two pattern languages for that string. A **Cucumber Expression** is the readable syntax with `{int}`-style placeholders, optional text and alternation. A **regular expression** is the older syntax, with character classes, quantifiers and capture groups. Both are compiled into a matcher that the runner applies to the step text when it collects the candidate definitions for a line, so the choice is about how you author the pattern, not about what the runner is able to match. ## How Cucumber-JVM decides Cucumber-JVM receives one string and has to infer which language it is written in. It reads the string as a **regular expression** when the string looks like one: - it begins with the anchor `^`; - it ends with the anchor `$`; - it is wrapped in forward slashes, `/like this/`. Anything else is parsed as a **Cucumber Expression**. Two consequences matter in practice. First, anchoring is how you opt in to regex deliberately — `@Given("^turbine ([A-Z]-[0-9]{3}) is offline$")` is unambiguously a regex, which is why teams that use regex anchor every one of them by convention. Second, the inference can misfire on ordinary prose: a step whose text genuinely ends in a dollar sign is read as an anchored regex and then behaves nothing like the sentence you wrote. Reword the step, or commit to regex and write the whole pattern as one. ## How the other implementations decide **cucumber-js** does not guess. The distinction is carried by the JavaScript type of the argument: `Given(/^the planner opens turbine [A-Z]-[0-9]{3}$/, fn)` passes a `RegExp` object and is a regex; `Given('the planner opens turbine {word}', fn)` passes a string and is a Cucumber Expression. Nothing about the contents changes that. **Behave** selects a matcher with a `use_step_matcher` call rather than inspecting each pattern: its default `parse` matcher understands `{name:Type}` placeholders, and a regular-expression matcher can be switched on instead, applying to the decorators that follow. **SpecFlow/Reqnroll** put the pattern in a `[Given]`-style attribute on a binding method; SpecFlow's attribute string is a regular expression, and Reqnroll also accepts Cucumber Expressions. ## What each one buys | Concern | Cucumber Expression | Regular expression | |---|---|---| | Readability | reads close to the step line itself | pattern punctuation buried in prose | | Typed arguments | `{int}`, `{float}`, `{word}`, `{string}`, custom types | captured text converted to the parameter's type | | Wording variants | `hour(s)`, `logs/records` built in | non-capturing groups written by hand | | Method arity | one argument per placeholder | one argument per capture group, accidental ones included | | Raw power | no lookahead, no backreferences | the whole regex engine | The arity row is the one that bites. In a regex every capture group becomes an argument, so adding `(s)` to pluralise a noun silently demands another parameter on the method; in a Cucumber Expression `(s)` is optional text and captures nothing. Conversely, a regex can express things an expression cannot — a strict identifier shape, a bounded repetition, an alternation of whole phrases. ## Choosing, and mixing, on a real suite On a 63-scenario wind-farm maintenance planner suite that two teams both edit, mixed syntaxes are the usual source of confusion, because both syntaxes register into the same candidate set for a step line: 1. **Default to Cucumber Expressions.** They read like the sentence, they give typed arguments without conversion code, and they make a reviewer's job easier. 2. **Reach for regex on purpose**, not by accident — a constrained identifier such as `[A-Z]-[0-9]{3}`, or a step whose text contains characters the expression syntax reserves (`{`, `(`, `/`). 3. **Write one convention down** so a reviewer can tell at a glance which language a line is in; on Cucumber-JVM that convention is usually "regex is always anchored". 4. **Keep patterns tight.** A deliberately loose regex written by one team is what later matches a sentence the other team wrote as an expression, and the run reports the step as ambiguous. 5. **Convert rather than accumulate.** When a regex exists only to capture a number or a quoted value, rewriting it as an expression usually deletes conversion code from the method body as well. The question is a good interview filter because the answer is not a preference. It is a rule you can state exactly, and the person who has hit the trailing-dollar case or the accidental capture group will say so unprompted.

  • Why do teams that use regex step definitions anchor every one of them?
    Anchoring pins the pattern to the whole step line rather than to part of it, so a broad pattern cannot also claim a longer sentence that merely contains the same words. On Cucumber-JVM the anchors do a second job: they are the signal that makes the runner read the string as a regex at all, so anchoring removes any doubt about which pattern language a reviewer is reading.
  • Can one suite mix Cucumber Expression and regex step definitions?
    Yes — the choice is per definition, and both compile into the same candidate set for a step line. That is exactly why mixing needs a convention: a loose regex from one file and a readable expression from another can both match the same sentence, and the runner then reports the step as ambiguous rather than picking one.
  • What happens to a Cucumber-JVM step whose text legitimately ends in a dollar sign?
    The runner sees the trailing dollar as an anchor and reads the whole string as a regular expression, so the placeholders and optional text you wrote are no longer interpreted as an expression. The fix is to reword the step text, or to accept the inference and write the pattern deliberately as a complete anchored regex.

It is the difference between writing the sentence and writing a search pattern for it: one is meant to be read back by a person, the other by the regex engine.

saying these in an interview costs you the question

  • Thinks the annotation string is always a regular expression
  • Claims every Cucumber Expression should be anchored with caret and dollar
  • Assumes adding a regex capture group leaves the method signature unchanged
  • Believes cucumber-js inspects the string contents to detect a regex
  • Says a project must pick one syntax and cannot mix them