skip to content

In Kotest, does the string matcher shouldMatch require the whole string to match the regex or just part of it — and how does that differ from shouldContain with a Regex, shouldStartWith and shouldContainIgnoringCase?

level: middleimportance: should knowfreq 32%

answer

  1. shouldMatch = whole string
  2. shouldContain(Regex) = search anywhere
  3. Dot does not cross newlines by default
  4. Literal: startsWith/endsWith/containIgnoringCase
  5. Raw strings to avoid double escaping

basics

~20 s

Kotest's shouldMatch requires the entire string to match the regex — it is a full match, not a search. To assert that a pattern occurs somewhere inside a string, pass a Regex to shouldContain, which looks for a match anywhere. shouldStartWith/shouldEndWith are literal, case-sensitive; shouldContainIgnoringCase is a case-insensitive substring check.

solid answer

~50 s

The key semantic is anchoring. - `str shouldMatch "[a-z]+@example\\.com"` — the **whole** string must match. `"mail: [email protected]"` fails, because the prefix is not covered by the pattern. There are overloads taking either a `String` pattern or a `Regex`. - `str shouldContain Regex("[a-z]+@example\\.com")` — a **partial** match: the pattern must occur somewhere in the string. This is the one you want for log lines and messages. - `str shouldContain "substring"` — literal, case-sensitive substring. - `str shouldContainIgnoringCase "substring"` — same, case-insensitive. There are also `shouldBeEqualIgnoringCase` and `shouldStartWith` / `shouldEndWith` (literal and case-sensitive). The common bug is writing `shouldMatch(".*ERROR.*")` to search a string: it works only because of the wildcards, and it silently fails against multi-line input because `.` does not match a newline by default. Use `shouldContain(Regex(...))` for searching, and add `RegexOption.DOT_MATCHES_ALL` if you genuinely need `.` to span lines.

code

kotlin · 7 lines
kotlin
"[email protected]" shouldMatch "[a-z]+@example\\.com"           // whole string
"mail: [email protected]" shouldContain Regex("[a-z]+@example") // search

header shouldStartWith "Bearer "          // literal, case-sensitive
message shouldContainIgnoringCase "not found"

output shouldMatch Regex(".*ERROR.*", RegexOption.DOT_MATCHES_ALL)

go deeper

for a junior

Know that shouldMatch checks the whole string while shouldContain with a Regex searches, and that there are literal prefix/substring matchers.

for a middle

Explain the anchoring difference with an example, name the case-insensitive matchers, and mention that . does not cross newlines.

for a senior

Argue for literal matchers over regexes wherever the expectation is literal, and diagnose the multi-line and escaping failure modes on sight.

for a principal

Treat it as an assertion-hygiene rule: regexes in tests are code with their own defect modes, so keep them anchored, raw-stringed and reserved for genuinely structural expectations.

## The one fact to get right Kotest's `shouldMatch` is a **full-string** assertion. The regex has to describe the entire value, not a fragment of it. This trips people who come from tools where the equivalent method searches. ```kotlin "[email protected]" shouldMatch "[a-z]+@example\\.com" // passes "mail: [email protected]" shouldMatch "[a-z]+@example\\.com" // FAILS "mail: [email protected]" shouldContain Regex("[a-z]+@example\\.com") // passes ``` So: **`shouldMatch` = describe the whole thing; `shouldContain(Regex)` = find it somewhere.** Both accept the pattern as a `Regex`, and `shouldMatch` additionally accepts a `String` that it treats as a pattern. ## Why the distinction matters in practice Two failure modes come out of confusing them. **False failures.** A test asserts a log line matches `"user \\d+ created"` while the real line is `"2026-01-01 INFO user 42 created"`. The assertion fails, and the developer's instinct is to wrap the pattern in `.*` on both sides. That works, but it hides that the intent was a search all along — and it invites the next problem. **False passes and multi-line surprises.** By default in Kotlin regex, `.` does not match a line terminator. A full-match assertion of `".*ERROR.*"` against a multi-line captured output fails even though the text plainly contains `ERROR`. The correct expressions are either a search: ```kotlin output shouldContain Regex("ERROR") ``` or, if you really need a whole-string pattern over multiple lines, a regex constructed with the option that lets `.` span newlines: ```kotlin output shouldMatch Regex(".*ERROR.*", RegexOption.DOT_MATCHES_ALL) ``` ## The literal matchers Alongside the regex pair, Kotest gives literal string matchers: - `shouldStartWith(prefix)` / `shouldEndWith(suffix)` — literal, case-sensitive. No regex interpretation happens, so a prefix containing `.` or `[` is taken at face value. That is usually what you want and is faster to read than an anchored regex. - `shouldContain(substring)` — literal, case-sensitive containment. - `shouldContainIgnoringCase(substring)` — containment ignoring case, for assertions on human-facing text where capitalisation is not part of the contract. - `shouldBeEqualIgnoringCase(other)` — whole-string equality ignoring case. A useful rule: reach for the literal matcher whenever the expectation is literal. `header shouldStartWith "Bearer "` says exactly what it means; `header shouldMatch "Bearer .*"` says the same thing with more ways to go wrong (unescaped metacharacters, forgotten anchoring semantics, multi-line behaviour). ## Case-insensitivity: matcher versus regex option There are two ways to ignore case, and they are not interchangeable: ```kotlin message shouldContainIgnoringCase "not found" message shouldContain Regex("not found", RegexOption.IGNORE_CASE) ``` The matcher form is clearer for a literal substring; the regex option is what you need when the pattern itself has structure. Do not write a regex just to get case insensitivity for a plain literal. ## Escaping, in Kotlin Because patterns are written as Kotlin strings, a literal dot is `"\\."` in a normal string. Raw strings avoid the double escaping and read much better for anything non-trivial: ```kotlin id shouldMatch Regex("""[0-9a-f]{8}-[0-9a-f]{4}""") ``` Forgetting to escape `.` is the classic source of an over-permissive pattern that passes on wrong input. ## Choosing between them 1. Is the expectation literal? Use `shouldStartWith` / `shouldEndWith` / `shouldContain` / `shouldContainIgnoringCase`. 2. Is the expectation structural and about the whole value (an id format, a full URL shape)? Use `shouldMatch`. 3. Is the expectation structural but about a fragment (a log line inside captured output)? Use `shouldContain(Regex(...))`. 4. Multi-line input plus `.` in the pattern? Decide explicitly about `DOT_MATCHES_ALL`. ## Interview-ready summary `shouldMatch` is anchored to the whole string; `shouldContain(Regex)` searches. Literal prefixes and substrings have their own matchers, including a case-insensitive containment matcher, and those are preferable to regexes whenever the expectation really is literal. The `.*`-wrapping habit is the tell that someone has not internalised the distinction.

  • A test asserts a multi-line captured output with Kotest's shouldMatch(".*ERROR.*") and it fails even though ERROR is clearly in the text. Why?
    Two reasons compound. `shouldMatch` requires the whole string to match, and by default `.` in a Kotlin regex does not match line terminators, so a pattern of `.*ERROR.*` cannot span the newlines. Either search with `shouldContain(Regex("ERROR"))` or build the regex with `RegexOption.DOT_MATCHES_ALL`.
  • Why prefer shouldStartWith over an anchored regex for a literal prefix?
    Because it says exactly what is meant and cannot be broken by regex metacharacters. A literal prefix containing `.`, `(` or `[` would need escaping in a pattern, and forgetting that produces an over-permissive assertion that passes on wrong input. The literal matcher also produces a clearer failure message.

saying these in an interview costs you the question

  • Assuming shouldMatch searches for the pattern anywhere in the string
  • Wrapping every pattern in .* instead of switching to shouldContain(Regex)
  • Expecting `.` to match newlines in multi-line assertions
  • Writing a regex purely to get case-insensitive matching for a literal substring
  • Forgetting to escape a literal dot and shipping an over-permissive pattern

context