In a JMeter Response Assertion scanning a 200 KB body, why is Substring cheaper than Contains?
answer
- Two of the four rules are not regexes
- Compilation is cached; matching is not
- Literal scan versus a backtracking engine
- Substring maps to String.contains
basics
~10 sSubstring is a plain literal scan of the text. Contains hands the whole body to a Perl5 regular-expression matcher instead. Same text, much more work per sample, repeated on every sample every thread makes.
solid answer
~50 sJMeter's Response Assertion has four Pattern Matching Rules and they split into two families. `Equals` and `Substring` are plain, case-sensitive string comparisons — internally `String.equals` and `String.contains`. `Contains` and `Matches` are Perl5 regular expressions, evaluated by Jakarta ORO's `Perl5Matcher` by default, with `Contains` searching anywhere in the text and `Matches` requiring the whole text to match. Compilation is *not* the recurring cost: the compiled pattern is kept in a shared LRU cache sized by `oro.patterncache.size`, default 1000, so a plan with a handful of patterns compiles each one once. The recurring cost is the match itself, run once per sample, on a matcher each thread owns privately. On a 200 KB body an unanchored regex can be retried from many offsets, while a literal scan is a straight walk. When the check is a literal, `Substring` gives the same verdict for less injector CPU.
code
xml · 9 lines<ResponseAssertion guiclass="AssertionGui" testclass="ResponseAssertion" testname="RA" enabled="true">
<collectionProp name="Asserion.test_strings">
<stringProp name="657066065">Apache JMeter</stringProp>
</collectionProp>
<stringProp name="Assertion.test_field">Assertion.response_data</stringProp>
<boolProp name="Assertion.assume_success">false</boolProp>
<intProp name="Assertion.test_type">16</intProp>
<stringProp name="Assertion.custom_message"></stringProp>
</ResponseAssertion>go deeper
Know that Contains and Matches are regular expressions while Equals and Substring are plain case-sensitive text. Picking the plain rule for a literal check is the everyday habit to form.
Explain the mechanics: the compiled pattern is cached, the match is not, and an unanchored regex may restart at many offsets in the body while a literal scan is a single walk.
Bring in scale: the same choice multiplied by every sampler and every thread in a 500-thread plan, and the field-selection lever that shrinks the input before the rule ever runs.
Set the house style — literal checks use Substring, regexes must be anchored and reviewed — so that plans written by different people do not each rediscover this at load.
## The two families of matching rule The Response Assertion's *Pattern Matching Rules* radio offers four values, and the JMeter manual is explicit that they are not four flavours of the same thing: - **`Contains`** — true if the text contains the **regular expression** pattern. Unanchored. - **`Matches`** — true if the **whole** text matches the regular expression pattern. Anchored end to end. - **`Equals`** — true if the whole text equals the pattern **string**, case-sensitive. Plain text. - **`Substring`** — true if the text contains the pattern **string**, case-sensitive. Plain text. The manual states it directly: *"Equals and Substring patterns are plain strings, not regular expressions."* In the element's own code that difference is visible as two completely different call paths — `Substring` reduces to `toCheck.contains(stringPattern)`, while `Contains` reduces to a `Perl5Matcher.contains(toCheck, pattern)` call against a compiled ORO pattern. ## What is cached and what is not A common wrong answer is "the regex is slower because it has to be compiled every time." It does not. | Thing | Cached? | Where | |---|---|---| | The compiled ORO pattern | Yes | A shared LRU cache; size set by `oro.patterncache.size`, default 1000 | | The compiled `java.util.regex` pattern, when `jmeter.regex.engine` is not `oro` | Yes | A separate cache sized by `jmeter.regex.patterncache.size`, default 1000 | | The matcher object | Per thread | `Perl5Matcher` is held in a thread-local, one per JMeter thread | | The decoded response body | Per sample | `SampleResult` decodes the bytes once and reuses the string for later assertions on the same sample | | **The match itself** | **No** | **Re-run for every sample, on every thread** | So compilation is a one-off, and the interesting corollary of the cache being an LRU is that a plan carrying **more distinct patterns than the cache size** will start evicting and recompiling. With one assertion in a 500-thread plan that never happens; with generated per-thread patterns it can. ## Where the per-sample cost actually goes What repeats is the scan, and its shape depends on the rule you picked: - **`Substring`** walks the text looking for a literal. Cost tracks the length of the body. - **`Contains`** runs a backtracking Perl5 engine that is free to restart at successive offsets in the body. A pattern that begins with a wildcard, or that contains nested quantifiers, can do far more work than the body's length suggests. - **`Matches`** anchors the pattern to the entire text, which sounds cheaper but usually is not: to say "no", the engine still has to fail across the whole document, and authors typically wrap the real pattern in `.*` on both sides to make it fit, reintroducing exactly the wildcard scanning they were trying to avoid. Multiply any of those by the divergence that motivates this topic — an assertion attached to every sampler in a 500-thread plan — and you get up to 500 concurrent scans of a 200 KB body, per sampler, per iteration. ## Choosing the rule deliberately A short checklist that costs nothing to apply: 1. If the pattern contains no regex metacharacters, use `Substring`. `Contains` with the literal `Apache JMeter` is a regular expression that happens to contain no operators — you are paying the engine for nothing. 2. If you need case-insensitivity or alternation, you genuinely need a regex; keep the pattern anchored to something concrete rather than starting it with a wildcard. 3. Point *Field to Test* at the smallest field that can answer the question. `Response Code` is three characters; `Text Response` is the whole body. 4. Order the patterns in one assertion so the most likely failure is first — the manual notes that *"If a pattern fails, then further patterns are not checked"*, so an early failure short-circuits the rest. ## The trap in the small print `Substring` and `Equals` are **case-sensitive** with no option to change that, and there is no plain-text equivalent of the regex `(?i)` switch. If your check must ignore case, you are back on `Contains` and you should accept the cost knowingly rather than reach for `Substring` and quietly weaken the check.
- If compilation is cached, when would a JMeter plan actually recompile the same kind of pattern repeatedly?When the plan uses more distinct pattern strings than the cache holds. The ORO cache is an LRU sized by `oro.patterncache.size`, default 1000, so a pattern built from a variable — a per-thread or per-row value interpolated into the assertion — produces a new cache key each time and can evict entries faster than they are reused.
- Does picking Matches instead of Contains make the check cheaper because it stops at the first character?No. `Matches` requires the whole text to match, so a failing check still has to be refuted across the entire document. Authors also tend to pad the pattern with wildcards on both sides so it spans the body, which restores the scanning cost `Contains` would have had.
saying these in an interview costs you the question
- Says Contains is slower because the regex is recompiled per sample
- Treats Substring and Contains as synonyms
- Thinks Matches is cheaper because it can fail fast
- Assumes Substring is case-insensitive
- Ignores the size of the field being scanned