skip to content

In ReportPortal's `AnalyzerConfig`, what do `minShouldMatch` and `numberOfLogLines` control, and how does raising each one change the defect-group suggestions a project gets?

level: middleimportance: should knowfreq 50%

answer

  1. two knobs, one bar and one window
  2. how alike before a label is reused
  3. how much log the matcher sees
  4. exposed as analyzer.minShouldMatch
  5. raise it and more failures stay undecided

basics

~20 s

minShouldMatch is the similarity threshold a new failure must clear against an already-labelled one before that label is reused. numberOfLogLines caps how much of the failure's error log is compared. Raise either and you get fewer, safer suggestions.

solid answer

~40 s

Both are fields on `AnalyzerConfig`, and both are exposed as project settings under the `analyzer.` prefix — `analyzer.minShouldMatch` and `analyzer.numberOfLogLines`. `minShouldMatch` is the bar a candidate has to clear: how much of the incoming failure's indexed text must agree with an already-labelled failure before its group is reused. `numberOfLogLines` sizes the window — how many of the failure's error-level log lines are handed to the matcher in the first place. Raising the threshold yields fewer proposals and leaves more failures in `TO_INVESTIGATE`; lowering it empties the queue with labels nobody vouched for. They are two of eight fields: `searchLogsMinShouldMatch`, `isAutoAnalyzerEnabled`, `analyzerMode`, `allMessagesShouldMatch`, `indexingRunning` and `largestRetryPriority` sit beside them, and the whole config travels with the analysis request.

code

json · 10 lines
json
{
  "isAutoAnalyzerEnabled": true,
  "analyzerMode": "CURRENT_AND_THE_SAME_NAME",
  "minShouldMatch": 95,
  "searchLogsMinShouldMatch": 95,
  "numberOfLogLines": 5,
  "allMessagesShouldMatch": false,
  "indexingRunning": false,
  "largestRetryPriority": false
}

go deeper

for a junior

Know that ReportPortal has project settings controlling how similar two failures must be before a label is reused, and how much of the log is compared.

for a middle

Explain both directions: raising the threshold means fewer and safer proposals, and the log window can be too small as easily as too large.

for a senior

Demonstrate that you would tune by the share of proposals people accept rather than by how empty the undecided queue looks, and change one setting at a time.

for a principal

Frame these settings as a policy choice about how much unverified labelling the organisation is willing to carry, and decide who owns changing them across many projects.

## Two knobs that decide whether you get a suggestion at all ReportPortal's analysis reuses a defect group that a **similar past failure** already carries. Two settings on `AnalyzerConfig` govern that word "similar" from opposite ends: - **`minShouldMatch`** is the **threshold**. It sets how much of the incoming failure's indexed text has to agree with a stored, already-labelled failure before that failure counts as a candidate and its group is reusable. - **`numberOfLogLines`** is the **window**. It caps how many of the failure's error-level log lines are handed to the matcher in the first place. One decides how alike is alike enough; the other decides how much text "alike" is measured over. Change either and you change the population of suggestions, without touching a model. ## What raising minShouldMatch does Raising the threshold makes the analyzer more conservative: 1. Fewer candidates clear the bar, so fewer failures get a group proposed. 2. More failures stay in `TO_INVESTIGATE`, and the undecided queue grows. 3. The proposals you do get are closer matches, so the accept rate should rise. Lowering it does the opposite, and this is where teams hurt themselves. A queue that empties overnight after somebody dropped the threshold has not been triaged — it has been *coloured in*. Each of those weak labels is then indexed as evidence, so the next round of failures matches against them, and a mistake propagates outward instead of staying put. Because the labels carry a provenance flag, the damage is at least findable and resettable; without that flag it would not be. There is a related setting, `searchLogsMinShouldMatch`, that applies a threshold to a different operation. Do not tune one believing you are tuning the other. ## What numberOfLogLines changes The window is not a "more is better" dial: - **Too small** and the matcher sees only the opening of a failure. Two genuinely different failures whose first lines are identical framework noise look the same. - **Too large** and every failure drags in stack frames, thread dumps and teardown chatter that are shared by everything the suite runs. Common text swamps the distinguishing text, and unrelated failures start matching each other. The window also interacts with the threshold. Widening the window changes what a given percentage means, so a threshold that behaved well over a few lines can become far stricter or far looser once the window grows. Change one at a time, and change nothing without a way of judging the result. ## The other six fields `AnalyzerConfig` carries eight fields in total, and an interviewer will not thank you for implying it carries two: | field | what it governs | |---|---| | `minShouldMatch` | the similarity bar for reusing a label | | `searchLogsMinShouldMatch` | the bar for a different search operation | | `numberOfLogLines` | how much of the failure's log is compared | | `isAutoAnalyzerEnabled` | whether analysis runs at all | | `analyzerMode` | how wide a slice of past launches is searched | | `allMessagesShouldMatch` | whether every message must agree | | `indexingRunning` | server-maintained indexing state, not a knob | | `largestRetryPriority` | how retried attempts are weighed | `analyzerMode` is worth calling out separately because its accepted wire values are `ALL`, `LAUNCH_NAME`, `CURRENT_LAUNCH`, `PREVIOUS_LAUNCH` and `CURRENT_AND_THE_SAME_NAME`. They differ in how much history is in scope, which is a third lever on the same outcome — and it is the one that matters most on a project with little history. ## Why the whole config rides along with the request The analysis request carries the project's full `AnalyzerConfig` rather than leaving the analyzer to look it up later. The effect is that a suggestion is produced **under a stated configuration**, and the suggestion record echoes parts of it back — the `minShouldMatch` that was in force and the `usedLogLines` that were actually compared. A month later you can look at a proposal and know the settings it was made under, instead of assuming today's settings applied. That is what makes these two knobs tunable in any honest sense. Without the echo you would be changing numbers and looking at a queue size, which tells you how many labels were applied and nothing at all about whether they were right. ## Tuning them without a scoreboard A sound approach, and a good thing to say out loud: - Start conservative. A short undecided queue full of wrong labels is worse than a long one full of honest unknowns. - Judge by **acceptance**, not by queue size — what share of proposals a person keeps. - Move one setting at a time, and give it enough launches to show an effect. - Remember that neither knob touches the model. They shape the evidence and the bar; they are not training.

  • A team drops `minShouldMatch` and the undecided queue empties overnight. Why is that not a win?
    Because the queue emptied by lowering the bar, not by anyone deciding anything. Each weak label is then indexed as evidence for the next similar failure, so one bad match seeds more. Judge the change by the share of proposals people accept, and be ready to reset the machine-set labels if the accept rate fell.
  • Why does the analysis request carry the whole `AnalyzerConfig` instead of the analyzer reading it later?
    So the proposal is reproducible. The suggestion is produced under a stated configuration, and the record echoes back the `minShouldMatch` in force and the log lines actually used. A reader months later can tell which settings produced a label, rather than assuming today's settings were the ones that applied.

saying these in an interview costs you the question

  • Thinking minShouldMatch tunes the model itself
  • Assuming more log lines always mean better matches
  • Lowering the threshold to shrink the undecided queue
  • Believing these two are the whole analyzer config
  • Confusing minShouldMatch with searchLogsMinShouldMatch