skip to content

Machine-Suggested Buckets

A model trained on this project's own past labels proposing a defect class for a new failure, and why a suggestion is never a verdict. Interviewers probe the cold start and the feedback trap.

on this pageshow

explore

questions

6

In ReportPortal, every failed test item's issue carries a defect group from `TestItemIssueGroup`. What does `TO_INVESTIGATE` mean beside `PRODUCT_BUG`, `AUTOMATION_BUG`, `SYSTEM_ISSUE` and `NO_DEFECT`, and which group does a brand-new failure land in?

level: juniorimportance: must knowfreq 66%

answer

  1. the bucket nobody has judged yet
  2. where a failure starts, not where it ends
  3. a queue, not a category
  4. stable locators, not display labels
  5. ti001 beside pb001 and ab001

basics

~10 s

TO_INVESTIGATE is ReportPortal's undecided defect group, and every new failure is filed there automatically. PRODUCT_BUG, AUTOMATION_BUG, SYSTEM_ISSUE and NO_DEFECT each record a decision already taken, by a person or by the auto-analyzer.

solid answer

~40 s

ReportPortal files every failed test item's issue under one defect group from `TestItemIssueGroup`: `PRODUCT_BUG`, `AUTOMATION_BUG`, `SYSTEM_ISSUE`, `NO_DEFECT` and `TO_INVESTIGATE`. Each constant carries a stable locator string — `pb001`, `ab001`, `si001`, `nd001`, `ti001` — and it is the locator, not the display label, that the API and the analyzer exchange. `TO_INVESTIGATE` is the default a failure arrives in: it is the absence of a decision, not a kind of defect, and the failure sits there until a person or auto-analysis moves it. That makes the size of that bucket a backlog figure rather than a category. The enum also holds `NOT_ISSUE_FLAG` with an empty locator, which is a marker rather than a bucket you file failures into.

go deeper

for a junior

Be able to name the five groups and say plainly that a new failure starts in TO_INVESTIGATE until a person or auto-analysis moves it somewhere else.

for a middle

Explain that each group carries a stable locator such as ti001 or pb001, and that the API and the analyzer exchange locators rather than the labels shown in the interface.

for a senior

Be ready to read a project where the undecided bucket never shrinks across builds, and to say what that means for triage capacity instead of blaming the analyzer.

for a principal

Own the policy question of how large the undecided queue is allowed to grow, and who is accountable for draining it before anyone quotes the defect counts in a release conversation.

## What a defect group actually is In ReportPortal a failed test item does not merely carry a status. It carries an **issue**, and the issue carries a **defect group** — one value drawn from `TestItemIssueGroup`. That group is the axis the defect counters and the triage screens aggregate along, so it is the single field that turns a wall of red into a short list of decisions. The groups you can file a failure under are `PRODUCT_BUG`, `AUTOMATION_BUG`, `SYSTEM_ISSUE`, `NO_DEFECT` and `TO_INVESTIGATE`. The enum additionally declares `NOT_ISSUE_FLAG` with an empty locator; treat that as a marker inside the vocabulary rather than a sixth bucket you file a failure into. ## The locator is the real identifier Every constant carries a short, stable **locator** string alongside its name: | group | locator | what the label records | |---|---|---| | `PRODUCT_BUG` | `pb001` | the thing under test is at fault | | `AUTOMATION_BUG` | `ab001` | the test is at fault | | `SYSTEM_ISSUE` | `si001` | something around the test is at fault | | `NO_DEFECT` | `nd001` | nothing is broken | | `TO_INVESTIGATE` | `ti001` | nobody has decided yet | Those glosses are one line each on purpose. **Deciding which of them is true for a given failure is defect triage, and it is not something the results tool does for you** — the tool's job is to represent the decision, and to represent honestly the state where no decision exists. The locator matters because it is what travels. When auto-analysis proposes a group it names a locator, and the server resolves that locator to the project's issue type before writing it onto the issue. The indirection means a display label can be changed without breaking a client, an importer, or the analyzer, and it means an API caller that sends a human-readable label where a locator is expected simply fails to resolve. ## Why TO_INVESTIGATE is not like the other four The other four are **verdicts**. `TO_INVESTIGATE` is the **absence of one**, and that difference is the whole reason the group exists: - A failure arrives already carrying `TO_INVESTIGATE`. Nothing has to run for it to get there; it is where a failure starts. - Because it is a real value and not a null, the undecided failures are countable, filterable and reportable like any other group. - Because it is countable, the size of the bucket is an honest backlog number. A launch with a hundred failures all in `TO_INVESTIGATE` is telling you that a hundred decisions are outstanding, not that the tool failed to categorise anything. - It is also the working set for machine suggestion. The undecided failures are exactly the pool a suggestion surface is asked to propose groups for; the already-decided ones are the evidence it proposes from. A candidate who says "unlabelled failures are just missing data" has missed the design. The undecided state is modelled deliberately so that *not knowing* is a first-class, visible outcome instead of an empty field somebody has to notice. ## Who moves a failure out of it Two actors, and only two: 1. **A person**, editing the issue on the item and choosing a group. 2. **Auto-analysis**, which reuses a group that a similar, already-labelled past failure carries. Both write into the same field, which is why ReportPortal keeps a separate boolean recording which of the two did it. The group alone cannot tell you — a machine-set `PRODUCT_BUG` and a hand-set `PRODUCT_BUG` are the same value in the same column. ## What the group does not tell you Be explicit about the limits of this one field, because interviewers push here: - It does not say **who or what** set it. That is a separate flag on the issue. - It does not carry a **confidence**. There is no score in the group itself; a guess and a certainty look identical. - It does not rank the failure, assign it, or schedule it. Severity, priority and the triage lifecycle live in the tracker, not in this enum. - It does not survive as an explanation. If you want to know *why* a failure was called an automation bug, the group is not where that lives — the issue's comment, the linked ticket and the activity trail are. ## How to talk about it A good answer names the five groups, says plainly that a new failure lands in `TO_INVESTIGATE`, and then makes the sharper point: the undecided group is a queue, not a category. From there you can go in either direction — towards what drains the queue, or towards how you tell a machine-set label from a hand-set one — and both are the questions the interviewer is usually steering at.

  • Why do the defect groups carry locator strings such as `ti001` and `pb001` at all?
    Because the group has to be named in wire traffic and storage without depending on a display label. Auto-analysis returns a locator, and the server resolves that locator to the project's issue type before writing it onto the issue. A label edited in the UI therefore cannot break analysis, an import, or an API client.
  • Does moving a failure out of `TO_INVESTIGATE` require deciding the product is at fault?
    No. `NO_DEFECT` exists precisely so a failure can leave the undecided queue without blaming the product or the test — an environment reset, expected data, a deliberate abort. What the model cares about is that a decision was recorded, so the failure stops counting as outstanding work.

saying these in an interview costs you the question

  • Calling TO_INVESTIGATE a kind of defect like the others
  • Assuming the display label is the API identifier
  • Thinking the analyzer invents new defect groups
  • Reading a large undecided bucket as a tool malfunction
  • Expecting the group itself to say who set it
open as a page

In ReportPortal, when auto-analysis puts a defect group on a failure it also sets `autoAnalyzed`. What does that boolean record, where does it live, and what does it deliberately not say?

level: middleimportance: must knowfreq 58%

basics

~20 s

autoAnalyzed records provenance — analysis set this defect group, not a person. It is a boolean on the issue rather than on the test item, stored in the auto_analyzed column, and it carries no confidence at all.

open as a page

In ReportPortal's `AnalyzerConfig`, what do `minShouldMatch` and `numberOfLogLines` control, and how does raising each one change the defect-group suggestions a project gets?

level: middleimportance: should knowfreq 50%

basics

~20 s

minShouldMatch is the similarity threshold a new failure must clear against an already-labelled one before that label is reused. numberOfLogLines caps how much of the failure's error log is compared. Raise either and you get fewer, safer suggestions.

open as a page

A ReportPortal project has just been created and its first launches are all red. Auto-analysis is enabled, yet no failure ever gets a defect group. What is happening, and how do you get the project out of it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A cold start. Auto-analysis reuses defect groups from failures somebody already labelled, and a new project has none, so everything stays in TO_INVESTIGATE. The way out is to label a representative sample by hand, then re-run analysis over the undecided items.

open as a page

For a ReportPortal project, would you let auto-analysis write the defect group straight onto incoming failures, or keep it as a proposal a person accepts — and how do you stop wrong machine labels feeding the next round?

level: principalimportance: should knowfreq 40%

basics

~20 s

Keep it as a proposal a person accepts, and keep the record of who set every label. Applied groups become the evidence for the next round, so a wrong label compounds unless you can find the machine-set ones and reset them.

open as a page

In ReportPortal, a proposed defect group comes back as a `SuggestInfo` record carrying `matchScore`, `esScore`, `resultPosition`, `esPosition`, `usedLogLines`, `minShouldMatch` and `userChoice`. What is that record for?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A SuggestInfo record is the receipt for one proposal: two ranking scores and the positions they gave the candidate, the settings the match ran under, how long it took, which model produced it, and what the person eventually chose.

open as a page