A ReportPortal project has just been created and its first launches are all red. Auto-analysis is enabled, yet no failure ever gets a defect group. What is happening, and how do you get the project out of it?
answer
- no history, nothing to match
- the label pool starts empty
- seed it by hand first
- scope of past launches is configurable
- then re-run over the undecided set
basics
~20 sA cold start. Auto-analysis reuses defect groups from failures somebody already labelled, and a new project has none, so everything stays in TO_INVESTIGATE. The way out is to label a representative sample by hand, then re-run analysis over the undecided items.
solid answer
~40 sNothing is broken — the label pool is empty. ReportPortal's analysis proposes a defect group by finding an **already-labelled** past failure that the new one resembles; with no labelled history there are no candidates, so every failure stays in `TO_INVESTIGATE`. `analyzerMode` compounds it: its wire values `ALL`, `LAUNCH_NAME`, `CURRENT_LAUNCH`, `PREVIOUS_LAUNCH` and `CURRENT_AND_THE_SAME_NAME` differ in how wide a slice of past launches is in scope, and a narrow scope on a fresh project sees nothing at all. The way out is to label by hand first — a representative sample across the failure families you actually get — and then re-run analysis over the undecided items with `AnalyzeItemsMode` `TO_INVESTIGATE`, which leaves the hand-made labels untouched.
go deeper
Understand that ReportPortal's suggestions come from failures your own team already labelled, so a brand-new project has nothing for the analyzer to reuse yet.
Walk the diagnosis: analysis enabled, candidates available, how wide the search scope is, per-failure exclusions, and whether indexing has caught up.
Show the operating loop — seed a broad sample by hand, re-run over the undecided items only, and refuse to fake progress by lowering the similarity bar.
Decide what a new project owes the ones beside it: whether seeding is a launch checklist item, who reviews early labels, and how you avoid a per-team drift in what each defect group means.
## Diagnose it before you tune it The symptom — analysis on, nothing labelled — looks like a misconfiguration and almost never is. Work the chain in order: 1. **Is analysis enabled?** `isAutoAnalyzerEnabled` on the project's `AnalyzerConfig` is the on switch. Check it, then stop suspecting it. 2. **Is there anything to match against?** This is the real answer on a new project. Analysis reuses a group that a past, already-labelled failure carries. Zero labelled failures means zero candidates. 3. **How wide is the search?** `analyzerMode` decides the slice of history in scope, with wire values `ALL`, `LAUNCH_NAME`, `CURRENT_LAUNCH`, `PREVIOUS_LAUNCH` and `CURRENT_AND_THE_SAME_NAME`. A narrow scope over an empty history is doubly empty. 4. **Is anything excluded?** `ignoreAnalyzer` on an issue keeps that failure out of analysis. It is per-failure, so it explains gaps, not a total blank. 5. **Is the index still catching up?** The config carries `indexingRunning` as server-maintained state; freshly ingested data is not instantly matchable. Only after all five would I look at `minShouldMatch`, and even then a high threshold cannot be the cause when the candidate set is empty. ## Why this is not a defect The design is worth defending out loud, because the interviewer is often testing whether you will blame the tool. A suggestion surface that reuses *your project's own past decisions* has one unavoidable property: it cannot propose a decision nobody has made yet. The alternative — shipping opinions learned from other people's projects — would produce confident labels that mean nothing about your codebase, and they would be indistinguishable from labels your team actually stands behind. So the empty state is honest. Every failure sitting in `TO_INVESTIGATE` is the tool saying "I have no basis for an opinion", which is exactly what is true. ## Getting out of it The cold start is broken by people, and then handed back to the machine: - **Label a representative sample, not the easy ones.** Cover the failure families you actually get — the environment flake, the assertion mismatch, the timeout, the setup failure — rather than fifty instances of the same broken test. Coverage of *kinds* is what makes future matches possible; volume of one kind does not. - **Label carefully at the start.** Early labels are disproportionately influential because they are the only evidence there is. A sloppy first week becomes the project's whole notion of what an automation bug looks like. - **Then re-run over the undecided set.** An analyze run takes an `AnalyzeItemsMode`; `TO_INVESTIGATE` collects the undecided failures and tries to label them, and it does not disturb anything already decided. `AUTO_ANALYZED` and `MANUALLY_ANALYZED` exist for the other two populations, and a request may ask for those three. - **Widen the scope while history is thin, then narrow it.** The broadest `analyzerMode` value is the only one with anything in it early on; narrowing later trades recall for relevance once each launch name has its own history. - **Do not compensate by dropping `minShouldMatch`.** Lowering the bar against an empty or tiny candidate set does not produce good labels, it produces confident bad ones, and those become the evidence for everything that follows. ## What to watch as it warms up | signal | what it tells you | |---|---| | undecided queue shrinking | analysis is finding candidates at all | | share of proposals people accept | whether the candidates are the right ones | | labels concentrated in one group | the seed sample was probably lopsided | | queue empties immediately after a settings change | somebody lowered the bar rather than triaging | The first two are the ones to quote. Queue size on its own is the metric that flatters a badly configured project, because there are two ways to empty a queue and only one of them is triage. ## The shape of a good answer Name the cold start, explain why an empty label pool produces exactly this symptom, and describe the seed-then-re-run loop. Then show the judgement: seeding needs breadth rather than volume, early labels carry outsized weight, and the temptation to force suggestions by loosening the threshold makes the project permanently worse rather than temporarily better.
- Which `analyzerMode` value would you start a fresh project on, and why?The widest scope. The five wire values — `ALL`, `LAUNCH_NAME`, `CURRENT_LAUNCH`, `PREVIOUS_LAUNCH`, `CURRENT_AND_THE_SAME_NAME` — differ in how much history is searched, and early on the broad end is the only one with anything in it. Narrow it once each launch name has its own history, trading recall for relevance.
- The team has now labelled a few dozen failures by hand. What exactly do you run?An analyze run scoped with `AnalyzeItemsMode` `TO_INVESTIGATE`. That collects the still-undecided failures and tries to match them against the new labels, and it leaves the hand-made decisions alone. Running it over `MANUALLY_ANALYZED` instead would put the very work you just did back in play.
- Would you seed the labels by importing another project's history?No. Reused labels from a different codebase produce confident groups that say nothing about this product's failures, and once applied they are indistinguishable from decisions your team stands behind. The cold start is short; the credibility damage from labels nobody vouched for is not.
It is a triage nurse's first shift at a brand-new clinic. The protocol is sound and the nurse is competent, but there are no past charts to compare this patient against, so everything goes to the doctor until enough cases have been written down.
saying these in an interview costs you the question
- Blaming the model when the label pool is empty
- Widening the search scope instead of labelling anything
- Expecting proposals before any failure is labelled
- Lowering the threshold to force early suggestions
- Seeding with fifty copies of one failure family