For a ReportPortal project, would you let auto-analysis write the defect group straight onto incoming failures, or keep it as a proposal a person accepts — and how do you stop wrong machine labels feeding the next round?
answer
- who is accountable for the label
- an applied label becomes evidence
- a wrong one compounds silently
- per-issue exclusion beats a project switch
- re-run only the machine-set population
basics
~20 sKeep it as a proposal a person accepts, and keep the record of who set every label. Applied groups become the evidence for the next round, so a wrong label compounds unless you can find the machine-set ones and reset them.
solid answer
~40 sThe two postures differ in who is accountable. With `isAutoAnalyzerEnabled` on, analysis writes the resolved defect group onto the issue and sets `autoAnalyzed`; with the suggestion surface, `searchSuggests(...)` proposes, a person decides, and `handleSuggestChoice(...)` records the choice. I default to proposing, because an applied label becomes indexed evidence for the next similar failure — that is the feedback trap, and it is why the machine's output is a proposal and not a verdict: it knows this failure resembles one you judged, never why it failed. Three controls make either posture survivable: `ignoreAnalyzer` on an issue keeps a known-bad family out of analysis; `AnalyzeItemsMode` `AUTO_ANALYZED` re-runs exactly the machine-set labels without touching hand-made ones; and raising `minShouldMatch` or narrowing `analyzerMode` tightens what qualifies as a match.
go deeper
Understand that a machine-proposed defect group is a suggestion drawn from past decisions, and that somebody still has to agree with it before it means anything.
Explain the loop: an applied label is indexed and becomes evidence for the next similar failure, so mistakes spread rather than staying where they were made.
Show the recovery path — per-issue exclusion for a bad family, a re-analysis run aimed only at machine-set labels, and a similarity bar you would rather raise than lower.
Own the posture and the conditions for changing it: where auto-apply is permitted, who audits provenance, and what evidence would move a failure family from proposal to automatic.
## The two postures There are only two, and the difference is where accountability sits. **Auto-apply.** `isAutoAnalyzerEnabled` is on, analysis runs over the undecided failures, and a match writes the resolved defect group straight onto the issue with `autoAnalyzed` set. Nobody is asked. The queue drains itself. **Propose and accept.** The suggestion surface returns candidates through `searchSuggests(...)`, each wrapped with the past failure it came from and that failure's error logs, and a person decides. `handleSuggestChoice(...)` sends the decision back. Auto-apply is defensible for the narrow, boring, high-volume families — the environment reset everybody recognises, the known infrastructure timeout. It is indefensible as a default for a whole project, for one reason that has nothing to do with model quality. ## Why a proposal is never a verdict The machine's claim is precise and modest: **this failure resembles one that somebody already judged.** It does not know why the test failed. It has no access to the change under test, the environment, or the intent of the assertion. Resemblance of failure text to a previously-labelled failure is genuinely useful evidence, and it is not causation. So the honest reading of a proposal is "here is a decision you made before, about something that looks like this". That is worth a lot when the reviewer can see the earlier failure — which is exactly why the surface returns the past item and its logs beside the proposal, rather than a bare group. ## The feedback trap Here is the mechanism that makes carelessness expensive: 1. A failure is labelled — by a person or by analysis. 2. That labelled failure is indexed and becomes a candidate for future matches. 3. The next similar failure matches it and inherits the same group. 4. Repeat. A correct label is an asset that keeps paying. A wrong label is a liability that keeps spending, and it spends silently: the queue looks healthy, the counters look plausible, and the same wrong group spreads across an entire failure family. Nothing errors. Auto-apply shortens the loop from four steps to three by removing the only human in it. That is why the posture question is really a question about how fast you want your mistakes to compound. ## The controls, and what each is for - **`ignoreAnalyzer`** — a boolean on an individual issue. Analysis skips those failures when collecting candidates, and the flag travels with the record so they stop acting as evidence. This is the surgical instrument: use it on a known-bad family you never want proposed again. It is per-issue, not a project switch. - **`AnalyzeItemsMode` `AUTO_ANALYZED`** — a re-analysis run aimed at exactly the machine-labelled items. Those items are pulled out of the analyzer's index and their issues reset before the run, so a bad batch is undone rather than argued with. `MANUALLY_ANALYZED` does the same for hand-set labels, `TO_INVESTIGATE` for undecided ones, and a request may ask for those three. - **`autoAnalyzed`** — not a control but the precondition for all of them. It is the provenance bit that makes the two populations separable. Without it, "redo the machine's work and leave ours alone" is not an expressible instruction. - **`minShouldMatch` and `analyzerMode`** — the blunt instruments. Raising the similarity bar or narrowing the slice of history in scope makes matches rarer and better. Reach for them after the surgical options, not before. ## What I would actually decide | decision | position | |---|---| | default posture | propose and accept | | auto-apply | only for named, high-volume families with a measured accept rate | | provenance | mandatory, and audited — including on clients that update issues | | quality metric | share of proposals accepted, not size of the queue | | review cadence | sample machine-set labels regularly, not only when someone complains | The provenance row deserves emphasis. When a person edits an issue, the value written for `autoAnalyzed` comes from the request body, so a client that echoes back what it read can re-assert the flag on a hand-made label. Left unchecked, the two populations blur and every recovery tool above quietly stops meaning what it says. ## What would change my mind Be concrete about this, because the interviewer is testing whether you hold the position dogmatically: - A sustained, measured accept rate on a specific failure family, with the family narrowly identified. - A team that reviews machine-set labels on a schedule rather than never. - A reset path that has actually been exercised, so "we can undo it" is an observation and not a hope. Absent those, auto-apply is not automation. It is a way of making the undecided queue look smaller than the amount of undecided work.
- One known-bad failure family must never teach the analyzer again. What do you set?`ignoreAnalyzer` on those issues. Analysis skips them when collecting candidates and the flag travels with the record, so they stop acting as evidence for anything else. It is a per-issue exclusion rather than a project-wide switch, which is what you want — the rest of the project keeps working normally.
- Which single number tells you the suggestion surface is earning its place?The share of proposals a reviewer accepts. A high accept rate means the matcher is finding the right past decisions; a falling one with a stable queue means it is proposing noise people are now filtering by hand. Queue size cannot tell you this, because a queue empties just as fast from lowering the bar.
- How would you undo a week of bad machine labels without losing the team's own triage?Re-analyse with `AnalyzeItemsMode` `AUTO_ANALYZED`, which collects exactly the machine-labelled items, removes them from the index and resets their issues before running again. Hand-set labels are a separate population and stay untouched — which only works if the provenance flag was trustworthy in the first place.
saying these in an interview costs you the question
- Letting the machine label without recording who set it
- Treating a machine label as a triage decision
- Re-running analysis over hand-made labels too
- Judging the analyzer by how empty the queue looks
- Discarding declined proposals as wasted effort