A ReportPortal launch's cluster list is useless -- either one cluster holding nearly every failure, or dozens each holding one. How do you tell which it is and what do you actually change?
answer
- the distribution of counts is the signal
- one huge count versus all ones
- read the message text, not just counts
- the flag only normalises digits
basics
~20 sRead matchedTests and message together across the list. One huge count under short generic text means unrelated problems merged; a page of counts of one means varying text split one problem, and removeNumbers helps only when that text is digits.
solid answer
~50 sBoth failure modes are visible from the list itself, because every cluster carries `matchedTests` and `message`. One cluster whose count is near the launch's whole failure total, under a short generic `message`, means the grouping was too broad -- a shared wrapper or prefix gave the comparison nothing to separate on. A list where almost every count is 1 means it was too narrow, and then the messages tell you why: if they differ only in digits, turn `removeNumbers` on in the generation request and regenerate. If they differ in a hostname suffix, a generated path segment or an identifier containing letters, the flag will change nothing, because it normalises digits only -- the fix then belongs upstream, in what the message says by the time it reaches the report. Confirm by regenerating the same finished launch and comparing count distributions, not ids.
go deeper
Know the two ways grouping by message goes wrong -- too broad merges unrelated failures, too narrow gives one cluster per failure -- and that member counts show which happened.
Explain how the normalisation flag maps onto the narrow case, and why it cannot help when the varying token in the message is not a number.
Show a diagnosis loop: read counts and messages, regenerate the same finished launch with the flag flipped, compare distributions rather than ids, and open the biggest cluster before trusting it.
Decide what a cluster list is allowed to drive automatically given both failure modes, and where normalisation should be enforced so the grouping has something to work with.
## Two ways a cluster list goes wrong Grouping failures by message similarity has exactly two failure modes, and they are opposites: - **Too broad.** One cluster swallows most of the launch. Distinct problems sit together because their messages share enough text -- a common wrapper exception, a shared assertion-library prefix, a generic timeout sentence -- that the comparison could not tell them apart. - **Too narrow.** Almost every cluster has one member. One problem has been split into as many clusters as it had failures, because each message carries something that varies run to run. Both produce a useless list, and both are visible without opening a single cluster. ## Diagnosing from the list alone Every `ClusterInfoResource` carries `matchedTests`, the count of test items in that cluster, and `message`, the text the group formed around. Read the two together: | what you see | what it means | what to do | |---|---|---| | one cluster, `matchedTests` near the launch's whole failure count, short generic `message` | too broad | look at the message -- if it is a wrapper or prefix, the text has no discriminating power | | many clusters, `matchedTests` almost all 1, messages differing only in digits | too narrow, numerically | turn `removeNumbers` on | | many clusters, `matchedTests` almost all 1, messages differing in non-numeric tokens | too narrow, but not fixable here | the varying token has to leave the message upstream | | a handful of clusters with plausible counts and readable messages | working | nothing | The third row is the one people miss. It looks identical to the second at a glance, and it is the reason "we turned the flag on and nothing changed" is such a common report. ## The lever you have, and its limit The generation request exposes one knob: `removeNumbers` on `CreateClustersRQ`, beside the required `launchId`. Turning it on strips digits before messages are compared; leaving it off compares them verbatim. The project's stored `analyzer.uniqueError.removeNumbers` supplies the default, and the request value overrides it for that one run without persisting. So your options, in order: 1. **Flip `removeNumbers` and regenerate.** Cheap, reversible, and it settles the numeric case in one experiment. 2. **Read the actual messages in the biggest cluster.** If they are genuinely different problems, no normalisation setting will separate them -- the text does not distinguish them and clustering has nothing else to work with. 3. **Change what the message says.** If the discriminating detail is missing from the message, or the noise is a non-numeric token, the fix is in what gets reported, not in the request. This is the honest answer when the first two do not help, and interviewers are listening for whether you reach it. ## Confirming the change rather than assuming it A full generation rebuilds the launch's clusters from scratch, so re-running on the **same finished launch** is a clean A/B: generate, read the list, flip the flag, generate again, read again. Compare the `matchedTests` distributions, not the cluster ids -- the ids are new either way, because regeneration replaces the rows. Two guards worth stating: - Generation is refused while the launch is still `IN_PROGRESS`, so run the experiment on a completed launch. - Judge the result against how many distinct problems you *believe* the launch had. A cluster count is not a quality score on its own; the question is whether the groups match the failures a human would separate. ## What you must not conclude A well-shaped cluster list is not a defect list. Clustering proposes groups from text; it does not decide that two failures are the same defect, does not explain what the members have in common, and does not rank them. Reading a tidy list of five clusters as five bugs is exactly the overreach that makes people distrust the feature the first time one cluster turns out to contain two.
- You flip removeNumbers, regenerate, and the singleton clusters remain. What have you learned?That the token splitting them is not numeric. The flag strips digits and nothing else, so a random hostname suffix, a generated file path or an alphanumeric identifier survives normalisation intact. The lever has been exhausted at this layer, and the varying detail has to be gone from the failure message before it is reported.
- How do you judge whether a cluster list is good, without a score to compare against?Against how many distinct problems you believe the launch actually had. Cluster count alone is not a quality measure -- five clusters is good if there were five problems and bad if there were two. Read the biggest cluster's members: if a human would separate them, the grouping is too broad regardless of how tidy the list looks.
- Why is re-running generation on the same launch a safe experiment?Because a full generation replaces that launch's clusters rather than adding to them: the existing rows are removed and rebuilt. Nothing accumulates, so you can generate under one setting, read the list, flip the flag and generate again. The launch must be finished, since generation is refused while it is still `IN_PROGRESS`.
saying these in an interview costs you the question
- Reading a tidy cluster list as a list of bugs
- Treating cluster count as a quality score
- Assuming removeNumbers can fix non-numeric noise
- Comparing cluster ids across two generations
- Concluding one big cluster means one defect