skip to content

In ReportPortal, a launch's generated error clusters are returned as `ClusterInfoResource` objects. What do the `index`, `message`, `matchedTests` and `launchId` fields on one of them tell you?

level: juniorimportance: must knowfreq 58%

answer

  1. one group of similar failures
  2. text, count, identifier, owner
  3. matchedTests is a member count
  4. the scope field names a launch

basics

~20 s

message is the shared failure text the group formed around, matchedTests counts the test items in it, index is the identifier the analyzer gave the group, and launchId scopes the cluster to a single launch.

solid answer

~40 s

ReportPortal's unique-error clustering groups a launch's failing test items by how similar their failure text is, and each group comes back as a `ClusterInfoResource`. Its `message` holds the representative failure text the group formed around -- that text is the whole basis of the grouping. `matchedTests` is the count of test items that landed in the cluster, so it is the first thing to read when checking whether one cluster swallowed the launch. `index` is the identifier the analyzer assigned the group, stored on the entity as `indexId`, while `id` is the cluster row's own database key -- two different numbers. `launchId` is the field people skip: a cluster belongs to exactly one launch, so the same failure family in tomorrow's run is a different cluster row.

go deeper

for a junior

Be able to name the four fields and say what each holds, and state plainly that a cluster belongs to one launch. Knowing matchedTests is a count, not a list, is enough at this level.

for a middle

Explain that the grouping is a text operation over failure messages, that generation is asynchronous and read back separately, and that the resource's index is the entity's indexId under another name.

for a senior

Show you read the matchedTests distribution across the list as a health signal before opening anything, and that you never store a cluster id as a durable key in something downstream.

for a principal

Own the position that clustering is a disposable per-launch view, not a defect registry, and be able to say what you would and would not let a cluster drive automatically.

## What clustering produces ReportPortal's unique-error clustering takes the failing test items of **one launch** and groups them by how similar their failure text is, with nobody having written a rule first. Each group that comes out is one row in the clusters table, and the API renders that row as a `ClusterInfoResource`. Two calls bracket the feature. One **starts** generation for a launch and returns an acknowledgement that generation has started -- the work runs asynchronously in the analyzer service. A second call **reads back** the clusters that now exist for that launch. An empty list a second after the first call means "not finished yet", not "no clusters found". Generation is also refused outright for a launch that is still `IN_PROGRESS`, so the sequence is: finish the launch, start generation, then read. ## The fields, one at a time | field | what it holds | |---|---| | `id` | the cluster row's own database key | | `index` | the identifier the analyzer gave this group; the stored entity holds the same value as `indexId` | | `launchId` | the launch this cluster belongs to | | `message` | the representative failure text the group formed around | | `matchedTests` | how many test items fell into the group | | `metadata` | a free-form map carried alongside the cluster | - **`message` is the entire basis of the grouping.** Clustering here is a text operation over failure messages. Everything a cluster claims is a claim about strings, which is why a message that embeds a request id or a timestamp behaves so differently from one that does not. - **`matchedTests` is a count, not a list.** It is computed as the number of test items linked to that cluster. To see *which* items, you follow the cluster; to judge whether the grouping is any good, the count alone is usually enough. - **`index` and `id` are two different numbers.** `id` is the row's primary key inside ReportPortal. `index` is the identifier the analyzer assigned the group, and the persisted entity calls that field `indexId` -- the API just renames it on the way out. Quoting one where the other was meant is the classic confusion when someone scripts against this endpoint. - **`launchId` is the scope, and it is the field people skip over.** A cluster is a per-launch object. The entity also carries a `projectId`, but that is addressing, not scope: the row still names exactly one launch. ## Reading a cluster list Because `matchedTests` sits on every row, the *shape* of the list is diagnostic before you open a single cluster: 1. One cluster whose `matchedTests` is close to the launch's entire failure count, sitting under a short, generic `message` -- the grouping merged problems that are not the same problem. 2. A long list in which nearly every `matchedTests` is 1 -- the grouping split one problem into many, usually because the messages carry text that varies run to run. 3. A handful of clusters with plausible counts and messages you can read as distinct failures -- the shape you actually want. That reading is the everyday use of the endpoint. You are not looking for a number; you are looking at a distribution. ## What a cluster is not A cluster is a **proposal about text**, not a verdict about defects. Nothing in `ClusterInfoResource` says "these are the same bug", assigns a defect type, or links a ticket. Those are separate acts performed by people and by other parts of the product, on top of a grouping that clustering merely offered. Treating a cluster as a defect record is the single most common overreach, and it shows up as soon as two genuinely different failures land in one cluster because they share a wrapper message. Equally, a cluster is not a rule. Nobody wrote a pattern that produced it, so nobody can point at the pattern to explain why two failures are together -- the only explanation available is the `message` the cluster formed around, plus whatever normalisation was applied before the messages were compared. ## Entity versus resource, and why it matters The stored entity is deliberately narrow: `id`, `indexId`, `projectId`, `launchId`, `message`. It holds **no member list and no count** -- membership lives in the link between items and clusters, and `matchedTests` is derived when the list is read. That is why the count you see is always current for the launch's present state, and why the entity is cheap to rewrite when clustering is run again. So when you read a cluster in an integration, hold on to `launchId` plus `message`, treat `matchedTests` as the health signal, and treat `id` as a handle valid only for as long as this launch's clusters are not regenerated.

  • You start cluster generation for a launch and immediately read the cluster list, which comes back empty. What has happened?
    The start call only acknowledges that generation has begun -- the analyzer does the work asynchronously. An empty list read straight afterwards means generation has not finished, not that nothing grouped. Poll the list, and check that the launch was not still `IN_PROGRESS`, since generation is refused for a launch that has not finished.
  • Why do the stored cluster entity and the API resource use different names for the same value?
    The entity persists the analyzer's group identifier in a field called `indexId`, and the converter copies it onto the resource's `index`. It is a rename on the way out, nothing more. The confusion it causes is real, though: `index` and `id` on the resource are different numbers, and only `id` is the row's database key.
  • The resource carries matchedTests but the stored entity has no count field. Why?
    The entity holds only `id`, `indexId`, `projectId`, `launchId` and `message`; membership lives in the link between test items and the cluster. `matchedTests` is counted when the list is read, so it always reflects the launch's current state and the entity stays cheap to delete and rebuild when clusters are generated again.

saying these in an interview costs you the question

  • Calling a cluster a defect record or a bug
  • Treating index and id as the same number
  • Assuming a cluster spans the whole project
  • Expecting matchedTests to list the member items