A ReportPortal cluster row holds `id`, `indexId`, `projectId`, `launchId` and `message`. Given that, how would you follow one recurring failure group across successive launches?
answer
- read the entity for what is missing
- no cross-launch reference of any kind
- regeneration replaces, it does not merge
- the message text is the only join
basics
~20 sNot through any field the model offers. A cluster row names one launch and holds no pointer to a cluster in another, so the only join available is the message text itself, compared outside the product.
solid answer
~40 sThe cluster entity is deliberately narrow -- `id`, `indexId`, `projectId`, `launchId`, `message` -- and what matters is what is absent: no reference to a cluster in a previous launch, no series key, no first-seen timestamp. `projectId` is addressing, not scope; the row still belongs to exactly one launch. Worse for anyone holding an id, a full generation for a launch **deletes** that launch's existing clusters and unlinks their logs before rebuilding, so `id` is not stable even within one launch, and there is no history of the previous grouping to diff against. The only join available is the `message` text, compared yourself, and only when both launches were generated under the same `removeNumbers` setting -- otherwise you are comparing normalised text against raw text.
go deeper
Recall that a cluster row names one launch and one message, and that nothing on it points at a cluster in a different launch.
Explain that a full generation deletes the launch's clusters and rebuilds them, so ids are not stable, and that message text is the only thing you can line up yourself.
Show the downstream consequences: no cluster-id foreign keys, no trends computed over ids, and an explicit statement of message equality as the join you chose.
Take a position on whether durable cross-launch grouping belongs in this layer at all, and what asserting it would commit the team to maintaining.
## A cluster row is launch-shaped The stored cluster entity in ReportPortal has five fields: `id`, `indexId`, `projectId`, `launchId` and `message`. Read that list for what is **missing**. There is no field naming a cluster in a previous launch, no series key, no stable slug derived from the message, no first-seen timestamp. Every cluster row belongs to exactly one launch, and `projectId` is addressing rather than scope -- it tells you which project's data this is, not that the cluster spans the project. On the API the same row appears as a `ClusterInfoResource` with `id`, `index`, `launchId`, `message`, `metadata` and `matchedTests`. `index` is the entity's `indexId` under a different name; neither of them is a cross-launch handle you were promised. So the honest answer to "can I follow this cluster into next week's run?" is: **not through any field the model gives you.** ## Regeneration is destructive It gets sharper than that. Running a full generation for a launch does not merge into what is already there. The pipeline first **removes that launch's existing cluster rows and unlinks the logs that pointed at them**, then builds the launch's clusters again from scratch. Two consequences follow: 1. A cluster's `id` is not a durable handle even *within* one launch. If anyone regenerates -- to try the other normalisation setting, say -- every id you recorded is stale. 2. There is no history of previous groupings. You cannot diff last generation's clusters against this one's, because last generation's rows no longer exist. An incremental update scoped to particular items behaves differently from a full launch generation, but the full-launch path is the one behind "cluster this launch", and it is the one whose behaviour surprises people. ## What you can actually join on If two launches were each clustered and you want to line their groups up, the only material you have is the **`message` text**, plus knowledge of whether the same normalisation was applied on both runs. That is a string comparison you perform yourself, outside the model, with all the usual caveats: - Compare messages produced under different `removeNumbers` settings and you are comparing normalised text against raw text -- they will not line up. - A message that carries any run-varying token will not match across launches even when the underlying failure is identical. - Nothing in the model tells you the comparison is legitimate; you are asserting it. ## What this means for anything you build on top - **Do not store a cluster id as a foreign key** in a dashboard, a spreadsheet or a bot. Store the launch and the message, and resolve the cluster when you need it. - **Do not present a cluster as a long-lived object** with a lifecycle. It has no lifecycle; it is regenerated output for one launch. - **Do not compute a trend over cluster ids.** Any "this cluster keeps growing week on week" claim built on ids is measuring row creation, not recurrence. - **Do treat "the same message appeared in five consecutive launches" as evidence** -- it is real evidence, you just have to assemble it yourself and say out loud that message equality is the join condition you chose. ## Why the model is like this It is a defensible design rather than an oversight. Clustering is a **cheap, disposable view over one launch's failure text**, offered so a screen of red becomes a short list. Making clusters durable would require deciding that two groups in two launches are the same thing -- and that decision, with its consequences for what counts as a recurring problem, is not something a text-similarity pass over one launch is entitled to make on its own. The model declines to make it, and leaves the row scoped to the launch that produced it.
- A dashboard stores cluster ids and charts how each cluster grew week over week. What is wrong with it?It is charting row creation, not recurrence. Ids are minted per launch and replaced whenever a launch's clusters are regenerated, so no id survives to be tracked. The chart should key on launch plus message text and state that message equality is the chosen join condition, since nothing in the model asserts it.
- Is the missing cross-launch link an oversight worth working around, or a deliberate design?Deliberate. Linking two groups across launches means asserting they are the same recurring problem, and a text-similarity pass over one launch has no standing to make that call. The model keeps clusters as a disposable per-launch view and leaves the durable judgement to whoever is entitled to make it.
- Two launches were clustered with different normalisation settings. Can you still compare their clusters by message?Not safely. One set of messages had digits stripped before grouping and the other did not, so equal underlying failures produce unequal text. Regenerate one launch under the other's setting first, which is cheap because a full generation rebuilds that launch's clusters from scratch anyway.
saying these in an interview costs you the question
- Storing a cluster id as a durable foreign key
- Assuming regeneration merges into existing clusters
- Expecting a previous generation to remain readable
- Believing projectId widens a cluster's scope