skip to content

How would you govern score names and configs across teams in one Langfuse project?

level: principalimportance: should knowfreq 30%

answer

  1. the name is a join key
  2. one definition, not three codebases
  3. configs are the registry you already have
  4. source separates human from judge
  5. changed meaning means a new name

basics

~20 s

Treat the score name as a schema, not a label: it is the key every chart, filter and run comparison groups by. Register each name once as a score config with a fixed type and range, keep human, judge and code-written signals distinguishable, and make a changed meaning mean a new name.

solid answer

~60 s

A score name in Langfuse is a join key. Dashboards group by it, dataset runs compare by it, and alerts fire on it, so uncoordinated naming is not cosmetic — 'helpfulness', 'helpful' and 'helpfulness_v2' become three half-populated series that no one trusts. The governance that actually works is small. Register every score name as a **score config** with its data type and range or label set, so the definition lives in the project rather than in three codebases. Agree a naming convention and stick to it — lowercase, dimension-first, no team prefixes that fragment the same concept. Decide deliberately whether a dimension judged by a model and by a human share a name (they can, since `source` separates `EVAL` from `ANNOTATION` and `API`) or are named apart. And treat a change of meaning as a **new name**: historical scores keep their old values and never backfill, so silently retyping or rescaling a name corrupts every series that spans the change. Publish the small list, and review additions the way you review a database column.

go deeper

for a junior

Know that a score's name is what dashboards group by, so using an existing name consistently matters more than inventing a precise new one.

for a middle

Explain why score configs belong in the project rather than in code, and why two spellings of one dimension produce two useless half-series.

for a senior

Show that you plan for change: a redefined score gets a new name, environments and tags stay clean so scores can be sliced, and gated scores are kept few.

for a principal

Own the vocabulary as a schema with a review point, decide which signals are worth collecting against judge and reviewer cost, and separate what you gate on from what you merely watch.

## Why the name is the schema Everything Langfuse aggregates about quality is grouped by score name. A trend line is 'the average of scores named X over time'; a dataset-run comparison is 'run A's X against run B's X'; an alert is a threshold on X. The name is therefore the one identifier that has to mean the same thing to everyone writing it. It is also the identifier with the weakest natural defences, because writing a score requires nothing but a string, and any service, evaluator or reviewer can invent one. The failure is unglamorous and common. Two teams instrument 'the same' dimension six weeks apart, one calls it `helpfulness` on 0-1, the other `helpful_score` on 1-5. Both dashboards look fine in isolation. The organisation-wide view is nonsense, the cross-team comparison everyone wanted is impossible, and there is no migration that fixes it, because the old scores keep their old values forever. ## The governance that is worth having **A registry, not a wiki page.** Score configs already are the registry: each pins a name to a data type and to a range or label set, and the annotation UI renders from it. Requiring that every production score name has a config makes the definition a project object rather than tribal knowledge, and gives new names a natural review point. **A convention thin enough to follow.** Lowercase, underscore-separated, named for the dimension and not the team or the tool that produced it: `faithfulness`, not `platform_faithfulness_v2`. Team prefixes feel tidy and fragment exactly the comparisons that make one shared project worth having. **Deliberate handling of sources.** The same dimension can legitimately be written by code, by a judge and by a human. Langfuse's `source` field (`API`, `EVAL`, `ANNOTATION`) already separates them, so sharing one name is the right default and lets you compare them directly. Split into two names only when the two are measuring genuinely different things under one word — and then say so in the config descriptions. **New meaning, new name.** Rescaling a numeric score, retyping it, or quietly redefining what it judges breaks every aggregate spanning the change, and nothing warns you. Close the old series and open a new name. This is unpopular because it looks like clutter; it is far cheaper than a chart that has been silently wrong for a quarter. **Environment and slicing hygiene.** Scores carry the environment of the client that wrote them, and traces carry the tags, user ids and metadata you set. Governance of *those* matters as much as names, because a shared score name is only useful if you can also slice it by release, tenant and environment. Agreeing that staging traffic never lands in the production environment is a one-line rule that saves a lot of confusing dashboards. ## What not to over-engineer Resist a taxonomy nobody can recite. A dozen named dimensions that every team writes correctly beat forty that drift. Resist mandating that every service score everything: coverage costs judge money and reviewer hours, and a rarely-populated score name is worse than no name, because it invites conclusions from a handful of samples. And keep the gate small — the set of scores you actually block a release on should be short, explicitly chosen, and separate from the larger set you merely watch. ## Where this stops Governance keeps scores *comparable*; it does not make them *correct*. Whether a judge's number tracks reality, how many samples a difference needs before it means anything, and how to keep judgements stable as models change are questions of evaluation methodology that no naming convention answers. The value of getting the names right is that when you do answer those questions, the data you have collected is still usable.

  • Should a dimension judged by an LLM and by a human share one score name?
    Usually yes. Langfuse records source as EVAL for evaluator scores and ANNOTATION for human ones, so a shared name lets you chart both and compare them on identical traces while still being able to filter them apart. Split into separate names only when the human and the model are genuinely assessing different things, and record that difference in the configs.
  • A team wants to change a numeric score from a 0-1 scale to 1-5. What do you tell them?
    Use a new name. Existing scores keep their old values and never backfill, so any average or alert spanning the change silently mixes two scales. Open a new configured name, let the old series end, and keep both for the overlap period if the trend matters. The alternative is a chart that is wrong in a way nobody will notice for months.
  • How many scores should a release actually be gated on?
    Few — typically one or two that map to real user harm, with the rest watched rather than enforced. Every gating score is a source of false blocks, and judge-written scores carry non-determinism on top. Keep the gated set short, explicitly chosen and reviewed, and let the broader vocabulary serve investigation and trend-watching instead.

saying these in an interview costs you the question

  • Treats score names as free-form labels
  • Lets each team prefix its own copy of a dimension
  • Rescales an existing score name in place
  • Assumes historical scores can be retyped or backfilled
  • Mandates dozens of scores without costing judge or reviewer time

context