What do Langfuse's NUMERIC, CATEGORICAL and BOOLEAN score data types change?
answer
- three types, not free-form values
- string label versus number versus pass/fail
- what does averaging even mean here
- configs pin range and label set
- annotation UI renders from the config
basics
~20 sThe data_type decides what a Langfuse score value may be and how it aggregates: NUMERIC takes a number and averages, CATEGORICAL takes a string label and counts per label, BOOLEAN takes a pass/fail recorded numerically so it charts as a rate. A score config pins the type and allowed values for a score name.
solid answer
~50 sLangfuse scores carry a `data_type` of NUMERIC, BOOLEAN or CATEGORICAL. NUMERIC takes a float and is charted as an average or distribution, so it suits graded judgements like a 0-1 faithfulness score. CATEGORICAL takes a **string** value such as 'positive' or 'off_topic' and is charted as counts per label — use it when the labels are not ordered and averaging them would be nonsense. BOOLEAN is the pass/fail case, stored numerically so it still aggregates as a pass rate. On top of the type there are **score configs**, defined per project: a config fixes a score name's data type and its allowed range or category list, and `config_id` binds a written score to it. Configs are what make the annotation UI render the right control and what stops one name from carrying a 0-1 float in one place and a 1-5 rating in another, which would quietly corrupt every chart built on that name.
go deeper
Know the three data types by name and that the value must match: a number for NUMERIC, a string label for CATEGORICAL, pass or fail for BOOLEAN.
Explain how each type aggregates and pick the right one for a given judgement, and say what a score config pins down beyond the per-call data_type.
Show that you think about the analysis before the write: scale documented in the config, bounded inputs for annotators, and no reuse of a name whose meaning changed.
Own the score vocabulary as a schema — which names exist, what type and range each carries, and the rule that a changed meaning means a new name rather than a silent retype.
## The three data types Every Langfuse score has a `data_type`, and the API enum has exactly three members: `NUMERIC`, `BOOLEAN`, `CATEGORICAL`. The type is not decoration — it decides what the `value` may hold and how the platform aggregates the score. **NUMERIC** is the default shape: `value` is a number, and Langfuse charts averages and distributions over it. Judge outputs on a 0-1 scale, latency-derived quality proxies and star ratings all live here. If a number's scale is not obvious from the name, put the scale in the score config rather than in a comment. **CATEGORICAL** takes a string `value` such as 'refused', 'partially_answered' or 'complete'. Langfuse aggregates it as a count per label. The reason to reach for it is that the labels have no meaningful arithmetic: the mean of 'refused' and 'complete' is not a thing. Encoding unordered outcomes as numbers to 'keep it simple' is the classic mistake, because the resulting average moves for reasons nobody can interpret. **BOOLEAN** is the pass/fail case — did the guardrail fire, did the answer contain a citation, did the tool call succeed. It is recorded numerically so that the aggregate reads as a pass rate, which is exactly the number you want on a dashboard. Modelling the same thing as a two-value CATEGORICAL score also works, but you lose the free rate chart. ## Score configs A data type alone does not stop two teams from writing incompatible values under one name. That is what **score configs** are for. A config is defined once in the project and pins, for a given score name, the data type and the allowed values: a minimum and maximum for numeric scores, or the permitted label set for categorical ones. When you write a score from the SDK you can pass `config_id` to bind it to that definition. Configs earn their keep in two places. In the annotation UI they decide what a human reviewer is shown — a slider bounded to the configured range, or a fixed list of labels rather than a free-text box, so two reviewers cannot invent two vocabularies. And in analysis they guarantee that everything filed under one name is on one scale, which is the precondition for any comparison across time, releases or dataset runs. ## Choosing a shape A useful order of questions: is the judgement binary? Use BOOLEAN. Is it one of a small set of named outcomes with no order? Use CATEGORICAL. Is it a graded quantity on a scale you can state? Use NUMERIC, and write the scale into the config. Where a categorical label and a number both fit — a 1-5 rating, say — prefer NUMERIC so the rating averages, and add a config with min 1 and max 5 so nobody writes a 7. A second, softer rule: keep the `comment` field for the reasoning behind a value, not for a second value. Judge evaluators fill it with their explanation, which is what makes a low score debuggable. Cramming structured data into the comment defeats aggregation just as thoroughly as picking the wrong type. ## What goes wrong The damage from a wrong type is retroactive and quiet. Scores already written keep their original values, so if `helpfulness` was numeric for three months and someone starts writing categorical labels under the same name, the historical chart does not break loudly — it splits, and every average computed over the boundary is meaningless. Fixing it means choosing a new name and accepting that the old series ends. Deciding the type and the config *before* a score name goes into production is much cheaper than migrating it afterwards.
- Why not encode 'refused', 'partial' and 'complete' as 0, 1 and 2 in a NUMERIC score?Because the average of an unordered set is uninterpretable. A mean of 1.4 tells you nothing about which outcomes moved, and the number shifts when you add a fourth label. CATEGORICAL keeps the counts per label, which is the shape the question actually has, and Langfuse charts it that way.
- What does a score config give you that passing data_type on every call does not?A single project-level definition of a score name: its type plus the allowed numeric range or label set. That constrains what human annotators can enter, renders the right control in the annotation UI, and prevents two teams from writing a 0-1 float and a 1-5 rating under one name — which no per-call argument can catch.
- A score name has been numeric for months and someone starts writing labels under it. What now?Old scores keep their old values, so the series silently splits and any aggregate spanning the change is meaningless. The clean fix is a new score name with a proper config, leaving the old series closed. Retro-typing existing scores is not something you should plan around.
saying these in an interview costs you the question
- Says every score has to be a number
- Encodes unordered labels as 0, 1, 2 and averages them
- Thinks data_type is only a display hint
- Treats score configs as optional decoration
- Reuses one score name for two different scales