skip to content

Why report sliced eval scores by segment and failure mode instead of one aggregate?

level: seniorimportance: must knowfreq 50%

answer

  1. means compress away the cohort that matters
  2. two axes: who, and what broke
  3. small slices swing wildly per item
  4. protect the worst slice, not the mean
  5. many slices, some fail by chance

basics

~20 s

An aggregate averages cohorts together, so a large healthy segment hides a small broken one. Slicing by who the request came from and by what went wrong exposes concentrated regressions, and lets you gate on the worst slice rather than the mean.

solid answer

~50 s

Aggregates are a compression, and the thing they compress away is exactly what you need. On a prior-authorization assistant, a prompt change lifted the overall score three points while specialty-drug denials — four percent of items, and the cases with real patient consequence — dropped fourteen. The mean moved the right way, so nothing fired. I slice on two axes: a **who** axis (payer plan type, drug class, channel, language) and a **what-went-wrong** axis (hallucinated formulary rule, missing criterion, wrong plan rule, over-refusal). The report is a table of per-slice scores with per-slice item counts, and the gate is the worst slice, not the average. Two disciplines keep it honest: enforce a minimum item count per slice, because an eight-item slice swings twelve points on one changed answer, and confirm a suspicious slice before acting, since with thirty slices a few will look bad by chance.

go deeper

for a junior

Know that one overall eval score can hide a broken group, and that scores are normally reported per category as well as overall.

for a middle

Be ready to name the two slicing axes — properties of the input and types of failure — and to explain why each slice needs a minimum item count before its score means anything.

for a senior

Show the operational judgment: gate on protected slices rather than the mean, pair runs so you can inspect the items that flipped, and require a suspicious slice to reproduce before calling a regression.

for a principal

Own which cohorts are protected and why — that choice encodes where the organization refuses to trade edge-case quality for average gains — plus the labeling cost that per-slice resolution implies and how the taxonomy grows from incidents.

## The failure the aggregate hides A single number over an eval set answers one question — did overall quality move — and silently refuses several others: for whom, in what way, and at what cost. Because the aggregate is a weighted mean, a cohort's ability to influence it is proportional to its share of the set. A cohort that is four percent of items can collapse entirely and move the headline by well under a point, comfortably inside run-to-run noise. If that cohort is the one where errors cause harm, the aggregate is not merely uninformative, it is actively misleading: it reports a win on the change that broke your riskiest path. The distortion runs the other way too. A change that improves the head of the distribution and degrades the tail looks strictly good; a change that fixes a small, hard cohort at a tiny cost to the easy majority looks strictly bad. Optimizing against an aggregate therefore drives a system toward the average request and away from the edges — which is where trust is actually won or lost. ## Two slicing axes **Segment slices (who or what the input is)** partition the set by properties of the request that exist before the system runs: payer plan type, drug class, customer tier, document length, language, channel. These map to accountability — a broken segment is a named group of users you can go and tell. **Failure-mode slices (what went wrong)** partition the *errors*, not the inputs, using a taxonomy built from real incidents: hallucinated rule, missing required criterion, wrong policy applied, over-refusal of a valid request, malformed output. These map to remediation — each mode has a different fix, and a mode-level view tells you whether a change traded one failure for another rather than removing failures. The two axes answer different questions and you want both. Segment slices are computed on every item; failure-mode slices are computed on the failures, so their denominators differ and the report must say so. ## Designing the set so slicing is possible Slicing is a property of the dataset, not of the dashboard. Every item needs the slice labels attached at curation time — segment tags from the trace metadata, and a failure-mode tag assigned when the item is reviewed. Retrofitting tags onto an existing set is expensive and usually partial, which is why teams end up with only an aggregate. More importantly, slicing is what sets the minimum size of each cell. A slice you intend to gate on must have enough items that its score is stable across runs. A slice of eight items moves 12.5 points per changed answer; nobody can act on a metric with that granularity. This is the reason a stratified quota per cell is worth its labeling cost: it buys per-slice resolution the aggregate does not need but the decisions do. ## Reading a sliced report without fooling yourself Slicing multiplies the number of comparisons, and comparisons are where false alarms come from. With thirty slices, a handful will look meaningfully worse in any given run purely by chance, and the smallest slices are the ones most likely to do it. Three practical controls: - **Minimum slice size.** Below the threshold, display the slice as informational and never gate on it. - **Confirm before acting.** Re-run the suspicious slice — ideally with more items, or repeated runs of the same items — before declaring a regression. A slice that reproduces is a signal; one that does not was noise. - **Pair the comparison.** Run both versions over the same items and look at which items *flipped*, not just the two rates. A regression concentrated in a slice usually shows up as a handful of specific items that changed from pass to fail, and those items are directly readable. ## Gating on the worst slice Once slices exist, the natural policy is that a change must not regress any protected slice beyond a threshold, even if the mean improves. That is much stricter than a mean-based gate and it is the right default for cohorts with asymmetric cost — safety-relevant categories, regulated segments, your largest customers. For everything else, the mean plus a watch-list is enough. Deciding which slices are protected is a product judgment, not a statistical one: it encodes where you are unwilling to trade quality for average gains. ## When a slice appears Every incident that reaches production is evidence the taxonomy was incomplete. The healthy loop is: incident, new failure mode, new slice, items sampled or written for it, minimum size met, then it joins the gated report. Over time the slice list is a written record of everything the system has been caught doing wrong.

  • How do you decide which slices are allowed to block a release and which are only informational?
    By cost asymmetry, not by size. Slices where an error is expensive or irreversible — safety-relevant categories, regulated segments, cases with patient or financial consequence — get a hard threshold. Everything else feeds a watch-list. A slice also needs enough items for a stable score before it can gate anything; below that, it reports but does not block.
  • With thirty slices, some will look regressed in any run. How do you avoid chasing noise?
    Set a minimum item count per gating slice, pair the comparison so you inspect the specific items that flipped rather than two rates, and require a suspicious slice to reproduce on a re-run before you act. Treat a first-time slice alert as a request to investigate, not as a verdict.
  • Where do new failure-mode slices come from?
    From incidents and from reviewer notes during labeling. Every production failure that the existing taxonomy could not name is evidence the taxonomy is incomplete: define the mode, sample or write enough items to reach the minimum size, and add it to the gated report. The slice list becomes a written record of everything the system has been caught doing wrong.

saying these in an interview costs you the question

  • The overall score went up, so the change is good
  • Slice counts do not matter, percentages are comparable
  • Any slice that dips is a regression worth reverting
  • One aggregate number keeps the dashboard simple
  • Add slices later, the data can be tagged retroactively

context