skip to content

Evaluation & Testing

Measuring output quality when there is no single right answer: reference metrics like BLEU and ROUGE, an LLM as judge, RAG-specific scores such as faithfulness, plus regression sets and A/B tests over prompt versions. This is the area that separates shipping an LLM feature from guessing at one.

on this pageshow

questions

25

What is the difference between a public LLM benchmark and a task eval?

level: juniorimportance: must knowfreq 68%

answer

  1. ranks models vs measures your system
  2. someone else's construct, not yours
  3. whole pipeline, not the bare model
  4. transfer is empirical, never assumed
  5. prefilter for a shortlist, then decide

basics

~20 s

A public benchmark is a shared, fixed dataset that ranks models against each other on a general capability. A task eval runs your own inputs through your own system and scores your own success criterion. Only the task eval predicts what your users will see.

solid answer

~50 s

A public benchmark such as MMLU-Pro or GPQA-Diamond is a fixed, shared dataset with a shared scoring rule. Its job is to compare **models** on a general capability, so everyone can read one number and rank the field. A task eval is a set of examples drawn from **your** problem — your inputs, your expected outcome, your definition of success — run through your **whole system**, meaning the prompt, retrieval, tools and post-processing, not the bare model. The two answer different questions. The benchmark answers "which model is generally stronger?"; the task eval answers "does my feature work well enough to ship?". Benchmarks are useful as a cheap prefilter to build a shortlist, but transfer from a leaderboard win to your product is an empirical claim you have to test, not an inference you get for free.

go deeper

for a junior

Be able to say plainly that a benchmark compares models on a shared public dataset while a task eval measures your own system on your own examples, and that only the second tells you whether the feature works.

for a middle

Explain why the score does not transfer: different construct, different input distribution, different definition of success. Note that the task eval tests prompt, retrieval and tools together, not the model in isolation.

for a senior

Show the funnel in practice — benchmarks build the shortlist, the internal suite gates the release — and describe how you would construct enough of a suite to make a real ship decision under time pressure.

for a principal

Own the argument that the internal suite is a durable asset that outlives any model version, and that funding it is what stops model choice being re-litigated from headlines every quarter.

## Two artifacts, two questions A **public benchmark** is a dataset plus a scoring rule, published so that anyone can run any model against it and get a comparable number. Examples in current use include MMLU-Pro (multiple-choice knowledge and reasoning across many subjects), GPQA-Diamond (graduate-level science questions written to be hard to look up), ARC-AGI-2 (abstract visual-pattern puzzles), SWE-bench Verified (fixing real issues in real code repositories), Terminal-Bench (completing tasks in a command-line sandbox), and GDPval (deliverables for real occupational tasks, judged by expert humans against human-produced work). LMArena is a different shape: it ranks models from pairwise human preference votes on freely chosen prompts. A **task eval** is the private mirror image. You collect examples of the actual job your product does — real user questions, real documents, real intended outcomes — write down what counts as a correct response, and run your candidate configurations against it. Crucially, the unit under test is the **system**, not the model: the same model behind two different prompts or two different retrieval setups will score differently, and it is the system your users meet. ## Why a benchmark number does not transfer Three independent reasons, and an interviewer will want at least one named. **Different construct.** The benchmark may measure a capability adjacent to, but not the same as, the one you need. A model that fixes Python bugs well is not thereby good at refusing to answer when your documents do not contain the answer. The formal name for "does this measurement capture the thing I care about?" is *construct validity*, and for most products the honest answer about a public benchmark is: only loosely. **Different distribution.** Benchmark items are curated, well-formed and usually short. Your traffic is messy, domain-specific, sometimes hostile, and often depends on documents the model has never seen. Retrieval quality, prompt structure and tool wiring dominate outcomes at exactly the point where benchmarks hold all of that constant. **Different success criterion.** Most benchmarks score exact-match correctness against a reference. Real products care about a bundle: correctness, but also groundedness (did it cite something real?), safe abstention, tone, latency, and cost per request. None of those appear on a leaderboard. ## What a task eval actually contains At minimum: a set of representative inputs; for each, either a reference answer or a checkable property ("cites the section that governs this case", "does not invent a clause number", "escalates instead of guessing"); and a scoring procedure you can re-run. The scale is usually much smaller than a public benchmark — dozens to a few hundred items — because every item has to be written or verified by someone who knows the domain. That is the trade: far fewer items, far higher relevance. The payoff is discrimination on the axis you care about. Two models that public leaderboards rank within a point of each other can separate cleanly on a small internal suite, because that suite asks about the one narrow capability your feature depends on. The reverse also happens: the model that wins the leaderboard loses your suite, and it is the suite you should believe. ## How the two work together The practical arrangement is a funnel. Public benchmarks and community rankings tell you which handful of models are even worth the integration effort — they are cheap, already computed, and updated as the frontier moves. Your internal suite then decides among that shortlist, and it is the artifact that gates the release. Benchmarks are a *prefilter with a shelf life*; the task eval is the decision record. A second, subtler benefit: the internal suite outlives every model. Providers deprecate versions, prices change, a cheaper model becomes good enough. Each of those events is a re-run of the same suite rather than a fresh argument. Teams that never build one end up re-litigating model choice from vibes and leaderboard headlines every quarter. ## What interviewers listen for The weak answer is "benchmarks are standard tests and task evals are custom tests" — true and useless. The strong answer names the asymmetry: a benchmark ranks models on someone else's construct, a task eval measures your system on yours, and the link between them is an empirical question. Candidates who have actually shipped an LLM feature reach for this immediately, because they have all been burned once by a model that looked better on paper and was worse in the product.

  • Your internal suite and the public leaderboard disagree about two models — which do you act on?
    The internal suite, provided you trust its construction. It measures your inputs, your pipeline and your success criterion, which is the thing you are shipping. The disagreement is still information though: it usually means the leaderboard is measuring a capability your product does not lean on, or your suite is too narrow to see a real difference. Say which you think it is rather than just declaring a winner.
  • If task evals decide, is there any reason to keep reading public benchmarks at all?
    Yes. They are free, they cover capabilities you have not tested, and they show where the frontier is moving — a large jump on an agentic or occupational benchmark is a signal to re-run your own suite. Use them to build the shortlist and to notice when a new class of capability has arrived. Just never let them close the decision.
  • Why should a task eval run the whole system rather than the model alone?
    Because users meet the system. Prompt wording, retrieval quality, tool definitions and post-processing routinely move outcomes more than the model swap does, and a model-only comparison attributes those effects to the wrong component. Testing end to end also catches integration failures — malformed tool arguments, truncated context, broken parsing — that no model-level score would ever reveal.

saying these in an interview costs you the question

  • Treats the top-ranked leaderboard model as automatically the right choice
  • Assumes a benchmark win transfers to any downstream task
  • Thinks a task eval just means running a public benchmark yourself
  • Believes leaderboards also capture cost, latency and reliability
  • Cannot state a success criterion for their own feature

context

open as a page

What is LLM-as-judge evaluation, and why is a judge score not ground truth?

level: juniorimportance: must knowfreq 70%

basics

~20 s

LLM-as-judge means prompting a model with a rubric to score another model's output. The score is a noisy estimate, not truth: the judge has its own biases and blind spots, so it must be validated against human labels before anyone trusts the number.

open as a page

What is the difference between offline and online evaluation of an LLM feature?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Offline evaluation scores a fixed, frozen set of saved examples in a harness before release, so it is repeatable and cheap. Online evaluation measures real user traffic after release, where actual behaviour and business outcomes decide whether the change helped.

open as a page

How do you sample production traces into a golden eval set without losing rare cases?

level: middleimportance: must knowfreq 60%

basics

~20 s

Sample by strata, not uniformly. Bucket traces by the dimensions you care about — request type, customer segment, known failure mode — then take a quota from each bucket, so rare but costly cases appear in numbers large enough to score.

open as a page

Which biases distort LLM-as-judge scores, and how do you control each one?

level: middleimportance: must knowfreq 68%

basics

~20 s

The recurring four are position bias (order of presentation sways pairwise verdicts), verbosity bias (longer wins), self-preference (a judge rates its own family higher), and formatting halo (bullets and confident tone read as quality). Controls: swap orders, normalise or cap length, judge across families, and strip presentation from the rubric.

open as a page

Why doesn't temperature 0 make an LLM eval suite reproducible in CI?

level: middleimportance: must knowfreq 70%

basics

~20 s

Temperature 0 only removes sampling randomness. Serving-side variance (batched floating-point arithmetic, expert routing), silently repointed model aliases, and unpinned fixtures all still move scores, so a suite needs a full pin set and repeated runs rather than one greedy pass.

open as a page

A vendor memo ranks models for a permit-review assistant using SWE-bench Verified and Terminal-Bench — what is wrong, and what would you measure instead?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Those benchmarks score agentic coding and command-line work, and a construction-permit assistant writes no code — the memo ranks models on a construct the product does not use. Replace it with a suite of real permit questions scored on citation correctness and safe abstention.

open as a page

Why report sliced eval scores by segment and failure mode instead of one aggregate?

level: seniorimportance: must knowfreq 50%

basics

~20 s

An aggregate averages cohorts together, so a large healthy segment hides a small broken one. Slicing by who the request came from and by what went wrong exposes concentrated regressions, and lets you gate on the worst slice rather than the mean.

open as a page

How do you validate an LLM judge against human labels before trusting its scores?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Have humans label a sample against the same rubric, run the judge on that sample, and measure chance-corrected agreement such as Cohen's kappa rather than raw percent agreement. Below roughly moderate agreement the judge is not usable, and the usual cause is an ambiguous rubric, not a weak model.

open as a page

Why do offline eval wins often fail to reproduce online, as when a grocery search-query rewriter gains 9 points offline but loses 0.4% cart conversion?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Three causes dominate: the frozen set no longer matches live traffic, the offline metric proxies something users do not reward, and the harness has no user in it to react. A large offline gain on the wrong population or the wrong proxy buys nothing online.

open as a page

How do you set an eval gate threshold that catches regressions without failing on noise?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Measure the suite's run-to-run spread on unchanged code first, then set the fail margin above that band — for example baseline minus two points — and decide on repeated runs with a majority rule so a single unlucky sample cannot block a release.

open as a page

How do you gate an LLM prompt change in CI when outputs are not exact-match?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Run a fixed set of scored eval cases as a build step and compare the aggregate score against the last accepted baseline. The build fails on a score drop beyond an agreed margin, not on any difference in wording.

open as a page

Why do MMLU, GSM8K and HumanEval no longer separate frontier models?

level: middleimportance: should knowfreq 54%

basics

~20 s

Saturation. Frontier models score roughly 88-99% on all three, so the surviving gap is mostly ambiguous items and label errors rather than capability. Once scores sit at a benchmark's ceiling it stops discriminating, which is why harder successors replaced it.

open as a page

Two reviewers disagree on 31 of 400 eval labels — what is your adjudication protocol?

level: middleimportance: should knowfreq 42%

basics

~20 s

Route the disputed items to a third, independent adjudicator, then read the resolved cases together. Most disagreement is a symptom of an underspecified rubric, so the real output is a sharper rubric plus a re-label of the affected category — not just 31 settled labels.

open as a page

When would you use pairwise judging instead of a pointwise rubric score?

level: middleimportance: should knowfreq 50%

basics

~20 s

Use pairwise when you need to resolve which of two candidates is better, because models compare far more reliably than they assign absolute numbers. Use pointwise when you need a per-item score that is comparable over time, across releases and by criterion.

open as a page

In an LLM rollout, what can shadow traffic measure and what can it never measure?

level: middleimportance: should knowfreq 47%

basics

~20 s

Shadow traffic mirrors real requests to a candidate whose output is scored but never shown, so it measures behaviour on the true request distribution at real latency and cost with zero user risk. It can never measure user response, because no user ever sees the output.

open as a page

What is benchmark contamination, and how would you detect it in a reported score?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Contamination is test data leaking into a model's training corpus, so the score measures recall of seen items instead of the capability. Detect it by comparing performance on freshly authored or perturbed items of equal difficulty: a large drop is the signature.

open as a page

Your 60-item eval suite shows a 4-point win — is that difference real?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Almost certainly not. On 60 pass/fail items scoring around 80 percent, the 95 percent interval is roughly plus or minus 10 points, so a 4-point gap is well inside noise. Detecting it needs pairing, graded scores, or several hundred items.

open as a page

When should a guardrail metric stop a rollout whose primary metric is up?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Whenever a pre-declared harm metric breaches its threshold, even if the primary metric improves. Guardrails encode costs the primary metric ignores — refunds, escalations, safety violations, latency, spend — and they are set before the experiment precisely so a good headline number cannot argue them away.

open as a page

How do you triage eval cases that keep flapping in CI at temperature 0?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Diagnose the flap first — an ambiguous expected answer, a judge scoring near a rubric boundary, or an unpinned fixture. Fix what is fixable, move the irreducible remainder into an advisory tier that reports but cannot block, and delete cases that stabilise nothing.

open as a page

How much weight should an engineering org give vendor-reported benchmark scores?

level: principalimportance: should knowfreq 31%

basics

~20 s

Enough to build a shortlist, never enough to close a decision. Reported scores come from the party selling the model, often under undisclosed scaffolding and effort settings, and they omit cost, latency and reliability. Internal suites hold the deciding vote.

open as a page

How often should a golden eval set be refreshed, and what should trigger a refresh?

level: principalimportance: should knowfreq 28%

basics

~20 s

On a scheduled cadence plus event triggers. Schedule a review each quarter; trigger immediately when the rules the labels encode change, when the product gains a capability the set never covers, or when the traffic mix shifts. Version the set, never edit it silently.

open as a page

How do you keep LLM-judge scores comparable across judge upgrades and months?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat the judge as a versioned instrument: pin the judge model version and rubric text, record both with every score, and never compare numbers across a change without a bridge. When either changes, re-run a frozen human-labelled set and dual-run old and new judges over an overlap to measure the offset.

open as a page

How do canary rollouts and A/B tests differ in what they control for?

level: principalimportance: should knowfreq 33%

basics

~20 s

A canary controls risk: a tiny exposed slice bounds blast radius while you watch for breakage, and it is read as a safety check. An A/B test controls inference: randomized arms and sufficient sample size let you attribute a measured effect to the change.

open as a page

How do you split an LLM eval suite between a per-PR gate and a nightly run on a fixed budget?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat cost and wall-clock as gate design constraints. Put a small stratified subset covering every failure mode on pull requests so it finishes in minutes, run the full suite nightly and before release, and measure how many nightly regressions the subset would have caught.

open as a page