skip to content

Offline vs Online Evaluation

You learn the two halves of an eval program: offline runs against a frozen dataset before you ship, and online measurement of real traffic after. Each answers questions the other cannot, and treating an offline win as a production win is one of the most common mistakes in AI engineering.

on this pageshow

questions

5

What is the difference between offline and online evaluation of an LLM feature?

level: juniorimportance: must knowfreq 74%

answer

  1. two regimes, different questions
  2. frozen sample versus live population
  3. proxy metric versus real outcome
  4. cheap fast filter, slow expensive verdict
  5. shadow, split, canary as escalating exposure

basics

~20 s

Offline evaluation scores a fixed, frozen set of saved examples in a harness before release, so it is repeatable and cheap. Online evaluation measures real user traffic after release, where actual behaviour and business outcomes decide whether the change helped.

solid answer

~50 s

Offline evaluation runs the candidate system over a **frozen dataset** of inputs with known expected behaviour, scoring each output with programmatic checks or a judge model. It is fast, repeatable, costs nothing but tokens, and can be run on every change — but it only ever measures the system against examples someone chose in advance. Online evaluation measures the system on **live traffic**: shadow runs that score real requests without showing the output, A/B splits that compare arms on product metrics, canaries that expose a small slice first. It is the only place you learn what real users, on today's request mix, actually do with the output. The two answer different questions. Offline answers "did I break anything I already know about?" Online answers "did this help?" A mature program uses offline as the cheap gate before exposure and online as the decision of record — treating an offline number as proof of production value is a classic mistake.

go deeper

for a junior

Be able to say plainly that offline evaluation scores a fixed saved dataset before release and online evaluation measures real users after, and that a good offline score is not yet evidence the feature helps anyone.

for a middle

Explain the mechanics of each: how a frozen set is scored and why it is repeatable, and the online forms — mirrored shadow requests, randomized splits, small canary slices — along with what each one costs in exposure and time.

for a senior

Show the judgment of running both as a funnel: which failures you gate offline, when exposure is justified, how long an online read really takes, and how production failures get folded back into the offline set.

for a principal

Own the tradeoff at program level: how much of the release budget goes to fast cheap filters versus slow trustworthy measurement, when the organization is allowed to ship on an offline number alone, and how you keep offline metrics honest as predictors of online outcomes.

## The two halves of an eval program An LLM feature has no single correct output, so "is it good?" cannot be answered by a unit test. Practitioners split the question into two measurement regimes that run at different times, on different data, and answer genuinely different questions. **Offline evaluation** happens before exposure. You take a fixed collection of inputs — a *frozen dataset*, meaning it does not change between runs — feed each one through the candidate prompt, model or pipeline, and score the outputs. Scoring can be programmatic (does the JSON parse, does the SQL run, does the answer contain the required entity), reference-based (compare against a stored expected answer), or a judge model applying a written rubric. The output is a number per slice: accuracy on refunds, faithfulness on long documents, format-validity overall. **Online evaluation** happens on live production traffic. Its forms escalate in exposure and in what they can tell you: mirroring real requests to a candidate whose output is scored but never shown; splitting users between a control and a treatment arm and comparing product metrics; exposing a tiny slice first and ramping it if nothing burns. The scores here are not rubric grades but outcomes — did the shopper buy, did the ticket get resolved without escalation, did the agent finish the task. ## What offline is good at Offline evaluation is **repeatable, cheap, and safe**. The same dataset scored twice gives roughly the same number (modulo sampling nondeterminism), so you can compare a candidate against a baseline directly. No user sees a bad output. A run costs tokens and minutes, not customers, so you can afford it on every change and iterate ten times a day. It is also the only way to test rare and dangerous cases *on purpose*. Live traffic will not conveniently contain the prompt-injection attempt, the malformed invoice, or the multilingual edge case at the moment you need to test them; a curated set holds them permanently. And it lets you slice: overall quality can rise while one important segment collapses, and a frozen set with segment labels shows that immediately. ## What offline can never tell you An offline set is a **sample someone chose**, scored by a **proxy for value**. Both assumptions can be wrong. The inputs may not match today's traffic. A set assembled last quarter carries last quarter's request mix; live traffic drifts with season, campaigns, new user segments and new failure modes. The score is then honest about a population you no longer serve. The metric may not proxy what users reward. A rewriter can produce more relevant, better-formed queries by every rubric a judge applies, and still make people buy less because the results feel unfamiliar. Offline metrics measure output properties; businesses are paid in user behaviour, and the mapping between the two is an assumption, not a measurement. And offline cannot capture **response**: users react to outputs, learn the system, change how they phrase requests, sometimes stop trusting it. A frozen harness has no user in it to react. ## What online is good at, and what it costs Online evaluation measures the thing you actually care about, on the population you actually have, including effects nobody thought to encode in a rubric. It is the decision of record. Its costs are real. Exposure means someone gets the worse variant. Statistical power means waiting — small effects on noisy business metrics take days or weeks of traffic to resolve. Attribution is hard: outcomes are influenced by pricing, seasonality, other teams' experiments. Turnaround is measured in days, so you cannot iterate on prompt wording this way. And you cannot A/B-test something catastrophic; that is what the offline gate is for. ## How they compose in practice The standard shape is a funnel. Every change is scored offline against the frozen set; a regression on a tracked slice stops it there. Changes that pass are exposed narrowly — first as scored shadow runs on mirrored requests, then to a small canary slice, then to a randomized split with a pre-declared primary metric and a set of harm metrics that can stop the rollout. Whatever the online experiment discovers is fed back as new offline cases, so the frozen set slowly learns the failure modes production found. The two are not competing methods and neither substitutes for the other. Offline is a **fast, cheap filter with weak external validity**; online is a **slow, expensive measurement with strong external validity**. Confusing the roles produces both classic failures: shipping on an offline number and being surprised, or A/B-testing every prompt tweak and shipping nothing. ## What interviewers listen for A weak answer describes offline as "testing" and online as "monitoring". A strong answer states the epistemic difference — chosen sample and proxy metric versus real population and real outcome — and names the concrete online mechanisms (shadow, split test, canary, harm metrics) with the exposure and latency each one costs you.

  • If online evaluation is the decision of record, why keep an offline suite at all?
    Because online measurement is slow, expensive and exposes users. You cannot iterate on prompt wording in week-long increments, you cannot deliberately test dangerous inputs on customers, and rare failure modes may not appear in a week of traffic. Offline is the cheap filter that keeps obviously broken candidates from ever reaching an experiment, and the only place rare cases are guaranteed to be tested every run.
  • Which direction does information flow between the two?
    Both ways, but the important flow is online back into offline. Production surfaces failures nobody anticipated; each one becomes a new case in the frozen set so the same regression cannot ship twice. Offline flows forward only as a gate — a pass buys the right to be exposed, never a claim of value.
  • How does the choice of metric differ between the two regimes?
    Offline metrics are properties of the output — format validity, faithfulness to a source, rubric scores, task success against a stored expectation. Online metrics are properties of the user's response — completion, purchase, escalation, retry. Offline metrics are chosen because they are believed to predict online ones, and that belief should itself be checked whenever an experiment disagrees with the harness.

saying these in an interview costs you the question

  • Treats an offline score as proof the change improves the product
  • Describes online evaluation as merely uptime and error monitoring
  • Assumes a frozen eval set stays representative indefinitely
  • Thinks offline and online are alternatives, so a team picks one
  • Claims A/B tests can replace offline suites for every prompt change

context

open as a page

Why do offline eval wins often fail to reproduce online, as when a grocery search-query rewriter gains 9 points offline but loses 0.4% cart conversion?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Three causes dominate: the frozen set no longer matches live traffic, the offline metric proxies something users do not reward, and the harness has no user in it to react. A large offline gain on the wrong population or the wrong proxy buys nothing online.

open as a page

In an LLM rollout, what can shadow traffic measure and what can it never measure?

level: middleimportance: should knowfreq 47%

basics

~20 s

Shadow traffic mirrors real requests to a candidate whose output is scored but never shown, so it measures behaviour on the true request distribution at real latency and cost with zero user risk. It can never measure user response, because no user ever sees the output.

open as a page

When should a guardrail metric stop a rollout whose primary metric is up?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Whenever a pre-declared harm metric breaches its threshold, even if the primary metric improves. Guardrails encode costs the primary metric ignores — refunds, escalations, safety violations, latency, spend — and they are set before the experiment precisely so a good headline number cannot argue them away.

open as a page

How do canary rollouts and A/B tests differ in what they control for?

level: principalimportance: should knowfreq 33%

basics

~20 s

A canary controls risk: a tiny exposed slice bounds blast radius while you watch for breakage, and it is read as a safety check. An A/B test controls inference: randomized arms and sufficient sample size let you attribute a measured effect to the change.

open as a page