How would you build a regression eval that catches multi-turn instruction drift?
answer
- the unit is a conversation
- assert the same rule at several depths
- one run at turn fifty proves nothing
- a curve, not a pass/fail
- scripts that tempt the model off-rule
basics
~20 sReplay fixed conversation scripts and assert the same constraint at several depths — say turn 5, 20 and 50 — instead of testing one-shot. Run each script repeatedly because sampling is nondeterministic, and report an adherence curve per depth so a regression shows up as decay, not a single failure.
solid answer
~60 sSingle-turn test sets cannot see drift, because every case starts at turn one where adherence is at its best. A drift eval is structurally different: a **scripted conversation** — a fixed list of user messages replayed against the live prompt — with the same constraint **asserted at multiple checkpoints**, for example turns 5, 20 and 50. A suite of a dozen such scripts, each stressing a different rule and a different tempting derail, is usually enough. Three design points. Because sampling is nondeterministic, run each script several times and report an **adherence rate per checkpoint**, not pass/fail; the artefact you compare across prompt versions is the decay curve. Use **programmatic checkers** wherever the rule permits — a language detector for a Spanish-only tutor, a regex for a required unit, a schema check for output shape — and reserve a judge model for genuinely subjective things like persona. And include **adversarial scripts** where the user tempts the model off the rule mid-conversation, since that is where drift actually starts in production.
code
python · 12 linesdef run_drift_eval(respond, script, checkpoints, check, runs=5):
"""respond(history) -> assistant reply text; check(reply) -> bool"""
results = {turn: 0 for turn in checkpoints}
for _ in range(runs):
history = []
for turn, user_msg in enumerate(script, start=1):
history.append({"role": "user", "content": user_msg})
reply = respond(history)
history.append({"role": "assistant", "content": reply})
if turn in checkpoints and check(reply):
results[turn] += 1
return {turn: hits / runs for turn, hits in results.items()}go deeper
Know that a drift test replays a whole conversation instead of a single prompt, and that the same rule is checked more than once — early and late — to see whether it still holds.
Explain checkpointing at several turn depths, why repeated runs are needed under nondeterministic sampling, and why a deterministic checker beats a judge model whenever the rule can be expressed as code.
Demonstrate the operating view: derive checkpoint depths and history volume from real sessions, build adversarial derails into scripts, report per-rule decay curves, and set the regression alarm on depth-specific adherence rather than suite pass rate.
Own the economics and the feedback loop — how much long-conversation eval the organization buys, what runs on a prompt change versus a model upgrade, and how production adherence by conversation depth feeds back to keep the scripts honest.
## Why a normal test set misses this entirely A conventional prompt test set is a collection of inputs, each evaluated on its own. Every case is effectively turn one: short context, the standing instruction fresh and dominant, no history of the model's own output to imitate. Adherence there is near its ceiling. A prompt can score perfectly on that suite and still fall apart at turn forty in production. Catching drift requires an eval whose unit is **a conversation**, not a prompt. ## The scripted conversation The basic artefact is a fixed sequence of user messages — a script — replayed against the system under test. The harness sends message one, captures the reply, appends both to the history, sends message two, and so on. The script is deterministic; only the model's replies vary. This gives you comparable conversations across prompt versions and across models. A script must be long enough to reach the depth where drift appears. If your production conversations run to sixty turns, a ten-turn script proves nothing. Scripts should also grow the history realistically: if real sessions include pasted documents or long tool outputs, include comparable bulk, because token volume drives decay more than turn count. ## Checkpoints and the adherence curve The key idea is asserting the *same* constraint at several depths. Take a sixty-turn recipe-planning script where the standing rule is "always give quantities in grams." Place a question that forces a quantity at turn 5, turn 20 and turn 50, and check each reply for grams. What you get is not a boolean but a curve: 100% at turn 5, 96% at turn 20, 71% at turn 50. That curve is the deliverable. It tells you the depth at which the prompt stops holding, which is exactly the number you need to decide whether to re-anchor, and from which turn. A suite of roughly a dozen scripts, each carrying one or two rules and a different domain, keeps runtime affordable while covering the rules that matter. Resist growing it to hundreds of long scripts; long conversations are the most expensive eval cases you will ever run. ## Repetition under nondeterminism At any temperature above zero, one run of one script tells you almost nothing at turn fifty — a single miss may be noise. Run each script N times (five to ten is common) and report the fraction of runs that satisfied the constraint at each checkpoint. Compare versions on those rates with an eye to how noisy they are; a 3-point move on ten runs is not a regression. Even at temperature zero, results are not perfectly reproducible in practice, so repetition is still worth it. ## Checkers: programmatic first Prefer deterministic verification wherever the rule allows it, because it is cheap, stable and unarguable: - language identification for a "reply only in Spanish" rule; - a regex or unit parser for "quantities in grams"; - schema or JSON validation for output-shape rules; - string presence for a mandatory disclosure or reference; - length or sentence counting for concision rules. Reserve a model judge for things no checker can express — persona consistency, tone, whether the assistant "stayed in character." Keep that surface small; every judged dimension adds cost and its own reliability question. ## Adversarial scripts Happy-path scripts under-report drift. Real conversations contain pressure: the user switches language mid-chat, asks for "just a quick summary" that invites the model out of its format, changes topic entirely and comes back, or pastes content in a different style. Write scripts that include these derails deliberately, then check the constraint *after* the derail. A prompt that holds at turn fifty on a placid script and collapses two turns after a language switch has a very different production profile. ## Reporting and CI Store results as (script, checkpoint, rule, run) rows so you can slice by rule and by depth. The useful regression alarm is not "suite pass rate fell" but "adherence at depth 50 for rule X fell from 0.9 to 0.6" — drift regressions are localized. Because these runs are long and expensive, teams typically run the full suite on prompt changes and model upgrades rather than on every commit, with a short-script smoke subset in the fast path. ## Closing the loop with production The same measurement exists in production: bucket observed adherence by conversation depth or tokens of history. If production decays faster than the eval, your scripts are too clean — real sessions are longer, messier or more adversarial than you modelled. Feed those shapes back into the scripts.
- How do you keep the cost of a long-conversation eval suite under control?Keep the suite small and deliberate: about a dozen scripts, each carrying one or two rules, with depth chosen from real session lengths rather than padded arbitrarily. Cache what you can, run the full suite on prompt and model changes rather than every commit, and keep a short-script smoke subset in the fast path. Reduce repetitions on stable rules and spend the runs on the ones near the decay edge.
- What would make you distrust a drift eval that reports perfect adherence at turn fifty?Usually the scripts are too clean or too short. Check that history volume matches production, that the checkpoint questions actually force the constrained behaviour rather than allowing a reply that trivially satisfies the checker, and that at least some scripts include a derail — a language switch, a topic change, a request that invites the default format back. Then compare against production adherence bucketed by depth.
- Should the checkpoint questions be identical across scripts?The probe should reliably force the constrained behaviour, but identical probes across every script make the eval narrow and easy to overfit. Vary the surface wording and the surrounding context while keeping what the probe elicits constant, so a passing result means the rule held, not that the prompt was tuned to one sentence.
saying these in an interview costs you the question
- Tests adherence with a single-turn prompt suite
- Runs each conversation once and calls a miss a regression
- Reports pass/fail instead of adherence by depth
- Uses a judge model for rules a regex could verify
- Scripts are all happy-path with no mid-chat derail