How heavily should trajectory metrics gate an agent release without freezing it onto one path?
answer
- conformance is not the same as correctness
- block invariants, budget efficiency, watch the rest
- a reference is one expert's path
- references rot when tools change
- name the harm before you block on it
basics
~10 sLet outcome correctness and hard safety constraints block, and keep path-conformance metrics advisory or budgeted. Gating on similarity to a reference path selects for imitating yesterday's solution and scores genuine improvements as regressions.
solid answer
~50 sSplit trajectory signals into three tiers. **Blocking**: outcome correctness, plus constraints you can defend as invariants — no forbidden side-effecting call, no answer without the required retrieval, hard step and cost ceilings. **Budgeted**: efficiency measures such as median step count and redundant-action rate, gated against a moving baseline with a tolerance band rather than an absolute target, so drift is caught without punishing noise. **Advisory**: path-similarity scores and judge rubric criteria, tracked as trends and reviewed, not enforced per commit. The reason for the split is that a reference trajectory is one expert's solution, not a specification; a model that solves in three steps what the reference does in seven is an improvement that a conformance gate reports as a failure. The maintenance cost is real too — references rot as tools change, and a suite nobody can afford to update becomes a suite nobody trusts.
go deeper
Know that trajectory metrics are not all equal: some describe rules the agent must never break, and others merely describe how a particular expert solved the task.
Be ready to separate constraints you would block on — forbidden calls, missing required retrieval, runaway cost — from softer path-similarity scores that should only be tracked.
Show you have run gates in practice: budgets against a rolling baseline rather than fixed numbers, noise-aware thresholds, and a plan for the references that break every time a tool is renamed.
Own the whole tradeoff. Argue why conformance gating selects for yesterday's solution, cost the maintenance burden of reference suites honestly, decide which flows deserve full references at all, and define how the baseline is allowed to move.
## The failure this question is about Trajectory metrics are so much more informative than a pass/fail bit that the natural instinct is to gate on them. That instinct, followed all the way, produces an agent optimized to reproduce the trajectories your team wrote down eighteen months ago. Every improvement that arrives as a *shorter or different* path — a stronger model that skips a lookup because it can infer the value, a new bulk tool that replaces three calls with one — is scored as a regression. The eval stops measuring quality and starts measuring similarity to the past. ## A three-tier gate **Tier 1, blocking.** Only properties you would defend as invariants of correct behaviour: - outcome correctness on the task suite, against a baseline threshold; - forbidden-call violations — an unauthorized side-effecting tool, a call outside the tenant's scope; - required-call violations — answering a factual question with no retrieval in the trace; - hard ceilings on steps, wall-clock and cost per episode, set well above normal so they catch runaway behaviour rather than ordinary variation. These share a property: each is a rule about what must or must not happen, independent of how the agent chooses to work. **Tier 2, budgeted.** Efficiency signals — median and p95 step count, redundant-action rate, tokens per successful task. Gate these against a rolling baseline with an explicit tolerance (say, a 25% degradation in median steps) rather than an absolute number, so the check tracks the system instead of a snapshot of it. Budget breaches should require a human decision, not an automatic block: sometimes the extra steps buy a real quality gain and the right move is to move the baseline. **Tier 3, advisory.** Path-similarity scores against reference trajectories and judge rubric criteria. Publish them per run, alert on trends across runs, and review them when something else is already suspicious. Never let a single-commit movement in these numbers block a merge, because their run-to-run noise is comparable to the effects you would be reacting to. ## The maintenance economics nobody budgets for Reference trajectories are code, and they rot faster than code. Every tool rename, schema change, added capability or model upgrade can invalidate a batch of them. A suite of a few hundred references represents a standing maintenance obligation, and when it exceeds what the team will actually pay, the observable outcome is not that references get updated — it is that failures get waved through, which is worse than not having the check. Before committing to reference-based gating at scale, cost the upkeep honestly and prefer encodings that survive change: constraints and budgets referencing *capabilities* age far better than full call-by-call transcripts. A cheaper middle path is targeted references. Author full reference trajectories for the handful of flows where the path genuinely is the product — regulated procedures, destructive operations, anything with a compliance story — and use constraints plus outcome checks everywhere else. That concentrates the maintenance where it earns its keep. ## Tasks with no single right path Many tasks admit several equally good solutions; research and open-ended investigation tasks admit dozens. For those, a reference is not merely expensive, it is conceptually wrong — there is nothing to be the reference *of*. The honest instruments there are outcome checks on the artifact produced, invariant constraints, budgets, and a rubric applied by a judge. Trying to force a canonical path onto such tasks is how teams end up with evals that are simultaneously strict and uninformative. ## Deciding the weights A workable heuristic: a trajectory metric earns blocking status only if you can state the user-visible harm of violating it in one sentence. "It issued a refund it was not asked for" — block. "It exceeded the cost ceiling by 8x" — block. "It took a different but valid route to the same answer" — that is not harm, and gating on it buys nothing while costing you every future improvement that looks unfamiliar. Also plan for the gate's own evolution. When a model upgrade legitimately changes the shape of good trajectories, the correct response is to re-baseline deliberately — rescore a window of known-good runs, move the budgets, retire references the new capability makes obsolete — and to record that you did, so later comparisons are not silently made across an inconsistent baseline. A gate nobody is allowed to move eventually gets bypassed instead.
- How do you re-baseline trajectory budgets after a model upgrade legitimately changes path shape?Deliberately and on the record. Rescore a window of known-good runs under the new model, compare distributions rather than single tasks, move the budgets to the new medians with the same tolerance band, and retire references the new capability makes obsolete. Note the baseline change in the eval history so later comparisons are not silently drawn across two different regimes.
- When is full reference-trajectory gating actually the right call?When the path is the product: regulated approval chains, destructive operations, anything with a compliance or audit story where doing the right thing the wrong way is itself a failure. There the reference encodes a real requirement rather than one solver's habits, and the maintenance cost is justified because the procedure changes rarely and deliberately.
- Your suite's outcome pass rate is flat but median step count has climbed 40% over three releases. What do you do?Treat it as a leading indicator and investigate before it becomes an outcome regression. Slice by task type to see whether the growth is broad or concentrated, check redundant-action rate to separate genuine extra work from repetition, and look at whether recent changes added tools or context that made selection harder. Flat outcome with rising cost is usually slack being consumed, not stability.
saying these in an interview costs you the question
- Blocks releases on similarity to a reference trajectory
- Sets absolute step budgets that never move with the system
- Ignores maintenance cost when planning hundreds of references
- Insists every task has one canonical correct path
- Treats a shorter valid path as a regression to be fixed