skip to content

When should trajectory comparison be exact-match, any-order, or subsequence?

level: seniorimportance: should knowfreq 45%

answer

  1. strictness is a dial, not a switch
  2. would a human be wrong to swap them?
  3. extra steps in between may be fine
  4. false failures kill trust fastest
  5. constraints and budgets beat fixed lists

basics

~20 s

Match as strictly as the task genuinely constrains, and no stricter. Exact-match only where every step and its position are mandated; any-order for independent steps; subsequence when relative order matters but extra steps in between are acceptable.

solid answer

~50 s

Pick the comparison mode from the task's real ordering constraints. **Exact-match** demands the same calls in the same positions — appropriate only for rigid procedures, such as a regulated approval sequence, and it fails any agent that found a shorter valid path. **Any-order** compares call sets and ignores sequence, which is right when steps are independent: two lookups against different systems can happen in either order and neither is wrong. **Subsequence** requires that certain calls appear in relative order while tolerating other calls in between — the natural encoding of "create the ticket before attaching the log". Most real suites are hybrid: a few ordered constraints, some unordered groups, and a step budget instead of a fixed length. Over-strict matching is the more common mistake, because it manufactures false failures that erode trust in the eval faster than a missed regression does.

code

python · 12 lines
python
def ordered_subsequence(required, actual):
    it = iter(actual)
    return all(step in it for step in required)

actual = ["search_kb", "create_ticket", "notify_user", "attach_log"]

# relative order preserved, extra calls in between are fine
print(ordered_subsequence(["create_ticket", "attach_log"], actual))  # True
print(ordered_subsequence(["attach_log", "create_ticket"], actual))  # False

# independent group: order irrelevant, presence required
print({"search_kb", "notify_user"} <= set(actual))                   # True

go deeper

for a junior

Know that a trajectory can be compared strictly (same calls, same order), loosely (same calls, any order), or somewhere in between, and that the choice depends on whether the task really fixes the order.

for a middle

Be ready to pick a mode from a described task and justify it — independent lookups get any-order, a create-then-attach dependency gets subsequence — and to explain why a fixed-length exact match rarely fits.

for a senior

Show you build hybrid references: ordering constraints, unordered groups, required and forbidden calls, a step budget. Talk about validating the comparison against known-good trajectories before it gates anything.

for a principal

Own the asymmetry between the two errors. Argue why a noisy, over-strict eval degrades to being ignored, and why scoring conformance to one expert path quietly selects against agents that find better ones.

## The strictness dial Comparing an observed trajectory to a reference is not one algorithm but a family, ordered by how much freedom the agent is allowed. Choosing a point on that dial is a modelling decision about the task, not a technical detail — and getting it wrong produces either false failures (too strict) or blind spots (too loose). **Exact-match** requires the agent's call sequence to equal the reference: same tools, same order, same length. It is the strongest signal when it passes and almost always the wrong default. Reserve it for procedures where order is externally mandated — a regulated approval chain, a deployment runbook, a payment flow where the authorization must precede the capture and nothing may sit between them. **Any-order (set or multiset) comparison** discards sequence entirely and asks only whether the right calls were made, with the right multiplicity. Use it where the steps are genuinely independent: an internal expense-report portal agent that must read the employee's cost centre and the current per-diem policy can fetch either first, and penalizing one ordering is measuring nothing. **Subsequence matching** is the middle ground and the workhorse. It asks whether the required calls appear in the required relative order, allowing arbitrary other calls in between. "Create the ticket before attaching the log" is exactly a subsequence constraint: the agent may search the knowledge base, notify the user, and check permissions in between, and it is still correct — but attaching a log to a ticket that does not yet exist is not. ## Edit distance and partial credit Between exact-match and set comparison sits **sequence edit distance**: the number of insertions, deletions and substitutions that turn the observed trajectory into the reference, usually normalized by reference length. It gives graded credit instead of a pass/fail bit, which makes it useful for tracking gradual regression across model versions — a suite whose mean edit distance drifts from 0.8 to 1.6 is degrading before any test flips to fail. Its weakness is interpretability: a distance of 3 does not tell you whether the agent skipped a critical write or added three harmless reads. Pair it with the categorical checks rather than replacing them. ## Hybrids are what real suites use A well-built reference is rarely a flat list. It is closer to a small dependency graph plus budgets: - ordered constraints (`create_ticket` before `attach_log`), - unordered groups (the two independent lookups, in any order), - required calls (some retrieval tool must be called before answering), - forbidden calls (never call a refund tool on this task), - a step budget rather than a fixed length. Scored this way, an agent that solves the expense-report task in five steps where the reference took seven still passes — it satisfied every ordering constraint and came in under budget — while an agent that attached a log first fails on a specific, nameable constraint. That specificity is the practical payoff: the failure message says which rule broke, not "trajectory mismatch". ## Choosing the mode in practice Ask, per pair of steps: would a competent human doing this task be wrong to swap them? If no, the order is not a constraint and must not be scored as one. Ask, per step: is this step required, or was it merely how the reference author happened to proceed? Only required steps belong in the comparison at all. This is the discipline that turns a reference trajectory from a transcript into a specification. A useful sanity check is to run the suite against two or three *known-good* trajectories produced by different competent solvers — different models, or a human. If your comparison fails any of them, it is too strict, and you have learned that before it starts blocking releases. ## Why over-strictness is the more expensive error Both errors are real, but they fail differently. A too-loose comparison misses a regression, which you eventually catch in production or in the outcome metric. A too-strict comparison produces failures on correct behaviour, and the human response to a noisy eval is to stop reading it — after which it catches nothing at all. Strictness also actively penalizes improvement: a newer model that solves in three steps what the reference does in seven is scored as a regression, so the eval quietly selects for imitating an older, slower path. Start loose enough to admit every genuinely correct path you can produce, then tighten only where you can name the property being protected.

  • How would you encode a reference for a task that has several genuinely valid solution paths?
    Stop encoding a path and start encoding properties: required calls, forbidden calls, ordering constraints between specific pairs, and a step budget. Where several whole strategies are acceptable, keep a small set of reference trajectories and score against the best match. Where even that is impractical, fall back to outcome checks plus a rubric applied to the trace by a judge.
  • What does normalized sequence edit distance add over a pass/fail comparison?
    Graded credit, which makes slow degradation visible. A suite whose mean edit distance drifts upward is regressing before any individual test flips to fail, so it works as a leading indicator across model versions. Its weakness is that the number does not say which deviation happened, so keep the categorical constraint checks alongside it for diagnosis.
  • How do you tell whether your comparison is too strict before it blocks a release?
    Run it against trajectories you already believe are correct — from a different model, an earlier version, or a human solving the task. Any of those that fail exposes a constraint you asserted but the task does not actually impose. Doing this while authoring the reference is far cheaper than discovering it from a wall of red in CI.

saying these in an interview costs you the question

  • Defaults to exact-match because it is the strictest
  • Treats every reference step as a required step
  • Ignores order entirely and misses a write-before-create bug
  • Says a shorter correct path is a regression
  • Cannot name which constraint a failing trajectory broke

context