In a multi-step harmful-task suite where most attempts end half-finished, why does scoring each task as a binary completed / not-completed throw away most of what the run cost you, and what would you report instead?
answer
- most attempts never finish
- depth is ordinal, keep it
- per-task mean vs pooled steps
- easy prefix inflates partial credit
- depth histogram split by stop reason
basics
~20 sBinary scoring maps every unfinished attempt onto one bucket, so an agent that stopped at the first call and one that stopped one call short look identical. You paid for the whole chain, so report chain depth: which required steps were reached per task, plus the reason each attempt stopped.
solid answer
~50 sMost attempts in a suite of this shape do not finish, so a binary column is mostly zeros and the zeros are not comparable. An agent that refused immediately, an agent that got seven of eight steps right, and an agent that crashed the mocked tool all land in the same cell. Report per-step completion instead: for each task, which of its required steps the attempt reached, with the argument-level criteria that made each step count. Then aggregate deliberately rather than by reflex. A mean of per-task completion fractions weights a three-step task the same as a twelve-step one; a mean over pooled steps weights long tasks more. Publish both, or publish the depth histogram and let the reader pick. Always carry the stop reason alongside the depth. Refused, exhausted the step cap, hit a harness or parsing error, and completed are four different facts, and collapsing them makes the suite look safer than it is.
code
json · 12 lines{
"task_id": "t-041",
"required_steps": 8,
"steps": [
{"n": 1, "reached": true, "criterion": "search tool called"},
{"n": 2, "reached": true, "criterion": "target field populated"},
{"n": 3, "reached": false, "criterion": "transfer tool called"}
],
"depth_reached": 2,
"stop_reason": "refused",
"trace_ref": "runs/2026-08-31/t-041.jsonl"
}go deeper
Says most runs do not finish, so a pass/fail column is nearly all failures and hides how far each attempt got.
Proposes per-step completion with the stop reason kept alongside, and can explain why an early refusal and a near-completion must not share a bucket.
Names concrete aggregation traps — unequal chain lengths, benign easy prefixes, dependent steps — and defends a specific headline number plus the histogram behind it.
Decides which single number the organisation is allowed to quote, and what the reporting format must preserve so a later run is comparable to this one.
**The data you already have.** In a suite of this shape the binary verdict is not a primitive — it is derived. The scorer walks the task's ordered criteria, decides each one against the trace, and ANDs the results. A pipeline that stores only the AND has thrown away the per-criterion booleans it computed a microsecond earlier. That is what makes binary reporting a *reporting* defect rather than an instrumentation one: the graded record costs nothing extra to keep, and keeping it is what makes a re-grade an afternoon of CPU instead of the whole sweep budget paid twice. **Why binary loses the signal.** The steps of a multi-step task are dependent by construction: you cannot reach step four without step three. Depth is therefore an ordinal measurement carrying real information, and a threshold at "all steps" discards every gradation below the top. When the completion rate is low — which it usually is, because most attempts in these suites end part-way — the binary column is near-constant. A near-constant column cannot separate two model versions, two guardrail configurations, or two system prompts, and a floor effect looks exactly like safety. **What a graded record looks like.** Per attempt: task id, the ordered required steps, a per-step verdict with the criterion that decided it, the stop reason, the token and turn counts, and a reference to the raw trace. Aggregation then becomes a reporting choice made in the open, rather than something baked irreversibly into the data at scoring time. **The traps in aggregating it.** | Trap | What it does to the number | |---|---| | Unequal chain lengths | A mean of per-task fractions weights a 3-step task like a 12-step one; pooling steps lets long tasks dominate. Both are defensible; the choice moves the headline, so state it. | | Easy prefixes | Many chains open with a benign lookup almost any agent performs. An agent that always does step one and never step two scores a non-zero mean that reads like partial danger and is nothing of the kind. | | Non-independence | Steps within an attempt are not independent trials, so intervals computed as if they were are far too tight. The attempt, not the step, is the sampling unit. | | Stop reason folded into depth | A limit-truncated attempt and a refusal both read as "stopped at 4". Keep them as separate series or you cannot tell a safe model from a small budget. | **Where the graded number itself misleads.** This is the reading to guard hardest against: **partial credit is not partial harm.** In most of these chains the damage, if any, is carried by one terminal step — the send, the transfer, the publish. An attempt that reached seven of eight steps caused nothing, and "87.5% completion" invites a reader to hear "87.5% of the harm". The fix is to mark, per task, which step is harm-bearing, and to report reaching *that* step as a separate series from mean depth. Mean depth is a sensitivity instrument for comparing configurations; the harm-bearing-step rate is the number that describes outcomes. Conflating them is how a suite produces a scary headline from a run in which nothing was accomplished — or, in the other direction, a comfortable mean that hides a small number of fully executed chains. **What it costs to do properly.** Nothing in extra model calls; the per-step verdicts fall out of the scoring pass. It costs storage for traces (a few kilobytes to a few hundred per attempt, so gigabytes at most for a large sweep — trivial against the run's dollar cost) and it costs schema discipline: once a report quotes mean depth, later runs must compute it identically or the trend is meaningless. **How to present it.** A histogram of attempts by depth reached, split by stop reason, plus two scalars: the count of fully completed chains and the count that reached the harm-bearing step. The completed count is what people will quote; the histogram is what tells you whether a change actually moved anything. **What to check.** Whether any task's step-one criterion is satisfiable by an entirely benign action; whether the depth distribution has a spike exactly at the configured step limit; whether re-running one configuration reproduces the histogram shape rather than just the headline; and whether the harm-bearing step is identified for every task, or whether the report is silently treating all steps as equally consequential.
- Your average per-task completion fraction rose after a guardrail change, but the count of fully completed chains did not move. What is the likely explanation?Attempts are getting further into the easy prefixes without ever finishing, or a few short tasks shifted and are over-weighted by the per-task mean. Check the depth histogram and re-aggregate over pooled steps before believing the change.
- Why is the attempt, not the step, the right sampling unit for uncertainty on these numbers?Steps within one attempt are dependent — later ones are only reachable if earlier ones succeeded — so treating them as independent trials produces intervals that are far too narrow.
Seven of eight steps toward a wire transfer moves no money. Partial credit measures progress, not damage, the same way ninety per cent of a bridge carries no traffic.
saying these in an interview costs you the question
- Reports one completion percentage with no chain-depth distribution behind it.
- Averages per-task fractions without noticing that chain lengths differ.
- Treats each step as an independent trial when computing confidence intervals.
- Folds cap-truncated attempts and refusals into the same 'did not complete' bucket.
- Deletes the raw trace after scoring, making any re-grading a full re-run.