You own the partial-progress grading rubric for an agent red-team harness that several teams run against different agents. How do you decide how many progress tiers the rubric has, and who arbitrates a disputed grade?
answer
- tiers must be decision-bearing
- reversibility is the sharpest line
- world-state check, not intent
- outcome-shaped, not step-count
- one arbiter, rulings become examples
basics
~20 sTie tiers to decisions, not to narrative detail. Use the fewest checkpoints that change what someone does: attempted, reached the irreversible action, completed. Every tier needs an objective world-state check and a named arbiter for disputes, or graders drift and cross-team numbers stop comparing.
solid answer
~50 sTwo failure modes bracket the choice. Too few tiers and every stop looks alike, so you lose the signal that distinguishes a refusal at framing from a block at the last hop. Too many and each is a judgement call, graders disagree, and numbers stop being comparable between teams — which is the whole reason a shared rubric exists. My rule is that a tier must be **decision-bearing** (someone acts differently depending on which side of it a run lands) and **objectively checkable** (a query against the sandbox state, not a reading of the transcript). Three or four usually survives that test: attempt made, capability or target object acquired, irreversible action staged, action completed. Governance matters as much as the scale. Publish worked examples per tier, require an attributed stop cause with every grade, sample-audit grades across teams, and name a single arbiter for disputes with the ruling folded back into the examples. Version the rubric, because a changed scale silently breaks every trend line drawn across it.
go deeper
Understands that a rubric needs clear, checkable tiers and that graders must apply them the same way.
Argues the granularity tradeoff and insists each tier be checkable against environment state rather than the transcript.
Runs it: worked examples, required stop-cause and environment fields, sample audits, disagreement rate as a health metric, versioned rubric.
Sizes the scale from the decisions it feeds, anchors a tier on reversibility, names a single arbiter with rulings folded back into the rubric, and keeps the tier out of leaderboard use.
The rubric is an instrument, and choosing its resolution is a tradeoff between **signal** and **inter-rater reliability** — the same tradeoff as choosing how many gradations to put on a ruler that several people will read. ## Sizing the scale Start from the decisions the numbers feed: which findings get escalated, which controls get funded, which release gets held. If moving a run from tier two to tier three changes none of those, that tier is decoration and it will cost you in disagreement without buying anything. The durable boundaries are the ones with a physical meaning in the environment: 1. the agent obtained the sensitive object; 2. the agent staged the irreversible action; 3. the agent executed it. **Reversibility** is the sharpest line available and it deserves a tier of its own, because everything before it is recoverable and everything after is not, and that is exactly the distinction an incident responder cares about. Three or four tiers usually survives this test; ten never does. ## Objectivity before granularity Each tier needs an assertion someone can run against the sandbox after the fact: the file exists, the row is gone, the message left the outbox. Tiers defined by inferred intent ("the agent was clearly trying to") cannot be graded twice the same way and will not survive a dispute. This is also why the transcript is never the source of truth: a model that narrates a step it did not take, or omits one it did, corrupts every intent-based tier at once while leaving **world-state tiers** untouched. ## Comparability across teams Different agents have different tool sets, so an absolute step count does not transfer — step four of a mail agent's chain and step four of a finance agent's chain are not the same event, and averaging them produces a number with no referent. Define tiers by the structure of the harmful outcome rather than by the specific chain, and both agents can be graded "staged the irreversible action". Require every grade to carry three fields alongside the tier: - the environment snapshot version, - the trial count, - and the attributed stop cause. Without those, cross-team aggregation is fiction dressed as a dashboard. ## What the governance costs Be honest about this when you propose it, because an unfunded rubric rots. - **Worked examples** per tier are a day of writing and a re-write every time a tier definition changes. - **Sample audits** — independently re-grading a fixed percentage of runs, five to ten is typical — cost grader hours every cycle and are the only way you learn the disagreement rate. - **Arbitration** costs the arbiter's time and a rubric edit per ruling. - **Versioning** costs a reference-set re-grade at every change. Set the granularity you can actually afford to audit; a six-tier rubric nobody audits is less trustworthy than a three-tier one that is audited, because its extra precision is unverified. ## Where these numbers mislead Two ways, both structural. - First, an **unaudited disagreement rate** hides inside the aggregate: if two graders assign different tiers to the same trace 20% of the time, your tier distribution has 20% noise that no amount of sample size removes, and trend lines drawn through it are reading grader turnover as system change. - Second, and worse, the moment a tier distribution appears on a **team scoreboard** the instrument stops measuring the system and starts measuring the graders — tasks get scoped to land in comfortable tiers, disputes get argued in one direction, and the number improves while the posture does not. ## Arbitration Name one arbiter, not a committee — a rotating grading owner works well and spreads the context. The process: 1. the disputed grade is filed with the trace and the world-state check that was used; 2. the arbiter rules; 3. the ruling is appended to the rubric as a worked example so the same dispute is not re-litigated. Track the disagreement rate from the audits as a health metric of the rubric itself — a rising rate means the scale has grown too fine or a tier's definition has rotted. ## Versioning and blast radius Changing the scale invalidates comparisons across the change. Version the rubric, re-grade a reference set under the new scale, label charts by rubric version, and never renumber tiers in place. And hold the framing line: the tier is a **severity input** for the person who owns the risk decision, not a leaderboard column.
- How do you know the scale has too many tiers?Independent graders disagree on the same trace, and no decision changes when a run moves one tier. Track the disagreement rate from sample audits and collapse tiers that nobody acts on.
- Why not define tiers as a step count out of the chain length?Chain length varies by agent and tool set, and steps are not equal in severity. An outcome-shaped scale — object acquired, action staged, action completed — transfers across teams; a fraction does not.
- What happens to trend lines when you change the rubric?They break. Version the rubric, re-grade a reference set under the new scale, and label charts by rubric version rather than silently splicing two scales together.
saying these in an interview costs you the question
- A tier per tool call, giving a scale nobody can grade twice the same way.
- Tiers defined by inferred intent instead of a checkable environment fact.
- Comparing step fractions across agents with different tool sets.
- Changing the scale in place and continuing the same trend chart.
- Letting the tier distribution become a team scoreboard.
- No arbiter, so disputes resolve by whoever argues longest.