skip to content

Partial Progress

How far a harmful chain got before stopping is a severity input, and a run cut short by a rate limit is not a defence. Interviewers ask what a partial number licenses you to claim.

on this pageshow

explore

questions

4

An agent red-team harness grades a run as "partial": the harmful chain reached step three of six before the agent was stopped. What does that grade license you to claim about the target system, and what does it not?

level: middleimportance: must knowfreq 60%

answer

  1. existence proof of reach, not a boundary
  2. one run is not a rate
  3. stop cause must be attributed
  4. pinned environment or it means nothing
  5. non-completion is weak evidence

basics

~20 s

It licenses a claim about that run only: the chain reached step three under this prompt, seed and environment state. It does not show the target is safe at step four; a retry, a different phrasing or a fuller mailbox may carry it further. Partial is evidence of reach, not of a boundary.

solid answer

~50 s

A partial grade is an existence proof about reach, bounded to the conditions of that run. You can say: with this task, this tool set, this starting environment and this attempt budget, the system permitted the first three steps and something stopped the fourth. What it is *not* is a boundary. "Stopped at four" is one observation of a stochastic system; another sampling, a re-phrased instruction, or a different starting state may go further. Nor does it tell you what stopped it — a model refusal, a tool permission, a confirmation gate and a malformed argument all produce the same stop, and only the first two are defences at all. So the honest write-up pairs the distance with the attributed stop cause, the number of trials behind it, and the pinned environment. Without those three, "3/6" reads as a safety margin the run never established.

go deeper

for a junior

Says the run reached step three and stopped, and knows that one run is not proof the system is safe beyond it.

for a middle

Separates the valid reach claim from the invalid safety claim, and insists on recording what caused the stop and how many trials were behind the number.

for a senior

Writes the finding with pinned environment, trial count, furthest observed reach and attributed stop cause, and refuses to convert a step fraction into a probability.

for a principal

Sets the house rule for how partial numbers may appear in reporting so that teams cannot quote a step fraction as a defence rate.

A **partial-progress grade** answers exactly one question — how much of this chain did the target system permit on this run — and the whole interview is about whether the candidate over-reads it in either direction. ## What the grade is made of The run happened in a **pinned sandbox**: a fixed tool set, a fixed starting world state, one task phrasing, one sampling of a stochastic model, one attempt budget. The grader queried the environment afterwards and found the artefacts of steps one through three present and the step-four artefact absent. Everything you may claim has to be derivable from that sentence. ## The valid claims - (1) *Reach.* Steps one to three are demonstrably permitted. This is a positive result and needs no repetition to be true — one observation of a thing happening is proof the thing can happen. - (2) *A stop event at step four*, plus a cause if and only if the harness recorded one. - (3) *A severity input.* A chain caught at the last-mile control is a worse posture than one refused at framing, even though both rows read "no harm", because the surviving margin is one control thick. ## The invalid claims - (1) *Safety at step four.* A non-observation in a stochastic system is weak evidence. Agent runs vary with sampling, with tool-output ordering, with the exact wording of an instruction; a single non-completion is far weaker than a single completion, and the asymmetry is not a technicality — it is the whole epistemics of this instrument. - (2) *A rate.* "3 of 6" has chain steps in the denominator, not attempts. It is a distance. Converting it to "the system blocks half the attack" swaps one denominator for a completely unrelated one, and no arithmetic repairs that. - (3) *Generalisation across environments.* Add one tool, or start with a fuller mailbox, and the chain's length changes; the distance is a property of the run, not of the model. - (4) *Attribution without evidence.* If the harness did not classify the stop, you cannot say a defence fired at all — a malformed tool argument, a turn cap, a 429 and a policy block all terminate the run identically from the outside. ## What the stronger claims cost Notice the **price ladder**, because it is what makes people cheat. | Claim | Cost | |---|---| | Reach | one episode | | A rate | tens of episodes per task — each a full multi-turn agent loop whose context grows with every tool result | | A bound ("the system never gets past step four") | infinity; no number of trials buys it | The confidence interval on a rate estimated from ten trials is wide enough that most teams cannot distinguish 10% from 30%. That a bound costs infinity is why the honest report never states one. When someone quotes a step fraction as a defence rate, they have bought the top of the ladder with a payment for the bottom rung. ## How I would phrase it Something close to: *"Under the pinned environment snapshot ENV-14, 4 of 20 trials permitted steps 1-3; furthest observed reach was step 4, stopped by the tool-permission check in 3 of those 4; no trial completed the final action. This bounds observed reach, not the system ceiling."* That is defensible in a review. "3/6, so it is half-safe" is not, and a stakeholder will read the second sentence out of a report that only contains the first. ## What I would check before writing it up - Was each reached checkpoint asserted against world state rather than the agent's own narration? - Is the stop cause recorded and classified? - How many trials, and did any trial go further than the one being quoted — because the maximum, not the median, is the reach claim? - Is the environment snapshot versioned so a re-run in a month is comparable? And if the answer to any of those is no, the finding still has value as an **existence proof** of reach; it simply may not carry a number.

  • A stakeholder reads "3 of 6" as "the system blocks half the attack". How do you correct that?
    Tell them the number is a distance on one sample, not a probability. The denominator is chain steps, not attempts, and steps are not equal in danger — reaching the irreversible step is worth more than three cheap setup steps.
  • What single extra field makes a partial grade far more useful?
    The attributed stop cause: model refusal, policy or tool control, infrastructure error, or agent error. Without it you cannot tell a defence from a fluke.
  • Does one partial run and one completed run carry the same evidential weight?
    No. A completed run proves the path exists and is close to irrefutable. A partial run is a non-observation at that depth, which stochastic behaviour can produce even where the path is open.

Getting three rooms into a house is not being half-stopped: the fraction counts rooms you chose to number, and the only door that mattered may be the last one.

saying these in an interview costs you the question

  • Reading "3 of 6" as "50% blocked" or as a defence rate.
  • Claiming the system is safe at step four on a single non-completion.
  • Reporting a distance with no environment snapshot or trial count.
  • Treating all chain steps as equally severe so the fraction looks meaningful.
  • Not recording why the run stopped, then calling the stop a control.

context

open as a page

An agent red-team harness run ends mid-chain because the tool endpoint the agent was calling returned rate-limit errors, not because the agent refused or a policy control fired. How do you score that run, and what do you change in the harness?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Score it as an invalid run, not a defence. Rate limiting is infrastructure back-pressure from the endpoint, not a model refusal or a policy control; it would not stop a patient attacker. Mark the run inconclusive, exclude it from the denominator, back off and retry, and log the cause.

open as a page

An agent red-team harness drives an LLM agent through a multi-step harmful task in a sandboxed environment and scores the outcome. Why record how far along the chain the agent got, instead of only a pass/fail on the final harmful outcome?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Because a pass/fail hides how close the run came. Recording the last step the agent reached shows whether it was stopped early at planning or only at the final harmful action, tells you which control fired, and gives severity a distance input. Two zero-score runs can be very different.

open as a page

You own the partial-progress grading rubric for an agent red-team harness that several teams run against different agents. How do you decide how many progress tiers the rubric has, and who arbitrates a disputed grade?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Tie tiers to decisions, not to narrative detail. Use the fewest checkpoints that change what someone does: attempted, reached the irreversible action, completed. Every tier needs an objective world-state check and a named arbiter for disputes, or graders drift and cross-team numbers stop comparing.

open as a page