An agent red-team harness run ends mid-chain because the tool endpoint the agent was calling returned rate-limit errors, not because the agent refused or a policy control fired. How do you score that run, and what do you change in the harness?
answer
- content decision vs rate decision
- inconclusive, not defended
- stop-cause taxonomy
- successes over valid runs
- throttling biases toward long chains
basics
~20 sScore it as an invalid run, not a defence. Rate limiting is infrastructure back-pressure from the endpoint, not a model refusal or a policy control; it would not stop a patient attacker. Mark the run inconclusive, exclude it from the denominator, back off and retry, and log the cause.
solid answer
~60 sThe rule is that a stop only counts as a defence if it is a decision about the *content* of the request. A rate limit is a decision about the *rate* of requests, so it stops this run and not the attack: an adversary simply waits, or spreads the chain over hours. Operationally that means three things. First, classify the outcome — the harness needs a stop-cause taxonomy (refused, blocked by control, infrastructure error, agent error, timeout) so an error never silently lands in the "defended" bucket. Second, fix the denominator: inconclusive runs come out of the attack-success-rate base rather than counting as failures, or the number quietly understates risk. Third, make the harness resilient — respect back-off, retry the affected step, cap retries, and fail the run loudly as inconclusive if the endpoint stays unavailable. The one caveat: if throttling is a deliberate, documented control on that agent's tool use, it is worth reporting as a *mitigating* factor — but as friction that raises attacker cost, never as the reason a chain did not complete.
go deeper
Recognises that a rate-limit error is a technical failure of the run, not the agent being stopped, and that the run should be retried.
Names the distinction — content decision versus rate decision — and knows the run belongs in an inconclusive bucket outside the success-rate denominator.
Designs the stop-cause taxonomy, fixes back-off and concurrency in the harness, reports discard counts, and notices that throttling biases the result toward the most severe chains.
Makes stop-cause classification a required field for any number the organisation quotes, so no team can publish a rate built on unclassified terminations.
This is the classic silent miscount in agent-harness scoring: a run that never got a verdict is filed as if the target won it. **Why a rate limit is not a defence.** A defence makes a judgement about *what was asked* — a model refusal, a policy classifier, a tool permission, a human confirmation gate. Throttling makes a judgement about *how many requests arrived per unit time*, and it is completely indifferent to their content. An attacker with patience is unaffected: the same chain paced at one call a minute completes. Worse, the correlation runs backwards from where you want it. Throttling fires hardest against your own high-volume automated suite and softest against a slow, deliberate real attack — so the mechanism you are about to credit is precisely the one that will not be there when it matters. **The taxonomy.** Every terminated run needs exactly one label, and only two of them are defence evidence: | Stop cause | Counts as | Why | |---|---|---| | Completed the harmful action | Success | The path is open; existence proof | | Refused by the model | Defence evidence | A judgement about content | | Blocked by a named control | Defence evidence | Permission, classifier, confirmation gate | | Agent error (malformed call, planning loop, context exhausted) | Inconclusive | Capability artefact, not a control | | Infrastructure error (429, 5xx, network, sandbox crash) | Inconclusive | Nothing judged the request | | Budget exhausted (turn cap, wall-clock cap) | Inconclusive | A limit you imposed on the test | **The denominator, and which way the bias runs.** Attack-success rate must be computed over valid runs — successes divided by (completed + refused + blocked) — and the discard count published beside it. A suite where 40% of runs die on throttling does not have a safe target; it has an integrity problem, and the two look identical in a single headline number. The bias has a direction, too: throttling correlates with long chains and heavy tool use, which are exactly the most severe scenarios, so silently dropping throttled runs skews the reported rate optimistic. If you must choose one number to publish, publish the pair. **What it costs.** A throttled run is money already spent for no verdict — you paid the tokens of every turn up to the 429 and got nothing gradeable back, so a suite with a 40% throttle rate is running at roughly 60% capital efficiency. Back-off makes wall-clock the binding constraint rather than tokens: honouring a Retry-After of 60 seconds inside a 25-turn episode can stretch a two-minute run into half an hour, and a few hundred tasks then cross from an overnight job to a multi-day one. That is the real reason teams raise concurrency and then quietly accept the throttles — the fix costs schedule, and the miscount costs nothing visible. **Harness changes I would make.** Honour Retry-After and back off exponentially with jitter rather than fixed sleeps. Retry the failed step rather than replaying the whole chain where the environment allows resumption, and record that a retry happened, because a resumed chain is not quite the same observation as a clean one. Cap suite concurrency so the harness stops throttling itself, and separate the target's request budget from the tooling's so a busy scorer cannot starve the agent. Add a health probe before and during the suite so a dead endpoint aborts the run rather than emitting a page of fake defences. Finally, surface stop-cause counts on the run report itself — if the miscount is only visible by grepping logs, it will not be visible. **The caveat worth stating.** If throttling is a deliberate, documented control on that agent's tool use, report it — but as attacker-cost friction with the bypass named (wait; spread the chain over hours; use parallel sessions), not as the reason the chain did not complete. Friction that raises cost is a legitimate mitigation to record in a report; it is never the answer to "was this attack blocked?". **What I would check.** Does the report show a stop-cause histogram? Is the success rate labelled with its base and its discard count? Do throttled runs get retried, and does the retry policy have a cap? And is the discard rate itself trended, since a rising one usually means the harness, not the target, changed.
- Where else does the same miscount happen in agent harness scoring?Turn-budget exhaustion, sandbox crashes, context-window overflow, and a tool that errors on a malformed argument. Each ends the chain without any control having judged the request.
- Throttling is a documented control on this agent's tool calls. Does that change the score?It changes the report, not the score. Note it as friction that raises attacker cost, state the bypass (wait, or spread the chain), and still exclude the run from the defence evidence.
- Why is discarding throttled runs without reporting the count dangerous?Throttling hits long, tool-heavy chains hardest — the severe cases. Dropping them silently biases the attack-success rate optimistic. Publish the discard count alongside the rate.
A rate limiter is a turnstile counting how many people pass per minute, not a guard checking who they are. Anyone willing to arrive slowly is never inspected at all.
saying these in an interview costs you the question
- Counting the throttled run as a blocked attack.
- Retrying blindly with no cap so the suite hammers the endpoint.
- Dropping inconclusive runs silently and quoting the rate as if the base were unchanged.
- Calling rate limiting a security control with no mention of the patience bypass.
- Having no stop-cause field at all, so every non-completion looks alike.