skip to content

When an agent exhausts its budget mid-task, should it return partial results or continue?

level: principalimportance: nice to knowfreq 35%

answer

  1. treat it as a degradation policy
  2. does the task accumulate?
  3. never dress partial as complete
  4. continuing spends unagreed money
  5. check budget before irreversible steps

basics

~20 s

Return the partial result with an explicit exhausted marker whenever partial work is usable — a bug list missing the last level beats nothing at all. Continue only when the task is checkpointable and someone has knowingly authorised the extra spend.

solid answer

~60 s

Treat exhaustion as a **degradation policy**, not an error path. Ask two questions. First, is partial output valuable on its own? For accumulative tasks — bugs found, records reconciled, files reviewed — it usually is, so return what exists with a status that says the budget ran out and, ideally, what was left uncovered. For all-or-nothing tasks, partial output is worse than nothing because it invites the caller to act on an incomplete picture. Second, is continuing safe and paid for? Summarise-and-continue keeps a long run alive past the context window, but it silently spends more money, adds latency, and loses detail from the summarised history — so it belongs behind an explicit decision, not as an automatic default. The pattern that survives contact with users is: structure the tool contract so a partial result is a first-class shape with a status field and a coverage description, expose depth as an up-front choice so the caller sizes the budget before the run, and never let a partial result be presented as a complete one.

go deeper

for a junior

Know that an agent can run out of budget before finishing, and that returning what it found so far — clearly labelled as incomplete — is often more useful than returning an error.

for a middle

Explain the two branches: accumulative tasks where partial output is genuinely useful, and all-or-nothing tasks where a partial answer misleads. Say why the incomplete status belongs in the result structure rather than in prose.

for a senior

Show you have handled it in a live system — the status field and coverage note in the contract, why continuing costs money nobody agreed to, and what you do when the budget ran out midway through irreversible writes.

for a principal

Own it as a degradation policy across the product: who authorises extra spend, how depth tiers let callers size runs up front, how partial-result contracts keep incomplete work from reading as complete, and why a rising rate of hard-cap endings is a design defect rather than a tuning problem.

## Exhaustion is a product decision When an agent hits its iteration cap, its token budget, or its cost ceiling with work outstanding, the harness has three options: return what it has, keep going somehow, or fail. Most teams discover only in production that they picked the wrong default. This is a policy question about how your product degrades, and it deserves the same thought as any other partial-failure design. ## Is partial output valuable? The deciding property is whether the task **accumulates**. - **Accumulative tasks** produce independently useful units: bugs found, tests written, records reconciled, documents classified, files reviewed. A QA agent that explored six of ten levels and found eleven bugs has produced eleven real bugs. Returning them with a `budget_exhausted` marker and a note that levels seven to ten were not visited is straightforwardly better than returning an error, because the consumer can act on the eleven and re-run for the rest. - **All-or-nothing tasks** produce one answer that is either right or misleading: a migration plan, a diagnosis, a financial reconciliation total. Half a reconciliation is not half correct — it is a number that looks authoritative and is wrong. Here the honest response is a failure with the trace attached, not a partial artifact. The failure to avoid is returning partial output *shaped like* complete output. If the caller cannot distinguish "here is the full bug list" from "here is what I found before running out", you have converted a budget event into a silent correctness bug. This is why the status belongs in the result type rather than in a sentence of prose the model wrote. ## Is continuing legitimate? The usual mechanism for continuing past a window limit is to compress the history and carry on with a summary plus the working artifacts. It genuinely works for long-horizon tasks and is standard practice. But as an *exhaustion policy* it has three costs worth stating plainly: - It spends more money than the budget the caller agreed to. If your cost ceiling exists to protect a per-user margin, continuing silently defeats it. - It adds latency, often multiplying the wait for a user already waiting. - It loses information. What survives is what the summariser chose to keep, and a detail dropped at step forty cannot be recovered at step ninety. So continuing is reasonable when the task is genuinely long-horizon and checkpointable, when the extra spend is authorised (by a caller-set tier, an operator policy, or a human in the loop), and when the run is asynchronous enough that latency does not matter. It is unreasonable as a silent default on a user-facing budget cap. ## Design so exhaustion is rare and graceful The strongest answer moves the problem upstream: - **Give the model an advisory budget** so it paces itself and lands on an answer rather than being cut off. Most exhaustion events should be soft landings, not hard truncations. - **Expose depth as a caller choice.** Named effort tiers — a quick pass versus a thorough investigation — let the caller size the run before it starts, which is far better than discovering at the cap that they wanted more. - **Externalise progress.** If the agent writes findings to a file or a todo list as it goes, partial results exist by construction and a resumed run can pick up where the last one stopped, rather than repeating everything. - **Bound tool payloads** so one enormous observation cannot consume a budget the model was pacing against. ## Irreversibility changes the calculus One more axis: what state did the agent leave behind? A read-only research agent can be stopped anywhere. An agent halfway through a sequence of writes — three of seven migrations applied, half a batch of tickets updated — cannot simply be abandoned. Either make the steps idempotent so a re-run converges, or record enough checkpoint state that a resumed run knows exactly where it stopped, or refuse to begin the sequence unless the remaining budget can plausibly finish it. Checking the remaining budget *before* starting an irreversible sequence is the cheap version of this and is often skipped. ## What a strong answer sounds like A principal-level response does not pick one option. It says: classify the task as accumulative or all-or-nothing, make the partial shape explicit in the contract so nothing masquerades as complete, keep continuing behind an authorisation decision rather than a silent default, and push the real fix upstream into pacing and depth selection so exhaustion becomes the rare tail rather than the routine ending.

  • How should a partial result be represented so a caller cannot mistake it for a complete one?
    Put the status in the type, not the prose: a status field naming the exhausted budget, the units of work completed, and an explicit description of what was not covered. Callers then branch on the status. A model-written sentence saying "I ran out of time" is easy to lose, easy to omit, and impossible to check programmatically.
  • When is summarise-and-continue clearly the right call at exhaustion?
    When the task is genuinely long-horizon and checkpointable, the run is asynchronous so extra latency is cheap, and the additional spend is authorised by policy or a human. A nightly batch reconciliation fits; a user waiting on a chat response with a per-session cost ceiling does not.
  • The agent exhausts its budget after applying three of seven irreversible steps. What now?
    Do not simply return. Record exactly which steps completed, leave the system in a state a resumed run can recognise, and surface the incomplete sequence prominently to a human. Better still, check the remaining budget before starting a sequence of irreversible actions and decline to begin one you cannot plausibly finish.
  • How do you keep exhaustion rare rather than routine?
    Give the model an advisory budget so it paces and lands gracefully, let callers pick a depth tier up front, have the agent externalise findings as it works so partial results exist by construction, and bound tool payloads so one huge observation cannot consume the budget. Then measure the share of runs ending at a hard cap and treat a rising number as a defect.

saying these in an interview costs you the question

  • Returns partial output shaped exactly like a complete answer
  • Always continues past the budget by summarising, silently spending more
  • Treats exhaustion as a generic error with no partial payload
  • Ignores whether the task accumulates or is all-or-nothing
  • Abandons a run halfway through irreversible writes

context