skip to content

What does an LLM agent harness own that the model itself never will?

level: middleimportance: must knowfreq 70%

answer

  1. who keeps state between calls
  2. the loop owns what weights cannot
  3. event log, context builder, tool router
  4. checkpoints, permissions, budgets live outside
  5. model as stateless reasoning control plane

basics

~20 s

The harness owns everything durable and enforceable: the session event log, the context builder that assembles each call, the tool router that actually executes calls, checkpoints, permissions and observability. The model only reasons over whatever context the harness hands it.

solid answer

~50 s

Think of the model as a stateless reasoning control plane and the harness as the runtime around it. Between calls the model retains nothing; the harness keeps a durable **session event log**, a **context builder** that decides which events, tool results and instructions go into the next request, a **tool router** that validates arguments and actually executes calls, plus **checkpointing**, permission enforcement, budgets and tracing. The model decides *what* to do next; the harness decides *what the model can see*, *what it is allowed to run*, and *what is recorded*. This split is why most production quality work is context engineering rather than prompt wording: you can swap models and keep the harness, but a rule that exists only as a sentence in a prompt is a suggestion, whereas a rule enforced in the router is a guarantee.

go deeper

for a junior

Be able to say plainly that the model is stateless: the application resends what it wants the model to see, and any memory of earlier steps is something your code kept, not something the model retained.

for a middle

Name the harness components and what each does — session log, context builder, tool router, checkpoints — and explain that the model proposes a tool call while your code validates and executes it.

for a senior

Show the judgement of moving guarantees out of the prompt into code: permissions, budgets, retries and irreversible actions belong in the router, and you should be able to justify that with a failure you have seen.

for a principal

Own the portability and evaluation argument: a clean harness/model boundary is what lets you swap models, A/B two context builders, and reconstruct exactly what the model saw during an incident. Be ready to defend build-vs-adopt for the runtime itself.

## The two halves of an LLM feature An LLM application is not "a prompt plus an API call." End to end it is two distinct pieces: the **model**, which maps a context window to a next output, and the **harness** (also called the agent runtime), which is ordinary software that decides what goes into that window, executes whatever the model asks for, and remembers everything across calls. The model is stateless. Every call starts from nothing but the bytes you send. It cannot store a fact, enforce a permission, retry a failed call, or observe itself. Anything that must survive between calls, be guaranteed rather than encouraged, or be inspected later, must live in the harness. ## What the harness owns **Session event log.** An append-only record of everything that happened: user turns, model outputs, tool invocations and their results, checkpoints, errors, approvals. This is the durable state of the session. The context window is a *view* derived from it, not the state itself. **Context builder.** The component that turns the log plus external material into the next request. It decides which past turns survive verbatim, which are summarized, which tool results are replaced by handles, what system instructions and tool definitions are attached, and what retrieved material is injected. Most measurable quality movement in a mature LLM app comes from changes here. **Tool router.** The component that receives the model's requested call, validates it against the schema, checks whether this session is permitted to run it, applies timeouts and concurrency limits, executes it, and turns the result into something the context builder can use. Note the ownership: the model *proposes* a call; the router *disposes*. **Checkpointing.** Markers in the log that let a run be paused, resumed after a crash, or rewound to a prior decision point. In an operational setting — say an on-call copilot that is about to run a destructive rollback — a checkpoint before the decision is what makes the run recoverable rather than a one-way door. **Cross-cutting concerns.** Budgets (tokens, wall clock, spend), retries and idempotency, tracing and metrics, redaction, and multi-tenant isolation. All of these are properties of software, not of weights. ## What the model owns The model owns judgement: interpreting an ambiguous request, choosing the next action, composing arguments, deciding when the goal is met, and writing the final prose. That is genuinely hard to replicate in code, and it is why you pay for it. It does not own correctness of execution, safety of side effects, or memory. ## Why the split is the interview answer Three consequences make this more than terminology. **Guarantees vs suggestions.** "Never delete production data without approval" written in a system prompt is a probabilistic constraint. The same rule expressed as a router check is deterministic. Whenever a candidate proposes to fix a safety or correctness problem by adding a sentence to the prompt, the better answer is usually to move the rule into the harness. **Portability.** Models change every few months. A harness with a clean split — log, builder, router, checkpoints — absorbs a model swap as a configuration change. A harness whose behaviour is smeared into one giant prompt does not. **Debuggability.** When something goes wrong, you want to answer "what exactly was in the window on the call that went wrong?" That question is only answerable if the context builder is a real, testable component with recorded inputs, rather than string concatenation scattered through request handlers. ## Where teams get the split wrong The common failures are symmetrical. One is putting harness responsibilities in the model: asking the prompt to "remember" earlier facts, to "only call this tool when authorised," or to "stop after ten steps." These are state, authorization and control-flow — code's job. The other is putting model responsibilities in the harness: hard-coding branching logic that tries to anticipate every user intent, which produces a brittle decision tree that the model could have handled by reasoning. A useful test when designing: for each requirement, ask whether you would accept it being satisfied 95% of the time. If yes, prompting is fine. If no, it belongs in harness code. ## Architecture shapes along the spectrum The same components compose into different shapes. A **prompt chain** is a fixed sequence of calls where the harness fully controls the order — cheap, predictable, easy to evaluate. An **agent loop** hands control flow to the model: call, execute tools, append results, call again, until a stop condition. The loop is more capable and less predictable, and it consumes context on every iteration, which is why context-budget techniques matter far more there. Most real products are a chain with one or two agentic stages inside it, not a pure agent.

  • If you swap the underlying model for a newer one, which parts of the harness should change?
    Ideally almost none. The log, router, checkpointing and observability are model-agnostic. What typically needs retuning is the context builder — how much history to keep, how aggressive compaction is, how tool definitions are worded — plus effort or reasoning-budget settings, since newer reasoning models consume very different token profiles. If a model swap forces changes to your permission or state code, the split was leaking.
  • Where does a prompt-level instruction stop being enough and become harness code?
    The moment the requirement is a guarantee rather than a preference. Irreversible side effects, authorization, spend and step caps, PII redaction, and anything you would have to defend in an incident review all belong in code, because the model will violate them at some rate. Prompts remain the right home for style, tone, prioritisation and heuristics where an occasional miss is acceptable.
  • How do you decide between a fixed prompt chain and an agent loop for a new feature?
    Ask whether the sequence of steps is knowable in advance. If it is — extract, classify, format — a chain is cheaper, lower-latency, easier to evaluate and easier to debug. A loop earns its cost when the path depends on what earlier steps discover, and when the space of possible paths is too large to enumerate. Many teams ship a chain first and promote a single stage to a loop once they have evidence the fixed path fails.

The model is a brilliant consultant with total amnesia between meetings. The harness is the office around them: the filing cabinet, the briefing pack assembled before each meeting, the assistant who actually places the calls, and the badge system that decides which doors open.

saying these in an interview costs you the question

  • Says the model remembers previous turns on its own
  • Puts tool permission checks in the system prompt only
  • Treats the context window as the session's durable state
  • Calls any multi-step LLM app an agent regardless of who controls flow
  • Assumes swapping models requires rewriting application logic

context