AI agents
LLM agents that decide what to do next: call a tool, read the result, and loop until the goal is met or the budget runs out. Interviewers use this area to separate people who have shipped an agent — with its stopping conditions, cost ceilings, and failure modes — from people who have chained two prompts.
on this pageshowhide
guide
overview
~1 minAn AI agent is a language model placed inside a loop. On each pass it chooses what to do next, usually by requesting a tool; your code carries out the request, feeds the outcome back into the model's input, and the cycle repeats until the task is finished or something outside the model ends it. Interviewers use this subject to separate candidates who have run an agent against real systems from those who have chained a couple of prompts together. The questions are operational: what the model can see, how the loop stops, what a run costs, which actions need a person's sign-off, and how you prove the job was done. The subject splits into six sections. [Tool use and function calling](/topics/found-ai-agents-tool-use) is the contract between model and code, including what happens when a tool fails. [Planning and reasoning](/topics/found-ai-agents-planning) covers choosing the next move, step by step or from an upfront plan, and repairing a plan that stops matching reality. [Agent memory](/topics/found-ai-agents-memory) is about what survives between steps and between sessions. [Agent loops and control flow](/topics/found-ai-agents-loops) is the operational layer: termination, budgets, self-correction and human checkpoints. [Multi-agent orchestration](/topics/found-ai-agents-multi-agent) splits work across isolated workers, and [agent evaluation](/topics/found-ai-agents-eval) asks how you measure any of it. Learn tool use first, because every other section assumes you know what a single request-and-result round trip looks like. The loop comes next, since stopping conditions and spending limits are where production incidents concentrate. Planning and memory build on that; evaluation and orchestration come last, once you can describe one agent end to end. Questions run from a junior asked what the model actually emits to a principal asked when a many-agent system is worth its token bill.
primer
A few ideas carry most of this subject. Hold these and the questions below read as consequences rather than trivia. - **The model proposes; the harness acts.** A model never runs anything. It produces a structured request, and ordinary code decides whether to execute it, with what permissions, and what to report back. Validation, retries, access control and audit all live in that surrounding code, so most questions about agent safety are really questions about the harness. - **The window is the model's entire world.** Tool names, descriptions and parameter schemas are all it knows about its capabilities; tool results are its only fresh view of the environment; the conversation history is rebuilt by your code on every call. Anything left out does not exist for the model, and anything put in competes for a finite budget. - **A loop needs an exit the model does not control.** The model finishing its answer is the happy path. Iteration caps, deadlines, spending limits and detection of repeated steps are what end the other runs, and interviewers expect you to name them before being asked. - **Autonomy should track reversibility.** Reading and cheap, undoable writes can run unattended; actions that move money, delete data or reach outside parties warrant a checkpoint. The line is drawn by consequence, and the checkpoint has to survive a person taking days to answer. - **Correction needs outside evidence.** Asking a model to review itself mostly re-samples the same judgment. Revision improves results when something external, such as a failing test, a validator or a tool error, points at the defect. - **Success is a state of the world, not a sentence.** An agent announcing that it finished proves nothing. Check the end state independently, look at the path it took, and repeat runs, because the same agent on the same task will not behave identically twice. - **Add structure only when the task demands it.** A fixed workflow is simpler to test than an agent when the steps are known in advance, and one agent is simpler than several unless the work genuinely splits. Each extra layer of autonomy has to pay for its cost and its new failure modes.
- Agent loop
- The repeating cycle of model call, tool execution and appending the result, run by ordinary code until a stop condition fires.
- Tool definition
- The name, natural-language description and parameter schema that tell a model a capability exists and how to request it.
- Tool call
- A structured request emitted by the model instead of prose, naming a tool and supplying arguments, which the host code decides whether to execute.
- ReAct
- A loop pattern that interleaves a reasoning step, an action and an observation, choosing each action from the most recent result.
- Plan-and-execute
- A pattern in which the model writes the full sequence of steps first and an executor then runs them with little further reasoning.
- Working memory
- What the agent holds for the task in progress, typically the current context plus any scratchpad or plan file it keeps.
- Semantic memory
- Durable facts about a user, account or domain, stored outside the model and retrieved when relevant.
- Stop condition
- Any rule that ends the loop: the model finishing, an explicit done signal, or a forced limit on steps, time, tokens or errors.
- Human-in-the-loop checkpoint
- A point where the loop pauses for a person to approve, edit or reject a proposed action before it runs.
- Orchestrator
- The agent that owns the overall goal, divides it among workers, and combines what they return into the final result.
- Subagent
- A worker agent running its own loop in a separate context window on one assigned piece of work, returning a condensed result.
- Trajectory
- The full sequence of reasoning, tool calls and observations an agent produced on one task, as opposed to its final output alone.
- pass^k
- The share of tasks an agent solves on every one of k repeated attempts; a measure of reliability rather than best-case ability.
- Goal drift
- An agent gradually pursuing a different objective than the one it was given, as early instructions lose influence over a long run.
The six sections describe one system from different angles, and each depends on the ones before it. **Tool use is the unit of work.** One request, one execution, one result appended to the conversation. Everything else is a policy about how to repeat, order, remember or check that unit. The schema and description decide whether the model picks the right tool; the error path decides whether a failure becomes information or a crash. **The loop repeats the unit and decides when to stop.** [Agent loops and control flow](/topics/found-ai-agents-loops) wraps tool use in stopping rules, budgets and human checkpoints. Reflection and self-correction live here too, because a revision round is just another pass through the loop with a verifier's output as its input. **Planning decides which unit comes next.** Step-by-step reasoning picks one action per observation; an upfront plan picks many at once and has to be repaired when a step contradicts it. The [replanning and plan repair](/topics/found-ai-agents-planning-replanning-repair) material overlaps with the loop section, since replanning without a cap is one of the ways a loop fails to end. **Memory decides what is in the window.** Every model call sees a context assembled by your code. [Agent memory](/topics/found-ai-agents-memory) covers what gets written to durable storage and what gets read back into that context, which is the same scarcity problem tool results and tool catalogs create, approached from the storage side. **Orchestration is the loop applied recursively.** An orchestrator can be read as an agent whose tools are other agents. It inherits every concern above and adds two of its own: what each worker is allowed to change, and how much detail each worker hands back. **Evaluation reads all of it.** [Agent evaluation](/topics/found-ai-agents-eval) scores the trajectory as well as the answer, so it needs the vocabulary of every other section: whether the tool choice was right, whether the loop ended for a good reason, whether memory supplied a stale fact. Its failure taxonomy doubles as an index back into the other sections.
- Tool Use and Function Calling →
The request-and-result round trip is the unit every other section repeats, orders or measures, so learn exactly what crosses the boundary first.
- Agent Loops and Control Flow →
Stopping rules, budgets and approval gates turn a single call into a safe loop; most production questions start here.
- Planning and Reasoning →
Once the loop is familiar, compare step-by-step reasoning with upfront plans and learn how a broken plan is repaired.
- Agent Memory →
Long tasks outgrow the window; this is where you decide what to store, what to recall and how to keep it trustworthy.
- Agent Evaluation →
With one agent understood end to end, learn to measure it: trajectories, repeated runs, verifiers and a taxonomy of failures.
- Multi-Agent Orchestration →
Last, because orchestration inherits every earlier concern and adds coordination cost; judge when isolation is worth it.
Describing the model as executing tools; it only emits a request, and forgetting that hides where validation, permissions and retries belong.
Relying on the model to decide when it is finished, with no iteration cap, deadline or spending limit enforced by the harness.
Letting a tool failure surface as an unhandled exception instead of a result the model can read and recover from.
Returning raw, unbounded tool output into the context and then trying to trim history afterwards instead of reducing the output at its source.
Gating approvals by tool name rather than by what the action can break, or recording an approval that is not tied to the exact arguments executed.
Treating the context window as memory; it is rebuilt every call, so anything not stored and deliberately read back is gone.
Reading from a shared memory store without a scope filter derived from the authenticated user, which can leak one user's facts into another's run.
Adding reflection rounds with no external check and expecting accuracy to rise; without new evidence a second pass can make a correct answer worse.
Scoring only the final answer from a single run, which hides lucky successes, wasted steps and the variance repeated attempts would reveal.
Reaching for several agents when one would do; subagents add isolation, not intelligence, and cost far more tokens and coordination.
The same handful of choices recurs across the sections. Naming the one you are making, and what would change your mind, is most of a senior answer. - **Autonomy versus control.** A free-running loop adapts to surprises; a fixed workflow or an upfront plan is cheaper, faster and easier to audit. The deciding factor is how predictable the steps are, and many good designs mix the two at different levels of the same task. - **Context richness versus context cost.** More tool definitions, fuller results and longer history give the model more to work with and also more to be confused by, at a price paid on every call. Loading tools on demand, summarising results and compressing history all trade some fidelity for focus. - **Speed versus oversight.** Every human checkpoint adds latency and depends on someone answering. Too many and people approve by reflex. The consequence of the action, not its frequency, should set the level. - **Freshness versus cost in memory.** Capturing facts continuously costs a model call per turn; batching them at the close of a session saves calls but risks losing work. Recall has the same shape: injecting memories upfront is fast, fetching on demand is precise. - **Isolation versus coordination.** Separate workers keep noisy exploration out of the main context and can run in parallel, but they work blind to one another and multiply token spend. The work has to split cleanly for the trade to pay.
Several shapes turn up in many sections under different names. Recognising one is often the quickest route to an answer. - **Turn failure into input.** Tool errors returned as results, failing tests fed to a revision round, a stale-plan signal that triggers repair: the loop keeps running because the problem arrives as information the model can use. - **Enforce limits outside the model.** Iteration caps, token budgets, replan caps per subgoal, reflection round limits and repeated-step detection are all counters kept by the harness, because the model is not a reliable judge of when to give up. - **Externalise state the window cannot hold.** A plan file, a memory store, a durable checkpoint for a paused run and a subagent's short summary all move information out of the context so it survives truncation, restarts or delegation. - **Load on demand instead of upfront.** Searching a large tool catalog, recalling memories through a tool call and fetching detail only when a step needs it keep the always-present context small. - **Check against something the agent did not write.** Postcondition checks on a step, programmatic verifiers in evaluation, a separate critic prompt and an authenticated approver all supply judgment that does not share the agent's blind spots. When a question looks unfamiliar, ask which of these it is testing; the [failure taxonomy](/topics/found-ai-agents-eval-failure-taxonomy) section maps many of them to the failures they prevent.
explore
- Tool Use and Function Calling20 questions
- Tool and Function Schemas5 questions
- Invocation Protocol and Parallel Calls5 questions
- Result Handling and Tool Errors5 questions
- Tool Selection at Scale5 questions
- Planning and Reasoning14 questions
- ReAct vs Plan-and-Execute5 questions
- Task Decomposition4 questions
- Replanning and Plan Repair5 questions
- Agent Memory15 questions
- Working, Episodic and Semantic Memory5 questions
- Memory Write Policies5 questions
- Memory Retrieval Policies5 questions
- Agent Loops and Control Flow13 questions
- Reflection and Self-Correction5 questions
- Termination, Budgets and Loop Detection4 questions
- Human-in-the-Loop Checkpoints4 questions
- Multi-Agent Orchestration5 questions
- Agent Evaluation14 questions
- Trajectory Scoring5 questions
- Eval Harness Design5 questions
- Agent Failure Taxonomy4 questions
questions
81 · 6 sectionsIn LLM function calling, what does the model emit and what turn must you append?
basics
~20 sInstead of prose the model returns a structured tool-call block naming the tool, its arguments and a call id. Your code executes the tool, appends a tool-result turn carrying that same id, and calls the model again with the extended history.
In a function-calling tool definition, which parts does the model actually see?
basics
~20 sThe model sees only three things: the tool's name, its natural-language description, and the JSON Schema describing its parameters. Implementation code, docstrings and internal comments never reach it, so everything it needs must live in those three fields.
In LLM tool calling, how are parallel tool results returned and paired to their calls?
basics
~20 sOne assistant turn can carry several independent tool-call blocks. All of them are answered in a single following turn holding one result block per call, each matched by the call id it echoes — never by array position or completion order.
When a tool call fails, why return an is_error tool result rather than letting the exception propagate?
basics
~20 sAn exception ends the agent's turn; a tool result flagged as an error keeps the loop alive and hands the model a fact it can act on. The model can then fix an argument, pick another tool, or report honestly that the step failed.
Why does an agent's tool-selection accuracy fall as its catalog grows from 20 to 400 tools?
basics
~20 sEvery tool definition is loaded into the prompt at once, so hundreds of them cost tens of thousands of tokens and force the model to discriminate between many near-identical options in a single pass. Crowding plus semantic overlap pushes selection accuracy down.
In a ReAct agent loop, what do the thought, action and observation steps each do?
basics
~20 sThought is the model's reasoning about what to do next, action is the tool call it emits, and observation is the real tool result appended back into the prompt. The loop repeats until the model answers instead of acting.
Why do agents write the plan to an external todo file instead of keeping it in context?
basics
~20 sA written plan survives what the conversation does not. Context gets truncated, summarized or crowded out on long runs, so an external plan file keeps the goal and per-step status stable and re-readable, and doubles as an audit trail of what the agent actually did.
How does plan-and-execute differ from ReAct in when the LLM makes decisions?
basics
~20 sReAct calls the model once per step and picks each action from the latest observation. Plan-and-execute calls the model once upfront to write the whole step list, then runs those steps with little or no further reasoning.
How does an LLM agent detect that its plan has gone stale mid-execution?
basics
~20 sDetection comes from checking each step against an expected outcome instead of assuming success. Three signals dominate: an explicit tool error, tool output that contradicts an assumption the plan was built on, and a postcondition check on the step's result that fails.
How do you pick subtask granularity when an agent decomposes a goal?
basics
~20 sSize each subtask so its completion can be checked objectively and it still fits one focused stretch of work. Too coarse and nobody can tell whether it succeeded; too fine and per-step overhead costs more than the work itself.
In an AI agent, what are working, episodic, semantic and procedural memory?
basics
~20 sWorking memory holds what the agent is using right now for the current task. Episodic memory records what happened in past sessions. Semantic memory stores durable facts. Procedural memory holds learned how-to — reusable skills the agent can apply again.
When is agentic memory recall better than pre-injecting top-k memories in an agent?
basics
~20 sAgentic recall — the agent calling a memory tool when it notices it needs a fact — wins when most turns need no memory and precision matters. Pre-injected top-k wins for a small always-relevant profile and tight latency budgets.
Why is an AI agent's context window not the same thing as its memory?
basics
~20 sThe context window is the input assembled for one model call — bounded, rebuilt by your code every time, and gone afterwards. Memory is durable state stored outside the model that the agent deliberately writes and reads back. Continuity comes from the store, not the window.
When should an agent write to long-term memory: per turn, at session end, or on an explicit remember tool?
basics
~20 sWrite triggers trade freshness for cost. Per-turn extraction catches everything but adds a model call to every turn. End-of-session extraction is cheaper and better informed. An explicit remember tool is precise but captures only what someone thought to flag.
How do scope and metadata filters on agent memory reads prevent cross-user leakage?
basics
~20 sEvery memory read must be constrained by a scope key — user, tenant, session or project — derived from the authenticated session, applied inside the query before ranking. Semantic search has no notion of ownership, so nothing but an explicit filter keeps one user's memories out of another's context.
Which AI-agent tool calls need human approval, and how do you tier the rest?
basics
~20 sGate by blast radius, not by tool name. Anything irreversible, externally visible, or above a value threshold — money movement, production schema changes, outbound messages, deletions — needs a human. Reads and cheaply reversible writes run automatically.
Why does agent self-correction improve results with a test suite but often not without one?
basics
~20 sSelf-correction needs a signal the model did not produce. Failing tests, compiler errors and schema violations are outside evidence of a defect. Pure self-judgment re-samples the same model that wrote the output, so it rarely catches what it already missed.
What stop conditions end an agent loop besides the model returning no tool call?
basics
~20 sAgent loops end naturally when the model returns a final answer with no tool call, or calls an explicit done tool. They end forcibly on a max-iteration cap, a wall-clock deadline, a token or cost budget, or an unrecoverable error.
How do you pause an AI agent for human approval that may take days?
basics
~20 sPersist the loop's state to a durable checkpoint, emit the approval request, and let the process exit. When the human answers — hours or days later, on a different machine — load the checkpoint, inject the decision, and resume from the interrupt point. Never block a thread.
Why can asking an LLM "are you sure?" turn a correct answer into a wrong one?
basics
~20 sA challenge like "are you sure?" supplies doubt, not evidence. Models tend to go along with implied disagreement, so the second pass often swaps a correct but unusual detail — an exact date, an odd spelling — for a more ordinary-sounding one, lowering accuracy.
Why is context isolation, not extra reasoning, the main gain from LLM subagents?
basics
~20 sSpawning subagents adds no reasoning capacity — it is the same model. The gain is that each worker's noisy intermediate output stays in its own window, so the orchestrator attends to a small set of distilled findings instead of a bloated transcript full of dead ends.
In a multi-agent LLM system, how does an orchestrator agent differ from a subagent?
basics
~20 sThe orchestrator owns the goal: it splits the work, delegates, and assembles the answer. Each subagent runs its own loop in a separate context window, sees only its assigned task, and hands back a short summary rather than its full transcript.
How would you choose between LangGraph, CrewAI and a vendor agent SDK?
basics
~20 sMatch the tool to the control you need. LangGraph gives explicit graph control flow with durable checkpointed state for production. CrewAI reaches a working prototype fastest with role-based crews. Vendor SDKs are the shortest path to that provider's newest capabilities.
Why do multi-agent systems need a single-writer rule for each shared artifact?
basics
~20 sIsolated workers cannot see each other's decisions, so two of them editing the same artifact make incompatible implicit choices that the orchestrator cannot reconcile from summaries. Keep writes with one designated agent; let the rest return read-only findings or proposals.
What are the main failure modes of a tool-using LLM agent, and what fixes each?
basics
~20 sTool-using agents fail in recognizable categories: invented tool names, invalid arguments, the wrong real tool, non-progressing loops, drift from the original goal, and false claims of success. Each has a different fix, so classifying the failure is the first debugging step.
In agent evals, what does pass^k measure that a single run per task hides?
basics
~20 spass^k is the share of tasks an agent solves on all k repeated attempts, so it measures reliability rather than best-case ability. A single run hides variance: roughly 90% per-attempt success clears all 8 attempts on only about 43% of tasks.
Why score an AI agent's trajectory and not just its final answer?
basics
~20 sFinal-answer scoring cannot tell a clean four-step solution from a forty-step flail that stumbled onto the same output, and it credits answers reached by luck. Trajectory metrics score the path: tool choices, ordering, and wasted steps.
How do you score tool selection and argument correctness in an agent trajectory?
basics
~20 sTreat the reference trajectory's tool calls as ground truth. Precision is the share of the agent's calls that appear in the reference; recall is the share of reference calls the agent made. Then, on matched calls only, score arguments separately.
An agent claims success after its tool returned HTTP 502 — what failure mode is this?
basics
~20 sA false success claim: the agent narrates completion over a failed side effect. It is dangerous because nothing crashed, so the run looks green while the work never landed. Only an independent post-condition check catches it.