skip to content

AI agents

5 roadmaps81 questionsupdated

LLM agents that decide what to do next: call a tool, read the result, and loop until the goal is met or the budget runs out. Interviewers use this area to separate people who have shipped an agent — with its stopping conditions, cost ceilings, and failure modes — from people who have chained two prompts.

on this pageshow

guide

overview

~1 min

An AI agent is a language model placed inside a loop. On each pass it chooses what to do next, usually by requesting a tool; your code carries out the request, feeds the outcome back into the model's input, and the cycle repeats until the task is finished or something outside the model ends it. Interviewers use this subject to separate candidates who have run an agent against real systems from those who have chained a couple of prompts together. The questions are operational: what the model can see, how the loop stops, what a run costs, which actions need a person's sign-off, and how you prove the job was done. The subject splits into six sections. [Tool use and function calling](/topics/found-ai-agents-tool-use) is the contract between model and code, including what happens when a tool fails. [Planning and reasoning](/topics/found-ai-agents-planning) covers choosing the next move, step by step or from an upfront plan, and repairing a plan that stops matching reality. [Agent memory](/topics/found-ai-agents-memory) is about what survives between steps and between sessions. [Agent loops and control flow](/topics/found-ai-agents-loops) is the operational layer: termination, budgets, self-correction and human checkpoints. [Multi-agent orchestration](/topics/found-ai-agents-multi-agent) splits work across isolated workers, and [agent evaluation](/topics/found-ai-agents-eval) asks how you measure any of it. Learn tool use first, because every other section assumes you know what a single request-and-result round trip looks like. The loop comes next, since stopping conditions and spending limits are where production incidents concentrate. Planning and memory build on that; evaluation and orchestration come last, once you can describe one agent end to end. Questions run from a junior asked what the model actually emits to a principal asked when a many-agent system is worth its token bill.

primer

A few ideas carry most of this subject. Hold these and the questions below read as consequences rather than trivia. - **The model proposes; the harness acts.** A model never runs anything. It produces a structured request, and ordinary code decides whether to execute it, with what permissions, and what to report back. Validation, retries, access control and audit all live in that surrounding code, so most questions about agent safety are really questions about the harness. - **The window is the model's entire world.** Tool names, descriptions and parameter schemas are all it knows about its capabilities; tool results are its only fresh view of the environment; the conversation history is rebuilt by your code on every call. Anything left out does not exist for the model, and anything put in competes for a finite budget. - **A loop needs an exit the model does not control.** The model finishing its answer is the happy path. Iteration caps, deadlines, spending limits and detection of repeated steps are what end the other runs, and interviewers expect you to name them before being asked. - **Autonomy should track reversibility.** Reading and cheap, undoable writes can run unattended; actions that move money, delete data or reach outside parties warrant a checkpoint. The line is drawn by consequence, and the checkpoint has to survive a person taking days to answer. - **Correction needs outside evidence.** Asking a model to review itself mostly re-samples the same judgment. Revision improves results when something external, such as a failing test, a validator or a tool error, points at the defect. - **Success is a state of the world, not a sentence.** An agent announcing that it finished proves nothing. Check the end state independently, look at the path it took, and repeat runs, because the same agent on the same task will not behave identically twice. - **Add structure only when the task demands it.** A fixed workflow is simpler to test than an agent when the steps are known in advance, and one agent is simpler than several unless the work genuinely splits. Each extra layer of autonomy has to pay for its cost and its new failure modes.

Agent loop
The repeating cycle of model call, tool execution and appending the result, run by ordinary code until a stop condition fires.
Tool definition
The name, natural-language description and parameter schema that tell a model a capability exists and how to request it.
Tool call
A structured request emitted by the model instead of prose, naming a tool and supplying arguments, which the host code decides whether to execute.
ReAct
A loop pattern that interleaves a reasoning step, an action and an observation, choosing each action from the most recent result.
Plan-and-execute
A pattern in which the model writes the full sequence of steps first and an executor then runs them with little further reasoning.
Working memory
What the agent holds for the task in progress, typically the current context plus any scratchpad or plan file it keeps.
Semantic memory
Durable facts about a user, account or domain, stored outside the model and retrieved when relevant.
Stop condition
Any rule that ends the loop: the model finishing, an explicit done signal, or a forced limit on steps, time, tokens or errors.
Human-in-the-loop checkpoint
A point where the loop pauses for a person to approve, edit or reject a proposed action before it runs.
Orchestrator
The agent that owns the overall goal, divides it among workers, and combines what they return into the final result.
Subagent
A worker agent running its own loop in a separate context window on one assigned piece of work, returning a condensed result.
Trajectory
The full sequence of reasoning, tool calls and observations an agent produced on one task, as opposed to its final output alone.
pass^k
The share of tasks an agent solves on every one of k repeated attempts; a measure of reliability rather than best-case ability.
Goal drift
An agent gradually pursuing a different objective than the one it was given, as early instructions lose influence over a long run.

The six sections describe one system from different angles, and each depends on the ones before it. **Tool use is the unit of work.** One request, one execution, one result appended to the conversation. Everything else is a policy about how to repeat, order, remember or check that unit. The schema and description decide whether the model picks the right tool; the error path decides whether a failure becomes information or a crash. **The loop repeats the unit and decides when to stop.** [Agent loops and control flow](/topics/found-ai-agents-loops) wraps tool use in stopping rules, budgets and human checkpoints. Reflection and self-correction live here too, because a revision round is just another pass through the loop with a verifier's output as its input. **Planning decides which unit comes next.** Step-by-step reasoning picks one action per observation; an upfront plan picks many at once and has to be repaired when a step contradicts it. The [replanning and plan repair](/topics/found-ai-agents-planning-replanning-repair) material overlaps with the loop section, since replanning without a cap is one of the ways a loop fails to end. **Memory decides what is in the window.** Every model call sees a context assembled by your code. [Agent memory](/topics/found-ai-agents-memory) covers what gets written to durable storage and what gets read back into that context, which is the same scarcity problem tool results and tool catalogs create, approached from the storage side. **Orchestration is the loop applied recursively.** An orchestrator can be read as an agent whose tools are other agents. It inherits every concern above and adds two of its own: what each worker is allowed to change, and how much detail each worker hands back. **Evaluation reads all of it.** [Agent evaluation](/topics/found-ai-agents-eval) scores the trajectory as well as the answer, so it needs the vocabulary of every other section: whether the tool choice was right, whether the loop ended for a good reason, whether memory supplied a stale fact. Its failure taxonomy doubles as an index back into the other sections.

  1. Tool Use and Function Calling →

    The request-and-result round trip is the unit every other section repeats, orders or measures, so learn exactly what crosses the boundary first.

  2. Agent Loops and Control Flow →

    Stopping rules, budgets and approval gates turn a single call into a safe loop; most production questions start here.

  3. Planning and Reasoning →

    Once the loop is familiar, compare step-by-step reasoning with upfront plans and learn how a broken plan is repaired.

  4. Agent Memory →

    Long tasks outgrow the window; this is where you decide what to store, what to recall and how to keep it trustworthy.

  5. Agent Evaluation →

    With one agent understood end to end, learn to measure it: trajectories, repeated runs, verifiers and a taxonomy of failures.

  6. Multi-Agent Orchestration →

    Last, because orchestration inherits every earlier concern and adds coordination cost; judge when isolation is worth it.

  • Describing the model as executing tools; it only emits a request, and forgetting that hides where validation, permissions and retries belong.

  • Relying on the model to decide when it is finished, with no iteration cap, deadline or spending limit enforced by the harness.

  • Letting a tool failure surface as an unhandled exception instead of a result the model can read and recover from.

  • Returning raw, unbounded tool output into the context and then trying to trim history afterwards instead of reducing the output at its source.

  • Gating approvals by tool name rather than by what the action can break, or recording an approval that is not tied to the exact arguments executed.

  • Treating the context window as memory; it is rebuilt every call, so anything not stored and deliberately read back is gone.

  • Reading from a shared memory store without a scope filter derived from the authenticated user, which can leak one user's facts into another's run.

  • Adding reflection rounds with no external check and expecting accuracy to rise; without new evidence a second pass can make a correct answer worse.

  • Scoring only the final answer from a single run, which hides lucky successes, wasted steps and the variance repeated attempts would reveal.

  • Reaching for several agents when one would do; subagents add isolation, not intelligence, and cost far more tokens and coordination.

The same handful of choices recurs across the sections. Naming the one you are making, and what would change your mind, is most of a senior answer. - **Autonomy versus control.** A free-running loop adapts to surprises; a fixed workflow or an upfront plan is cheaper, faster and easier to audit. The deciding factor is how predictable the steps are, and many good designs mix the two at different levels of the same task. - **Context richness versus context cost.** More tool definitions, fuller results and longer history give the model more to work with and also more to be confused by, at a price paid on every call. Loading tools on demand, summarising results and compressing history all trade some fidelity for focus. - **Speed versus oversight.** Every human checkpoint adds latency and depends on someone answering. Too many and people approve by reflex. The consequence of the action, not its frequency, should set the level. - **Freshness versus cost in memory.** Capturing facts continuously costs a model call per turn; batching them at the close of a session saves calls but risks losing work. Recall has the same shape: injecting memories upfront is fast, fetching on demand is precise. - **Isolation versus coordination.** Separate workers keep noisy exploration out of the main context and can run in parallel, but they work blind to one another and multiply token spend. The work has to split cleanly for the trade to pay.

Several shapes turn up in many sections under different names. Recognising one is often the quickest route to an answer. - **Turn failure into input.** Tool errors returned as results, failing tests fed to a revision round, a stale-plan signal that triggers repair: the loop keeps running because the problem arrives as information the model can use. - **Enforce limits outside the model.** Iteration caps, token budgets, replan caps per subgoal, reflection round limits and repeated-step detection are all counters kept by the harness, because the model is not a reliable judge of when to give up. - **Externalise state the window cannot hold.** A plan file, a memory store, a durable checkpoint for a paused run and a subagent's short summary all move information out of the context so it survives truncation, restarts or delegation. - **Load on demand instead of upfront.** Searching a large tool catalog, recalling memories through a tool call and fetching detail only when a step needs it keep the always-present context small. - **Check against something the agent did not write.** Postcondition checks on a step, programmatic verifiers in evaluation, a separate critic prompt and an authenticated approver all supply judgment that does not share the agent's blind spots. When a question looks unfamiliar, ask which of these it is testing; the [failure taxonomy](/topics/found-ai-agents-eval-failure-taxonomy) section maps many of them to the failures they prevent.

explore

report an issue with this guide →

questions

81 · 6 sections

In LLM function calling, what does the model emit and what turn must you append?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Instead of prose the model returns a structured tool-call block naming the tool, its arguments and a call id. Your code executes the tool, appends a tool-result turn carrying that same id, and calls the model again with the extended history.

open as a page

In a function-calling tool definition, which parts does the model actually see?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The model sees only three things: the tool's name, its natural-language description, and the JSON Schema describing its parameters. Implementation code, docstrings and internal comments never reach it, so everything it needs must live in those three fields.

open as a page

In LLM tool calling, how are parallel tool results returned and paired to their calls?

level: middleimportance: must knowfreq 66%
basics
~20 s

One assistant turn can carry several independent tool-call blocks. All of them are answered in a single following turn holding one result block per call, each matched by the call id it echoes — never by array position or completion order.

open as a page

When a tool call fails, why return an is_error tool result rather than letting the exception propagate?

level: middleimportance: must knowfreq 68%
basics
~20 s

An exception ends the agent's turn; a tool result flagged as an error keeps the loop alive and hands the model a fact it can act on. The model can then fix an argument, pick another tool, or report honestly that the step failed.

open as a page

Why does an agent's tool-selection accuracy fall as its catalog grows from 20 to 400 tools?

level: middleimportance: must knowfreq 72%
basics
~20 s

Every tool definition is loaded into the prompt at once, so hundreds of them cost tens of thousands of tokens and force the model to discriminate between many near-identical options in a single pass. Crowding plus semantic overlap pushes selection accuracy down.

open as a page

In a ReAct agent loop, what do the thought, action and observation steps each do?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Thought is the model's reasoning about what to do next, action is the tool call it emits, and observation is the real tool result appended back into the prompt. The loop repeats until the model answers instead of acting.

open as a page

Why do agents write the plan to an external todo file instead of keeping it in context?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A written plan survives what the conversation does not. Context gets truncated, summarized or crowded out on long runs, so an external plan file keeps the goal and per-step status stable and re-readable, and doubles as an audit trail of what the agent actually did.

open as a page

How does plan-and-execute differ from ReAct in when the LLM makes decisions?

level: middleimportance: must knowfreq 78%
basics
~20 s

ReAct calls the model once per step and picks each action from the latest observation. Plan-and-execute calls the model once upfront to write the whole step list, then runs those steps with little or no further reasoning.

open as a page

How does an LLM agent detect that its plan has gone stale mid-execution?

level: middleimportance: must knowfreq 62%
basics
~20 s

Detection comes from checking each step against an expected outcome instead of assuming success. Three signals dominate: an explicit tool error, tool output that contradicts an assumption the plan was built on, and a postcondition check on the step's result that fails.

open as a page

How do you pick subtask granularity when an agent decomposes a goal?

level: middleimportance: must knowfreq 65%
basics
~20 s

Size each subtask so its completion can be checked objectively and it still fits one focused stretch of work. Too coarse and nobody can tell whether it succeeded; too fine and per-step overhead costs more than the work itself.

open as a page

In an AI agent, what are working, episodic, semantic and procedural memory?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Working memory holds what the agent is using right now for the current task. Episodic memory records what happened in past sessions. Semantic memory stores durable facts. Procedural memory holds learned how-to — reusable skills the agent can apply again.

open as a page

When is agentic memory recall better than pre-injecting top-k memories in an agent?

level: middleimportance: must knowfreq 62%
basics
~20 s

Agentic recall — the agent calling a memory tool when it notices it needs a fact — wins when most turns need no memory and precision matters. Pre-injected top-k wins for a small always-relevant profile and tight latency budgets.

open as a page

Why is an AI agent's context window not the same thing as its memory?

level: middleimportance: must knowfreq 58%
basics
~20 s

The context window is the input assembled for one model call — bounded, rebuilt by your code every time, and gone afterwards. Memory is durable state stored outside the model that the agent deliberately writes and reads back. Continuity comes from the store, not the window.

open as a page

When should an agent write to long-term memory: per turn, at session end, or on an explicit remember tool?

level: middleimportance: must knowfreq 62%
basics
~20 s

Write triggers trade freshness for cost. Per-turn extraction catches everything but adds a model call to every turn. End-of-session extraction is cheaper and better informed. An explicit remember tool is precise but captures only what someone thought to flag.

open as a page

How do scope and metadata filters on agent memory reads prevent cross-user leakage?

level: seniorimportance: must knowfreq 52%
basics
~20 s

Every memory read must be constrained by a scope key — user, tenant, session or project — derived from the authenticated session, applied inside the query before ranking. Semantic search has no notion of ownership, so nothing but an explicit filter keeps one user's memories out of another's context.

open as a page

Which AI-agent tool calls need human approval, and how do you tier the rest?

level: middleimportance: must knowfreq 72%
basics
~20 s

Gate by blast radius, not by tool name. Anything irreversible, externally visible, or above a value threshold — money movement, production schema changes, outbound messages, deletions — needs a human. Reads and cheaply reversible writes run automatically.

open as a page

Why does agent self-correction improve results with a test suite but often not without one?

level: middleimportance: must knowfreq 72%
basics
~20 s

Self-correction needs a signal the model did not produce. Failing tests, compiler errors and schema violations are outside evidence of a defect. Pure self-judgment re-samples the same model that wrote the output, so it rarely catches what it already missed.

open as a page

What stop conditions end an agent loop besides the model returning no tool call?

level: middleimportance: must knowfreq 62%
basics
~20 s

Agent loops end naturally when the model returns a final answer with no tool call, or calls an explicit done tool. They end forcibly on a max-iteration cap, a wall-clock deadline, a token or cost budget, or an unrecoverable error.

open as a page

How do you pause an AI agent for human approval that may take days?

level: seniorimportance: must knowfreq 58%
basics
~20 s

Persist the loop's state to a durable checkpoint, emit the approval request, and let the process exit. When the human answers — hours or days later, on a different machine — load the checkpoint, inject the decision, and resume from the interrupt point. Never block a thread.

open as a page

Why can asking an LLM "are you sure?" turn a correct answer into a wrong one?

level: juniorimportance: should knowfreq 56%
basics
~20 s

A challenge like "are you sure?" supplies doubt, not evidence. Models tend to go along with implied disagreement, so the second pass often swaps a correct but unusual detail — an exact date, an odd spelling — for a more ordinary-sounding one, lowering accuracy.

open as a page

Why is context isolation, not extra reasoning, the main gain from LLM subagents?

level: middleimportance: must knowfreq 70%
basics
~20 s

Spawning subagents adds no reasoning capacity — it is the same model. The gain is that each worker's noisy intermediate output stays in its own window, so the orchestrator attends to a small set of distilled findings instead of a bloated transcript full of dead ends.

open as a page

In a multi-agent LLM system, how does an orchestrator agent differ from a subagent?

level: juniorimportance: should knowfreq 55%
basics
~20 s

The orchestrator owns the goal: it splits the work, delegates, and assembles the answer. Each subagent runs its own loop in a separate context window, sees only its assigned task, and hands back a short summary rather than its full transcript.

open as a page

How would you choose between LangGraph, CrewAI and a vendor agent SDK?

level: middleimportance: should knowfreq 50%
basics
~20 s

Match the tool to the control you need. LangGraph gives explicit graph control flow with durable checkpointed state for production. CrewAI reaches a working prototype fastest with role-based crews. Vendor SDKs are the shortest path to that provider's newest capabilities.

open as a page

Why do multi-agent systems need a single-writer rule for each shared artifact?

level: seniorimportance: should knowfreq 45%
basics
~20 s

Isolated workers cannot see each other's decisions, so two of them editing the same artifact make incompatible implicit choices that the orchestrator cannot reconcile from summaries. Keep writes with one designated agent; let the rest return read-only findings or proposals.

open as a page

Multi-agent research systems burn ~15x the tokens of a chat — when does that pay?

level: principalimportance: should knowfreq 50%
basics
~20 s

It pays when the work is read-heavy, splits into independent parallel pieces, and the result can be checked — research, broad search, breadth-first exploration. It does not pay for latency-sensitive, high-volume or tightly sequential tasks where a single agent or plain workflow is adequate.

open as a page

What are the main failure modes of a tool-using LLM agent, and what fixes each?

level: middleimportance: must knowfreq 70%
basics
~20 s

Tool-using agents fail in recognizable categories: invented tool names, invalid arguments, the wrong real tool, non-progressing loops, drift from the original goal, and false claims of success. Each has a different fix, so classifying the failure is the first debugging step.

open as a page

In agent evals, what does pass^k measure that a single run per task hides?

level: middleimportance: must knowfreq 68%
basics
~20 s

pass^k is the share of tasks an agent solves on all k repeated attempts, so it measures reliability rather than best-case ability. A single run hides variance: roughly 90% per-attempt success clears all 8 attempts on only about 43% of tasks.

open as a page

Why score an AI agent's trajectory and not just its final answer?

level: middleimportance: must knowfreq 72%
basics
~20 s

Final-answer scoring cannot tell a clean four-step solution from a forty-step flail that stumbled onto the same output, and it credits answers reached by luck. Trajectory metrics score the path: tool choices, ordering, and wasted steps.

open as a page

How do you score tool selection and argument correctness in an agent trajectory?

level: middleimportance: must knowfreq 58%
basics
~20 s

Treat the reference trajectory's tool calls as ground truth. Precision is the share of the agent's calls that appear in the reference; recall is the share of reference calls the agent made. Then, on matched calls only, score arguments separately.

open as a page

An agent claims success after its tool returned HTTP 502 — what failure mode is this?

level: seniorimportance: must knowfreq 55%
basics
~20 s

A false success claim: the agent narrates completion over a failed side effect. It is dangerous because nothing crashed, so the run looks green while the work never landed. Only an independent post-condition check catches it.

open as a page