skip to content

What does AutoGen's UserProxyAgent do, and when should you use it?

level: juniorimportance: should knowfreq 42%

answer

  1. a human wearing an agent's interface
  2. no model client, no system prompt
  3. the default blocks on the console
  4. supply your own input function
  5. it does not run code any more

basics

~20 s

UserProxyAgent represents a human as a participant. It has no model client and no system prompt: when it is its turn, it calls its input_func to obtain a human reply and returns that as a message. By default it blocks on console input.

solid answer

~50 s

In `autogen-agentchat` 0.7.x, `UserProxyAgent` is the agent whose "model" is a person. It takes `name` and an optional `input_func`; when the runtime gives it a turn, it invokes that function with a prompt and returns whatever comes back as the agent's message. The default `input_func` is Python's blocking `input()`, which is fine in a notebook or a script and wrong everywhere else — inside a team it holds the entire run while it waits, and in a web service it would block the event loop. For anything real you supply an async `input_func` that resolves against your UI or queue, and you pass a `CancellationToken` so an abandoned conversation can be torn down instead of waiting forever. It is a common misconception, carried over from AutoGen's pre-0.4 API, that this agent executes code — in the current AgentChat design code execution belongs to a dedicated agent type, and the user proxy only relays human text.

code

python · 12 lines
python
import asyncio

from autogen_agentchat.agents import UserProxyAgent
from autogen_core import CancellationToken


async def ask_human(prompt: str, token: CancellationToken | None = None) -> str:
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(None, input, prompt)


user = UserProxyAgent(name="human", input_func=ask_human)

go deeper

for a junior

Be able to say plainly that it stands in for a person: it asks for input and returns whatever the human typed as its message, with no model behind it.

for a middle

Explain the input_func parameter, the blocking console default, and why an async implementation plus a cancellation token is required once the agent runs inside a real application.

for a senior

Discuss the operational shape of a suspended run — resource hold, timeouts, disconnect handling — and separate the turn-taking human from the narrower approval-gate design.

for a principal

Own the boundary question: which decisions genuinely need a human inside the loop versus at its edges, and what an unbounded human wait costs in throughput and in resources held by a parked run.

## The idea AutoGen models a conversation as agents taking turns. A human participating in that conversation needs to look like an agent to the runtime, and `UserProxyAgent` is that adapter. It holds no model client, has no system prompt and does no reasoning. Its entire job is: turn arrives → ask the human → return the human's words as this turn's message. ## Constructor surface The parameters that matter are `name`, an optional `description` (which is how other components see its role), and `input_func`. The input function can be synchronous — `Callable[[str], str]`, matching the built-in `input` — or asynchronous, taking the prompt and an optional cancellation token and awaiting a string. The default is the built-in `input()`, which is why the quickstart examples "just work" in a terminal. ## Why the default is a trap outside a notebook Blocking `input()` inside an async application blocks the event loop: nothing else in the process progresses while the prompt sits there. Inside a team, the consequences are worse than slow — the whole run is suspended on the human, so a user who closes the tab leaves a run parked indefinitely, holding whatever resources it holds. Two mitigations belong in any serious answer: 1. **Async input.** Supply an `input_func` that awaits your real input channel — a websocket message, a queue, a database row a UI writes. If you must use console input inside an async app, at least push it to a thread executor rather than calling `input()` on the loop. 2. **Cancellation.** Pass a `CancellationToken` into the run and cancel it on disconnect or timeout. Human-in-the-loop without a cancellation path is an unbounded wait by construction. ## What it is not Developers who learned AutoGen before the 0.4 redesign remember `UserProxyAgent` as a dual-purpose object: it collected human input *and* executed code blocks the assistant produced, with a `human_input_mode` setting choosing between always asking, never asking, and asking only at termination. That was the old `ConversableAgent`-based API. The current AgentChat design splits those responsibilities: the user proxy relays human text, and execution of generated code is a separate agent type with its own executor configuration. Claiming the user proxy runs code is one of the clearest signals that a candidate is describing a version of the framework that no longer exists. ## Where it fits in a design Use it when the human is genuinely a *participant* — approving a step, supplying a missing fact, choosing between options the agents surfaced. Do not use it as a general application input mechanism: if your UI already has a chat box, your service can simply call the team with the user's message as the task and never instantiate a user proxy at all. The proxy earns its place when the human must speak *in the middle of an agent run*, not at its boundaries. A useful design question follows from that: is the human a turn-taker or a gate? If they only ever approve or reject, an approval gate around the specific action is usually simpler and safer than a full conversational participant, because it scopes the wait to one decision. If they are genuinely conversing with the agents, the proxy is the right abstraction. ## Operational notes - Because it produces ordinary messages, everything the human types becomes part of the conversation history the models see. Whatever redaction or validation you would apply to user input before it reaches a model applies here too — there is no filtering layer inside the proxy. - The prompt string passed to `input_func` is what you render to the human. If your UI needs richer context than a prompt line, capture it from the surrounding run rather than expecting the proxy to supply it. - Timeouts belong in your `input_func`, not in hope. An implementation that awaits forever is indistinguishable from a hung system. ## Interview framing A good answer is short and lands three things: it is an adapter that makes a human look like an agent; the default blocking input is notebook-grade only; and it does not execute code in the current API.

  • What breaks if you leave the default input function in place inside a web service?
    The default is Python's blocking `input()`, which reads from the process's stdin and blocks the event loop while it waits — so a service either hangs or reads from a console nobody is watching. Supply an async `input_func` bound to your real input channel, and pair it with a `CancellationToken` and a timeout so a user who disconnects does not leave the run suspended indefinitely.
  • How was human input configured in AutoGen before the 0.4 redesign?
    The pre-0.4 API centred on `ConversableAgent`, where a `human_input_mode` setting chose between always prompting, never prompting, or prompting only around termination, and the user proxy also executed code blocks. The current AgentChat design replaces that with distinct agent classes — a model-backed assistant, a human-input proxy, and a separate code-executing agent — so the mode flag has no equivalent in current code.
  • Would you use a UserProxyAgent for simple approve-or-reject decisions?
    Often not. A proxy makes the human a conversational turn-taker, which suspends the whole run while it waits. If the human only ever approves or rejects one action, an explicit approval gate around that action is narrower, easier to time out, and easier to audit. Reach for the proxy when the human genuinely converses mid-run and their free-text answer becomes part of the context the models reason over.

saying these in an interview costs you the question

  • Says UserProxyAgent executes the code the assistant writes
  • Thinks it has its own model client or system prompt
  • Leaves the blocking default input function in an async service
  • Describes human_input_mode as current API rather than pre-0.4
  • Assumes a waiting human proxy does not stall the rest of the run

context