How do you configure AutoGen's OpenAIChatCompletionClient for an unknown model?
answer
- clients live in the extension package
- the client knows only listed models
- capabilities must be declared, not guessed
- a wrong declaration fails silently
- Azure addresses a deployment, not a model
basics
~20 sPass an explicit model_info dictionary alongside model and base_url. The client keeps a capability table for models it recognizes; for anything outside it — a self-hosted or newly released model behind an OpenAI-compatible endpoint — it cannot infer capabilities and raises unless you declare them.
solid answer
~40 sIn `autogen-agentchat` 0.7.x, model clients live in `autogen_ext.models.*`: `OpenAIChatCompletionClient` and `AzureOpenAIChatCompletionClient` under `autogen_ext.models.openai`, `AnthropicChatCompletionClient` under `autogen_ext.models.anthropic`. The OpenAI client ships a lookup table of known model names mapping to capabilities. Point it at a local or third-party OpenAI-compatible server via `base_url` with a model name it does not know, and it has no way to tell whether that model supports tools, images or JSON output — so you must supply `model_info` with `vision`, `function_calling`, `json_output`, `family` and, in recent versions, `structured_output`. Getting these wrong is worse than omitting them: declaring `function_calling=True` for a model without tool support produces confusing empty or malformed tool calls rather than a clean error. The Azure client needs its own wiring — `azure_deployment`, `api_version`, `azure_endpoint`, plus an API key or `azure_ad_token_provider`.
code
python · 14 linesfrom autogen_ext.models.openai import OpenAIChatCompletionClient
client = OpenAIChatCompletionClient(
model="my-local-model",
base_url="http://localhost:11434/v1",
api_key="placeholder",
model_info={
"vision": False,
"function_calling": True,
"json_output": False,
"family": "unknown",
"structured_output": False,
},
)go deeper
Know that the agent takes a model_client and that clients come from the extension package, and that a model the client does not recognize needs its capabilities spelled out.
Name the model_info fields and explain why they exist: AutoGen must decide before sending whether tools, images or JSON output are allowed against that endpoint.
Emphasise the silent-failure path of a wrong capability declaration, and the lifecycle split — clients shared process-wide and closed on shutdown, agents constructed per session.
Own the boundary contract: capability declarations are configuration that can be wrong without failing, so they need verification in CI against each endpoint, the same as any other integration assumption.
## What a model client is in AutoGen Agents in AgentChat do not talk to a provider directly. They hold a **model client** — an object with a uniform chat-completion interface — and swapping the client is how you swap models under an unchanged agent. Clients live in the extension package: `autogen_ext.models.openai` provides `OpenAIChatCompletionClient` and `AzureOpenAIChatCompletionClient`; `autogen_ext.models.anthropic` provides `AnthropicChatCompletionClient`. Because the client is a constructor argument, giving two agents different models is just passing two different client instances, and giving them the same model is passing the same instance — which is normal and desirable, since a client is a connection pool, not per-agent state. ## Why model_info exists AutoGen has to make decisions before it ever sends a request: can this model be given tool schemas at all? Can it accept image content? Can it be asked for JSON? What message conventions does its family use? For a recognized model name the client answers those from a built-in table. But the OpenAI client is deliberately usable against *any* OpenAI-compatible endpoint — a local inference server, a gateway, a vendor clone — through `base_url`. In that case the model name is meaningless to the table, and rather than guessing, the client requires you to declare capabilities via `model_info`. The fields: - `vision` — whether image content may be sent. - `function_calling` — whether tool schemas may be sent and tool calls expected. This is the one that matters most for agents: an `AssistantAgent` with tools is useless against a model where this is false. - `json_output` — whether a JSON-shaped response can be requested. - `family` — which model family this belongs to, used for family-specific handling; `"unknown"` is a legitimate value. - `structured_output` — present in recent versions, covering schema-constrained output. ## The dangerous failure mode A missing `model_info` gives you a loud error at construction time, which is the good case. A *wrong* `model_info` gives you a quiet, expensive one. Declare `function_calling: True` for a model that cannot do it and the agent will happily send tool schemas; the model responds with prose describing what it would call, no tool executes, and the agent returns a plausible-looking answer built on nothing. Declare `vision: True` and image content is sent to a text-only endpoint. Treat `model_info` as a contract you have actually verified against the endpoint, not a block copied from a tutorial. ## Azure wiring `AzureOpenAIChatCompletionClient` is a separate class because Azure's addressing is different: you name a **deployment** (`azure_deployment`) as well as the underlying `model`, and you must supply `api_version` and `azure_endpoint`. Authentication is either `api_key` or, preferably in production, `azure_ad_token_provider` for token-based auth so no static secret sits in configuration. Passing an OpenAI-style client with a rewritten base URL at Azure is a common mistake that produces confusing 404s on the deployment path. ## Lifecycle Clients are async and hold an underlying HTTP client, so close them: `await client.close()` when the application shuts down, or manage them with an async context manager. In a long-running service, build the client once at startup and share it across agents and sessions; building one per request leaks connections and adds handshake latency to every call. This is the inverse of the agent lifetime rule — agents are per-session state, clients are process-wide infrastructure. ## Streaming and other knobs The OpenAI client also takes ordinary request settings — `temperature`, `max_tokens`, `timeout`, `max_retries`, `parallel_tool_calls`. Note that token-level streaming for an agent is switched on at the *agent* (`model_client_stream=True` on `AssistantAgent`), which makes the agent emit streaming chunk events as the model produces them; the client must of course support streaming for that to work. Leaving it off means you still get the reply, just not incrementally. ## Interview framing The question behind the question is whether you understand that AutoGen cannot introspect an arbitrary endpoint. Capability declaration is a manual contract at the boundary, and the cost of getting it wrong lands as bad agent behaviour rather than a stack trace — which is exactly the class of bug that survives review and reaches production.
- What actually goes wrong if you declare function_calling=True for a model that cannot call tools?Nothing fails loudly. AutoGen sends tool schemas, the model answers in prose describing what it would do, no tool executes, and the agent returns a fluent answer with no grounding behind it. You discover it from missing tool-execution events or from wrong results, not from an exception. That asymmetry is why the capability declaration should be verified against the endpoint rather than copied.
- Should each agent get its own model client instance?No — share one client per model across agents. The client wraps an HTTP connection pool and carries no per-agent state, so sharing is both correct and cheaper; per-agent instances multiply connections and handshake latency. Distinct instances are for distinct models or distinct endpoints, which is exactly how you give a cheap model to a routing agent and an expensive one to a reasoning agent.
- Why is there a separate Azure client class rather than a base_url override?Azure's request addressing differs: you target a named deployment via azure_deployment, alongside azure_endpoint and an explicit api_version, and authentication can be an Entra token provider rather than a static key. The dedicated AzureOpenAIChatCompletionClient encodes that shape. Pointing the plain OpenAI client at an Azure host with a rewritten base URL typically yields 404s on the deployment path.
saying these in an interview costs you the question
- Assumes AutoGen can probe an endpoint for its capabilities
- Copies a model_info block without checking it matches the model
- Creates a new model client per request in a service
- Points the plain OpenAI client at Azure by rewriting base_url
- Never closes the client, then blames the framework for leaked connections