When do you build directly on autogen-core instead of autogen-agentchat?
answer
- Policy layer versus mechanism layer
- Does conversation describe the problem?
- Open reactor sets point downward
- You inherit termination and observability
- Mixing layers is the usual answer
basics
~20 sDrop to autogen-core when the conversation-and-team model stops matching the problem: event-driven topologies, agents deployed as separate services, non-chat message types, or bespoke control flow. The cost is that you write the message protocol, routing and stopping rules yourself.
solid answer
~50 sAgentChat is a policy layer over core's mechanism, so the question is whether its policy fits. Stay on `autogen-agentchat` when the work is genuinely conversational — a bounded set of agents taking turns, a message history the model consumes, a stopping rule expressible as a termination condition. Drop to `autogen-core` when you need something its abstractions do not model: publish/subscribe topologies where the set of reactors changes independently of the producer, typed non-chat messages between components, agents that must run in separate processes or scale independently, or control flow that is not a turn-taking conversation at all. The tradeoff is explicit: core gives you addressing, subscriptions, typed dispatch and a runtime, and nothing else — you own the message schema, the routing, the termination logic and the observability that AgentChat provided. A common middle path is a core-level topology that delegates bounded sub-problems to AgentChat teams, since teams already run on the core runtime.
go deeper
Know that autogen-agentchat is the layer you normally write against, and that autogen-core underneath it is the lower-level runtime you would only reach for in unusual designs.
Explain what AgentChat's team model assumes — known participants, a shared message history, a termination condition — and name a concrete case such as pub/sub fan-out where those assumptions stop holding.
Demonstrate the accounting: dropping to core means owning message schemas, loop bounds, aggregation and tracing, so justify the move with a specific constraint and describe the mixed topology that limits the hand-written surface.
Own the whole decision — deployment boundaries, team ownership and release cadence, the onboarding cost of a bespoke topology, and how to keep prompts, tools and domain logic outside framework classes so the layer choice stays reversible.
## The two layers are not alternatives in the usual sense `autogen-agentchat` is implemented on `autogen-core`'s runtime. Choosing core is not switching frameworks; it is declining a set of defaults. That framing matters, because it means the decision is reversible in one direction (a core application can host an AgentChat team) and it means "which is better" is the wrong question. The right question is: does the conversational model describe my problem, or am I fighting it? ## What AgentChat's policy assumes The team abstraction encodes a specific shape: a known set of participants, a message history that grows as they speak, a selection rule for who goes next, and a stopping condition evaluated over that history. When your problem has that shape, the layer is a large gift — the message protocol, the turn logic, the streaming surface and the stopping rules already exist and are debugged. It also assumes messages are conversational. The history is something a model reads. If your agents mostly exchange structured records that no model will ever see as prose, you are paying for a chat transcript you do not want. ## Signals that you have outgrown it - **The reactor set should be open.** You want to add an auditor, a metrics collector or a policy checker without editing the producer. That is publish/subscribe, and core's topics and subscriptions model it directly. - **Messages are not conversation.** Typed events with schemas, versioned across services, consumed by handlers rather than read by a model. - **Deployment boundaries are real.** Different agents need different hardware, different release cadences, different blast radii, or independent horizontal scaling. Core's runtime abstraction is what the distributed transport plugs into. - **Control flow is not turn-taking.** Fan-out with a collector, long-lived stateful actors keyed by tenant or session, event-sourced pipelines, workflows driven by external triggers rather than by the previous speaker. - **Lifecycle is not a run.** Agents that live indefinitely and react, rather than participating in a bounded task with a start and an end. If none of these is true, staying on AgentChat is the right call and saying so confidently is a better interview answer than reaching for the lower layer to sound sophisticated. ## What you take on The honest accounting is what separates a principal answer from a plausible one: 1. **Message design becomes yours.** You define every type, its schema, and its evolution story — which becomes a deployment concern the moment agents run in separate processes. 2. **Termination becomes yours.** Core has no notion of a run that ends. Loop bounds, budget caps and deadlines are logic you write, and getting them wrong costs tokens. 3. **Orchestration becomes yours.** Who acts next, how results are aggregated, how a fan-out joins — all of it is code, not configuration. 4. **Observability becomes yours.** Publishing is fire-and-forget and unhandled message types only log, so a core system without deliberate tracing is a system where nothing appears to happen and nothing raises. 5. **Onboarding cost.** A bespoke actor topology is harder for the next engineer than five lines of team construction. ## The composition escape hatch Because the layers stack, the pragmatic architecture is often mixed: core owns the topology — the topics, the long-lived stateful agents, the deployment boundaries — and hands bounded, genuinely conversational sub-problems to AgentChat teams, awaiting their results like any other unit of work. This keeps the hand-written surface small and confined to the parts that actually needed it, and it lets you migrate incrementally: start entirely on AgentChat, and pull only the piece that outgrew it down a layer. ## The framework-commitment question underneath A principal-level answer closes the loop on lock-in. Building on core means writing more of your own orchestration, which paradoxically leaves less of your logic shaped by any one vendor's conversation abstraction — your handlers become ordinary async functions over your own message types. Given that AutoGen's line has been folded into a broader agent framework effort, that portability has real value. The corollary is a design rule that applies at either layer: keep prompts, tool implementations, domain logic and evaluation cases outside framework classes, so whichever layer you chose is a shell you can replace rather than the substance of the system.
- What is the strongest single signal that a team should stay on autogen-agentchat?That the problem is genuinely a bounded conversation: a known set of participants, a transcript the model reads, and a stopping rule you can state in a sentence. When that holds, the team abstraction hands you turn selection, message protocol, streaming and termination already debugged, and rebuilding them on core buys nothing but maintenance. Reaching for the lower layer without a concrete constraint is complexity you will pay for on every onboarding.
- Can you adopt core incrementally rather than rewriting?Yes, and that is usually the right sequencing. AgentChat teams run on the core runtime, so a core-level application can own the topology — topics, long-lived stateful agents, deployment boundaries — and delegate a bounded conversational sub-problem to a team, awaiting its result. You pull down one piece at a time, keeping the hand-written surface confined to whatever actually outgrew the conversational model.
- What is the first thing you build yourself when you drop to core, and what does skipping it cost?Bounds and observability. Core has no concept of a run that ends, so loop caps, budget limits and deadlines are your code, and without them a reactive topology can spin indefinitely against a paid API. Publishing is fire-and-forget and unhandled message types only log, so tracing every publish, delivery and handler outcome is what makes the system debuggable at all — retrofitting it after an incident is far more expensive.
- How does the choice of layer affect lock-in?Building on core means your handlers are ordinary async functions over message types you defined, so less of the system is shaped by one vendor's conversation abstraction — which matters given AutoGen's line has been folded into a broader effort. But the real protection is independent of layer: keep prompts, tool implementations, domain logic and evaluation cases outside framework classes, so whichever layer you chose stays a replaceable shell.
saying these in an interview costs you the question
- Choosing core for sophistication rather than a constraint
- Assuming core also provides termination and turn-taking
- Believing you must pick one layer for the whole system
- Treating the two layers as competing frameworks
- Ignoring the observability you inherit when you drop down