How do you wire OpenTelemetry tracing into an AutoGen team, and what do the spans show?
answer
- instrument the runtime, not the team
- a provider passed into a constructor
- pass that runtime back to the team
- two logger names for the lighter path
- structure and timing, not content
basics
~20 sBuild an OpenTelemetry TracerProvider with an exporter, pass it as tracer_provider to a SingleThreadedAgentRuntime, and hand that runtime to the team constructor. The runtime then emits spans for message dispatch and agent processing, so one multi-agent run becomes a single nested trace.
solid answer
~40 sAutoGen's tracing hangs off the runtime, not the team. You configure an OpenTelemetry `TracerProvider` with whatever exporter you use, construct `SingleThreadedAgentRuntime(tracer_provider=tracer_provider)`, and pass that runtime into the team via its `runtime` argument. The runtime instruments message send/publish and agent handling, so a run appears as one trace with nested spans per agent turn and tool invocation rather than as a flat pile of unrelated LLM calls. That nesting is the point: in a group chat the interesting question is which agent triggered which call, and a trace answers it where a log line cannot. Alongside spans, AutoGen emits structured Python logging under two named loggers exported from `autogen_core` — `TRACE_LOGGER_NAME` for human-readable developer tracing and `EVENT_LOGGER_NAME` for structured events including LLM call records. Which backend you ship spans to is a separate platform decision.
code
python · 17 linesfrom autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_core import SingleThreadedAgentRuntime
from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
def build_traced_team(agents: list[AssistantAgent]) -> RoundRobinGroupChat:
provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(provider)
runtime = SingleThreadedAgentRuntime(tracer_provider=provider)
# Without runtime=, the team builds its own untraced runtime.
return RoundRobinGroupChat(agents, runtime=runtime)go deeper
Know that AutoGen can emit OpenTelemetry spans and standard Python log events, and that the simplest visibility is attaching handlers to the logger names the core package exports.
Explain that instrumentation lives on the runtime: you pass a tracer_provider to SingleThreadedAgentRuntime and that runtime to the team, after which a run becomes one nested trace instead of unrelated calls.
Show that you use traces for structure and latency and the message stream for token attribution, that you know the classic no-spans cause is a team building its own runtime, and that you set sampling for chatty teams.
Own the policy: what payload is captured versus redacted, how trace retention interacts with data-protection commitments, and how tracing survives when agents run across processes rather than in one runtime.
## Why a multi-agent run needs traces, not logs A single-agent call is easy to reason about from logs: one prompt, one completion. A team run is a tree — the orchestrator picks a speaker, that agent calls a model, the model asks for a tool, the tool returns, the agent speaks, control moves on. Flat logs interleave those and lose the parent/child relationship, which is precisely the relationship you need when a run cost far more than expected. Spans preserve it. ## The wiring Three steps, and the ordering matters: 1. **Build a provider.** Standard OpenTelemetry SDK: a `TracerProvider`, a span processor, an exporter (OTLP to whatever collector you run). 2. **Give it to the runtime.** `SingleThreadedAgentRuntime(tracer_provider=tracer_provider)` — the runtime is the component that actually dispatches messages between agents, so it is the natural instrumentation point. 3. **Give the runtime to the team.** AgentChat team constructors accept a `runtime` argument. If you do not pass one, the team creates its own untraced runtime — which is why "I set up OpenTelemetry and see nothing" is almost always "the team is not using my runtime". Because the instrumentation lives in `autogen-core`, the same wiring covers both the Core API and AgentChat teams built on top of it. ## What you get Spans cover the runtime's work: sending a message to an agent, publishing to a topic, an agent processing a message. Nesting means one run is one trace, with a span per agent turn and children for the calls made inside it. The practical payoff is attribution — you can see that of the eleven model calls in a run, seven came from the critic looping, and you can see the latency each contributed. What spans do not do is replace the message stream. The stream carries content and token usage per message; traces carry structure and timing. Mature setups use both: stream items become the transcript users and support see, spans become the performance and cost picture. ## Logging as the lighter path Not every deployment wants a collector. `autogen_core` exports two logger names for the standard `logging` module: - `TRACE_LOGGER_NAME` — developer-oriented tracing output, useful at DEBUG in local runs. - `EVENT_LOGGER_NAME` — structured events, including model-call records that carry prompt and usage details. Attaching handlers to those names gets you a lot of visibility for one function call, and it is the right first move before you commit to a tracing backend. The cost is that structure is lost: you get lines, not a tree. ## Operational cautions **Payload capture is a privacy decision.** Model-call events and some span attributes can carry prompt text. In a system where users paste customer data into agents, that text lands in your observability backend under a different retention policy than your database. Decide deliberately what is captured, and redact at the exporter if needed. **Overhead is real but small.** Span creation per message dispatch is cheap next to a model call. The cost that bites is export volume — a chatty team generates many spans per run, and sampling policy should be set with that ratio in mind. **Distributed runtimes.** The instrumentation point being the runtime is what makes trace context propagate when agents are not in one process. That is the argument for tracing over ad-hoc logging: you cannot reconstruct a cross-process conversation from local log files. ## Where the boundary sits Wiring AutoGen to emit spans is a framework question. Which tracing or evaluation product consumes them, how you score agent outputs, and what dashboards you build on top are platform choices that live outside the framework's surface — and interviewers usually expect you to name that boundary rather than blur it.
- You configured a TracerProvider but no AutoGen spans appear. What is the first thing to check?Whether the team is actually using your runtime. Team constructors accept a `runtime` argument and create their own untraced runtime when you omit it, so the provider you configured is never consulted. Confirm you built `SingleThreadedAgentRuntime(tracer_provider=...)` and passed that same object into the team.
- Traces or the message stream — which do you use for cost attribution?The stream, primarily. Token usage is attached per message as `models_usage`, so per-agent cost is a sum over the messages. Traces give you structure and latency: which agent's turn spawned which calls and how long each took. Real setups use both, because neither answers the other's question.
- What is the privacy risk of turning on full tracing for an agent system?Prompt and completion text can end up in span attributes and structured model-call events, which means user-supplied content is copied into an observability backend with its own retention and access rules. Decide explicitly what payload is captured, redact at the exporter where needed, and treat trace retention as part of your data policy.
saying these in an interview costs you the question
- Expecting spans without passing the runtime into the team
- Thinking tracing is configured on the agent or model client
- Believing spans carry token usage so the stream is unnecessary
- Ignoring that prompt text may be exported to the tracing backend
- Assuming a tracing product is required to get any visibility