In Koog, when do you install EventHandler versus the OpenTelemetry feature on an agent?
answer
- Callbacks versus spans
- One reacts, the other reveals
- Handlers run on the execution path
- Content capture is a PII switch
- Install both, for different jobs
basics
~20 sEventHandler gives you in-process Kotlin callbacks on agent lifecycle points — tool calls, model calls, errors — for custom logic like metrics or audit records. The OpenTelemetry feature emits spans through the OTel SDK so runs appear in your existing tracing backend. They compose; install both.
solid answer
~50 sBoth are installed on the agent with `install(...)`, and they answer different questions. `install(EventHandler) { ... }` registers callbacks on lifecycle points — handlers such as `onBeforeAgentStarted`, `onBeforeLLMCall`, `onAfterLLMCall`, `onToolCall`, `onToolCallFailure`, `onAgentRunError` and `onAgentFinished`. That is where you run *your* code: increment a counter, write an audit row, enforce a spend budget, assert in a test. `install(OpenTelemetry) { ... }` instead wires the run into OpenTelemetry: you set service info and add span exporters, and Koog emits spans for the agent run and its model and tool calls, so an agent's work is correlated with the rest of your service's traces. One knob to know is span verbosity — with content capture enabled, spans carry prompt and completion text, which is enormously useful in debugging and a PII exposure in production. Choose EventHandler for behaviour, OpenTelemetry for visibility.
go deeper
Know that both are installed on the agent, that one gives you Kotlin callbacks on events like tool calls and errors, and that the other emits traces.
Be able to name concrete handlers and say what each feature is for: callbacks for metrics, audit and tests; spans for seeing where a run spent its time.
Show the production judgment — handlers run on the execution path so they must stay cheap, and content capture on spans is a privacy and cost decision, not a default.
Own the standard: what every agent emits, retention and access rules for spans containing prompts, and which signals gate a rollout or trip an alert.
## Two different questions When an agent misbehaves you ask two kinds of question. "What did it do, in order, and how long did each step take?" is a tracing question. "What should my system *do* when it does that?" is a control question. Koog answers them with two separate installable features, and confusing them leads to people writing metrics collection inside a tracing exporter or, worse, trying to reconstruct traces from log lines. ## EventHandler: callbacks you own `install(EventHandler) { ... }` attaches Kotlin lambdas to the agent's lifecycle. The available hooks span the whole run: the agent starting and finishing, errors escaping the run, model calls before and after, tool calls, tool failures and validation errors. Each handler is your code running in your process with access to your dependencies. That makes it the right home for anything that is a *decision* or a *record*, not a trace: - **Metrics.** Increment a counter per tool invocation, time model calls, record failure rates by tool name. - **Cost and budget.** Accumulate token usage per run and per tenant; alert or abort when a run exceeds a threshold. - **Audit.** Write an immutable record of which tools ran with which arguments for a regulated workflow. - **Operational alerting.** Page on a spike in tool-call failures, which is often the first symptom of a downstream outage. - **Tests.** Assert that a strategy called the expected tools in the expected order, without inspecting logs. Two cautions. Handlers run inside the agent's execution path, so anything slow or blocking in them slows the run and anything that throws can disrupt it — keep them cheap and defensive, and push heavy work onto a queue. And handler arguments frequently contain user content, so what you log from them is a privacy decision. ## OpenTelemetry: spans for the system you already run `install(OpenTelemetry) { ... }` connects the agent to the OpenTelemetry SDK. You configure identity for the emitted telemetry — service name and version — and add one or more span exporters that ship spans wherever your platform collects them. Koog then produces spans for the run and its constituent model and tool calls. The payoff is correlation. An agent is rarely the whole request: an HTTP handler received a call, hit a database, ran the agent, and the agent hit three tools that themselves called your services. If all of it is OTel, one trace shows the whole causal chain with real latencies, and "the agent is slow" resolves into "one tool waits four seconds on a downstream call" without guesswork. The knob that matters operationally is content verbosity. Enabling it puts prompt and completion content onto spans, which is exactly what you want while developing a strategy and exactly what you do not want streaming into a shared observability backend in production: prompts contain user data, spans get retained, and access to the tracing UI is usually broader than access to the production database. Treat verbose span content as a debug setting with a deliberate policy behind it, and remember span size costs money at volume. ## Why you install both They are complementary rather than alternatives. Traces tell you what happened and where the time went; event handlers let your system react. A realistic production agent installs OpenTelemetry so runs land in the same backend as everything else, and EventHandler so token spend is metered, tool failures are counted, and an audit trail exists. Neither replaces ordinary application logging, which remains the fallback when the collector itself is broken. ## Boundaries worth naming in an interview These features are Koog's *emission* surface. Which backend consumes the spans, how you evaluate answer quality, and how you score agent trajectories are separate concerns with their own tooling; the framework's job here is to expose the run faithfully and let you choose where it goes. Similarly, event handlers observe and react — they are not a mechanism for editing the run's control flow, which lives in the strategy graph itself. ## A practical triage order When an agent misbehaves: read the trace first to see the shape of the run and where latency and retries went. If the shape is right but the outcome is wrong, turn on content capture in a non-production environment and read the actual prompts and completions. If the problem is systemic rather than one run — spend creeping up, a tool failing five percent of the time — that is a metrics question, and metrics come from event handlers, not from staring at individual traces.
- What is the risk of doing heavy work inside an EventHandler callback?Handlers execute inside the agent's run, so a slow or blocking callback adds latency to every affected step, and a throwing one can disrupt the run. Keep them to cheap in-memory work — counters, timers, a small record — and hand anything expensive, like a database write or an HTTP call, to a queue or a separate coroutine scope.
- Why would you disable content capture on spans in production?Because prompts and completions contain user data, and spans are retained in a backend that usually more people can read than can read your database. Full content also inflates span size and cost at volume. Keep verbose content for development and incident-scoped debugging, and rely on structure — step names, durations, tool names, token counts — in normal production traces.
- An agent looks slow in production. Which feature do you reach for first, and why?OpenTelemetry. The trace shows the run's shape — how many model calls, how many tool calls, how long each took — and because the agent's spans sit in the same trace as the surrounding request and downstream service calls, latency attributes itself. Event-handler metrics come next, to tell you whether this run was typical or an outlier.
saying these in an interview costs you the question
- Treats event handlers as a way to change the run's control flow
- Leaves prompt and completion capture on spans enabled in production
- Thinks one feature replaces the other
- Does blocking I/O inside a lifecycle callback
- Tries to reconstruct traces from log lines instead of emitting spans