skip to content

In Koog, what does a PromptExecutor do, and how does MultiLLMPromptExecutor differ?

level: middleimportance: must knowfreq 72%

answer

  1. One entry point for every model call
  2. Client per vendor, executor above it
  3. Routing key is the model's provider
  4. Single client versus provider map
  5. Stateless — history rides in the Prompt

basics

~20 s

A PromptExecutor is Koog's single suspending entry point for running a Prompt against an LLModel, hiding provider HTTP details behind an LLMClient. SingleLLMPromptExecutor wraps one client; MultiLLMPromptExecutor holds several and routes each call by the model's provider.

solid answer

~50 s

Koog splits three things: a `Prompt` (typed messages), an `LLModel` (which model, with declared capabilities), and an `LLMClient` (the provider adapter that speaks OpenAI's, Anthropic's, Google's, Ollama's or OpenRouter's wire format). `PromptExecutor` is the interface that joins them — you call `execute(prompt, model)` and suspend until response messages come back, plus a streaming variant returning a Flow. `SingleLLMPromptExecutor` is bound to exactly one client, so passing a model from another provider is an error, not a silent fallback. `MultiLLMPromptExecutor` is constructed with a map of `LLMProvider` to client and dispatches per call on `model.provider`, which is how you run cheap local Ollama models next to a hosted frontier model in the same process. Helpers like `simpleOpenAIExecutor(apiKey)` build the common single-client case. Agents take the executor, so it is your seam for fakes, retries and cost accounting.

code

kotlin · 16 lines
kotlin
suspend fun main() {
    val executor = MultiLLMPromptExecutor(
        LLMProvider.OpenAI to OpenAILLMClient(System.getenv("OPENAI_API_KEY")),
        LLMProvider.Anthropic to AnthropicLLMClient(System.getenv("ANTHROPIC_API_KEY")),
    )

    val answer = executor.execute(
        prompt = prompt("triage") {
            system("Answer in one sentence.")
            user("Why did order 4711 fail to ship?")
        },
        model = OpenAIModels.Chat.GPT4o,
    )

    println(answer.content)
}

go deeper

for a junior

Know that you build a Prompt, pick a model, and call an executor to get a response, and that the API key lives on the client you construct, not in the prompt.

for a middle

Be ready to explain the split between Prompt, LLModel and LLMClient, and how a multi-provider executor dispatches on the model's provider so per-call model choice works.

for a senior

Show that the executor is the choke point for retries, caching, cost accounting and test doubles, and that capability differences mean provider swaps still need re-evaluation, not just a config change.

for a principal

Own the policy: which tiers of models each workload uses, how you avoid vendor-specific behaviour leaking into strategies, and what your fallback story is when a provider degrades mid-incident.

## The three pieces Koog keeps separate Koog deliberately splits *what you say* from *which model* from *how you reach it*. A `Prompt` is an immutable, typed list of messages. An `LLModel` is a value object naming a provider, a model id, and the capabilities that model declares. An `LLMClient` is the provider-specific adapter that knows one vendor's HTTP API. `PromptExecutor` is the interface that ties all three together, and it is the only one of them that most of your code touches. You hand an executor a prompt, a model, and optionally the descriptors of tools the model may call; it suspends and returns the model's response messages. There is also a streaming form that returns a Flow so you can push tokens to a client as they arrive, and moderation support on providers that offer it. ## The client layer underneath Concrete clients include `OpenAILLMClient`, `AnthropicLLMClient`, `GoogleLLMClient`, `OllamaClient` and `OpenRouterLLMClient`. Each one owns credentials, base URL and HTTP settings, converts Koog's typed messages into that vendor's request shape, and converts the reply back into Koog's message types. That conversion is where provider differences are absorbed: system-message handling, tool-call encoding, and structured-output mechanics all differ between vendors, and the client is what makes them look the same upstream. ## Single versus multi provider `SingleLLMPromptExecutor` wraps exactly one client. It is the right default: one provider, one key, no routing logic. If you pass it a model belonging to a different provider, the call fails — Koog does not quietly substitute anything, and that strictness is deliberate, because a silent substitution would change cost, latency and capability behind your back. `MultiLLMPromptExecutor` is built from provider-to-client pairs and selects the client for each call from the model you passed. This is what makes per-call model choice real: a cheap small model for classification and routing, a large model for the final synthesis step, a local Ollama model in tests, all through one injected dependency. Convenience builders such as `simpleOpenAIExecutor`, `simpleAnthropicExecutor` and `simpleOllamaAIExecutor` exist for the one-provider case so demos and tests stay short. ## Capabilities are part of the model, not the executor Because `LLModel` carries declared capabilities — tool calling, structured/JSON schema output, vision and so on — Koog can reject an incompatible combination early rather than letting a tool-using agent silently degrade into a chat-only model that never calls anything. Model constants are provided per provider (for example `OpenAIModels` and `AnthropicModels` objects) so you are not typing model id strings by hand, though you can construct an `LLModel` yourself for a model Koog does not ship a constant for. ## What the abstraction buys Provider portability is the obvious win: swapping vendors is a change to executor construction, not to your strategy, tools or prompts. Beyond that, a single choke point for every model call is exactly where you want retry, caching, rate limiting, token accounting and tracing to live. Testing benefits most of all — a fake `PromptExecutor` that returns canned responses lets you test an agent's control flow with no network and no key. ## What it costs An abstraction over five vendors is necessarily close to a lowest common denominator: a provider-specific knob Koog does not model is not reachable through it, and you may need to configure it on the client or drop to the client directly. Capability drift is the practical failure mode — the same prompt against two providers can produce different tool-calling behaviour or different adherence to a JSON schema, so "we can swap providers" is true at the API level and only partly true at the behaviour level; you still re-evaluate. Also remember the executor is stateless. It does not accumulate conversation history for you: history lives in the `Prompt` you build and pass. Anything long-lived — sessions, memory, checkpoints — is a separate concern layered on top. ## Where it sits in an agent An agent is constructed with an executor and a model, and every model-calling step in it goes through that executor. So in a production service the executor is usually a single injected singleton, configured once with keys and timeouts, wrapped with whatever cross-cutting behaviour you need, and shared by every agent and every one-shot prompt call in the process.

  • You need retries with backoff on 429s across every model call. Where do you put that in Koog?
    At the executor or client layer, not inside strategies or tools. Because every model call funnels through one `PromptExecutor` instance, wrapping or decorating it (or the underlying `LLMClient`) applies the policy uniformly and keeps agent code free of transport concerns. The same seam is where token accounting and caching belong, and it is why the executor is normally a single injected singleton.
  • How would you unit-test an agent's control flow without calling a real provider?
    Inject a fake `PromptExecutor` that returns scripted `Message.Response` values, including a scripted tool call, and assert which nodes and tools ran. Since the executor is an interface and agents depend on it rather than on a concrete client, no HTTP, no API key and no network flakiness enter the test. Reserve real-provider runs for a small evaluation suite.
  • What breaks if you point the same prompt at a model that lacks a declared tool-calling capability?
    Koog surfaces the mismatch rather than degrading silently: capabilities are declared on `LLModel`, so a tool-using call against a model without tool support is rejected instead of returning plain prose that your agent then fails to parse. That early failure is the intended behaviour — the alternative is an agent that appears to run but never invokes a tool.

saying these in an interview costs you the question

  • Thinks the executor falls back to another provider automatically
  • Says the executor stores conversation history between calls
  • Assumes swapping providers needs no re-evaluation of behaviour
  • Confuses the LLMClient (transport) with the PromptExecutor (entry point)
  • Believes provider choice is fixed at startup rather than per call

context