What does Koog's Spring Boot starter auto-configure, and how do you use it in a service?
answer
- Config keys under ai.koog
- One bean: the executor
- Providers configured decide single or multi
- Strategies and state stay yours
- Suspend, do not block request threads
basics
~20 sKoog's Spring Boot starter reads provider settings under ai.koog in your application config, builds an LLM client for each configured provider, and exposes a ready PromptExecutor bean you inject. It wires model access only — strategies, tools and session state stay yours.
solid answer
~40 sAdd the `koog-spring-boot-starter` dependency and put provider credentials under `ai.koog` in `application.yaml` — for example `ai.koog.openai.api-key` or `ai.koog.anthropic.api-key`. Auto-configuration then constructs a client per configured provider and publishes a `PromptExecutor` bean: with one provider configured you get a single-client executor, with several a multi-provider executor that routes by the model you pass. You inject that bean into a `@Service` and either call `execute(prompt, model)` for a one-shot prompt or pass it to an agent you build yourself. Know the boundary: the starter does not define your strategy, your tools, or where conversation state lives, and the executor is a stateless singleton, so history is yours to carry. Koog's API is suspending, so call it from a suspending controller or service method rather than blocking a servlet thread on `runBlocking`.
code
yaml · 6 linesai:
koog:
openai:
api-key: ${OPENAI_API_KEY}
anthropic:
api-key: ${ANTHROPIC_API_KEY}go deeper
Know the shape: add the starter dependency, put provider keys under ai.koog in application config, and inject the PromptExecutor bean into your service.
Explain what auto-configuration actually creates — a client per configured provider and one executor bean — and that configuring several providers yields provider-routed dispatch.
Show the boundary and the discipline: strategies, tools and conversation state remain yours, calls are suspending, and long runs belong off the request path with timeouts and a failure policy at the executor seam.
Own the platform decisions — model tiering and spend caps applied once at the executor, secret handling and per-environment profiles, and how agent workloads are isolated from ordinary request traffic.
## What the starter actually does The starter is deliberately small. It performs the one piece of wiring that is pure boilerplate in every Spring service: reading provider credentials from configuration, constructing the right `LLMClient` for each provider you configured, and publishing a `PromptExecutor` bean into the context. Provider settings live under the `ai.koog` prefix in `application.yaml` or `application.properties`, one block per provider, so they participate in Spring's normal configuration story — profiles, environment variables, config server, secret injection. The bean it publishes adapts to what you configured. Configure a single provider and you get an executor bound to that one client. Configure several and you get a multi-provider executor that dispatches each call on the provider of the `LLModel` you pass, which is what makes per-call model tiering possible from a single injected dependency. ## Using it From there it is ordinary Spring. Inject `PromptExecutor` into a service by constructor. For a one-shot prompt, build a `Prompt` and call `execute` with the model you want. For an agent, hand the executor to the agent you construct, alongside your own strategy and tool registry. The pattern that scales is to keep a thin service layer between your controllers and Koog: one Kotlin class per capability that owns its prompt, its model choice, and its result parsing. Controllers then depend on your domain-shaped API rather than on prompt strings, which keeps prompts testable and swappable and keeps model-selection policy in one place. ## What it deliberately does not do It is a *model access* starter, not an agent framework in a box. It does not invent a strategy, register tools, decide which model each use case should use, manage conversation history, or give you session storage. Those are application decisions, and the starter leaves them alone on purpose — the alternative would be a set of opinionated beans you fight rather than use. That boundary matters for interview answers, because the follow-up is always about *state*. The executor is a stateless singleton shared by every request. Multi-turn conversations therefore need somewhere to live: a database keyed by conversation id, a cache, or Koog's own persistence and memory features installed on the agents you build. Nothing about injecting a bean solves that for you. ## Coroutines and the servlet model Koog's API is suspending from top to bottom. In a Kotlin Spring application, the clean route is a suspending controller handler or a suspending service method, which Spring supports for Kotlin handlers. The anti-pattern to name explicitly is wrapping every call in `runBlocking` on a servlet thread: a model call can take tens of seconds, an agent run can take minutes, and blocking a bounded request-thread pool on that turns a modest burst into thread starvation for your whole application, including endpoints that have nothing to do with the agent. If you must bridge from blocking code, do it onto a dedicated bounded dispatcher, not the request thread, and put a timeout on it. Related: long agent runs do not belong inside a synchronous request at all. The usual shape is to accept the request, start the run on a background scope with its own supervision and timeout, return an identifier, and let the client poll or subscribe over SSE or WebSocket for progress. ## Configuration and operational concerns Because keys are ordinary Spring configuration, the usual rules apply: no keys in the repository, inject from the environment or a secret manager, and use profiles so local development can point at a locally hosted model while production uses a hosted provider. Add timeouts consciously — a default HTTP timeout that is generous for a chat completion may still be far shorter than a long structured generation, and an unbounded one is worse. Also plan for provider failure. The starter gives you the seam — a single executor bean — but the retry, circuit-breaking, budget and fallback policy are yours to add around it. Do it once, at that seam, rather than in each service. ## When the starter is not the right tool If your application is not Spring, the same wiring is a few lines by hand: construct the clients, construct the executor, register it in whatever container you use. The starter saves boilerplate and standardises configuration keys; it does not unlock capability. Knowing that keeps the answer honest — the interesting engineering in embedding Koog into a Spring service is coroutine discipline, state ownership and failure policy, not the dependency line.
- You configure both an OpenAI and an Anthropic key. What bean do you get?A multi-provider executor. Auto-configuration builds a client for each configured provider and exposes a single `PromptExecutor` that dispatches each call on the provider of the `LLModel` you pass. That is what lets one injected dependency serve a cheap model for classification and a large model for synthesis, with the choice made per call rather than at startup.
- Where does conversation history live in a Spring service using this starter?Wherever you put it. The executor is a stateless singleton and holds nothing between calls; history rides in the `Prompt` you construct. Typically you store messages keyed by conversation id in your database or cache and rebuild the prompt per turn, or install Koog's persistence and memory features on the agents you build. The starter solves configuration, not state.
- Why is runBlocking around a Koog call in a controller a bad idea?Because model calls take seconds and agent runs can take minutes, and blocking a bounded request-thread pool on them starves unrelated endpoints under modest load. Use a suspending handler or service method instead. If you truly must bridge from blocking code, do it on a dedicated bounded dispatcher with an explicit timeout, and move genuinely long runs off the request path entirely.
- How would you add a spend cap across every model call in the service?At the executor seam, since every call passes through the single injected bean — wrap or decorate it, or install an event handler on your agents that accumulates token usage per tenant and rejects further calls past a threshold. Doing it once there beats scattering checks across services, and it keeps the policy auditable and testable in isolation.
saying these in an interview costs you the question
- Expects the starter to build a full agent with tools
- Blocks servlet threads with runBlocking on agent runs
- Assumes the executor bean remembers conversation history
- Hardcodes provider keys instead of using Spring configuration
- Runs multi-minute agent work inside a synchronous HTTP request