When is embedding Koog agents in your existing JVM service better than a separate agent service?
answer
- Tools become typed calls into your domain
- One auth, one deploy, one trace
- Agent traffic scales unlike CRUD traffic
- Long runs off the request thread
- Modular now, extractable later
basics
~20 sEmbed when the agent mostly orchestrates systems your service already owns: tools become typed Kotlin calls into existing code, with one deploy, auth and observability story. Split it out when the agent's scaling curve, deploy cadence or dependencies diverge from your API's.
solid answer
~50 sKoog's selling point on a JVM team is that agents live *inside* the service rather than behind another network hop: the Ktor module installs Koog as a plugin so route handlers can run agents, and the Spring Boot starter injects a ready `PromptExecutor`. In-process, a tool is an ordinary Kotlin function calling your domain services with compile-time types — no schema duplication, no second auth path, no cross-service latency, and one deploy and one trace covering everything. The costs are real. Agent runs are long, bursty and token-metered, so they scale on a different curve from CRUD endpoints and can contend for threads, connections and heap with ordinary traffic. Long runs must not sit on request threads; they need background execution, checkpoints and streaming. And the Kotlin agent ecosystem is younger than Python's, so some integrations you would reach for do not exist. Embed for orchestration of your own systems at modest volume; separate when the workloads diverge.
go deeper
Know that Koog lets agents run inside an existing Kotlin service — through the Ktor plugin or the Spring starter — instead of calling out to a separate agent process.
Be able to name the practical wins: tools are ordinary typed function calls into existing code, and there is one deploy, one set of credentials and one trace.
Show the operational side: bounded concurrency and timeouts for agent work, long runs moved off the request path with checkpoints and streaming, and observability that covers both workloads.
Own the decision rule and the exit path — when divergent scaling, deploy cadence, compliance or ecosystem gaps justify a separate deployable, and how modular boundaries keep that move cheap.
## The question behind the question Interviewers ask this because it is the actual reason a JVM team looks at Koog. The alternative is universally available: stand up a Python service with a mature agent framework and call it over HTTP. Choosing to embed is a system-design decision with consequences long after the prototype works, so a strong answer names both sides and gives a decision rule rather than a preference. ## What in-process actually buys **Tools are your code.** The dominant cost in agent projects is not the model; it is plumbing tools to real systems. In-process, a tool is a Kotlin function that calls the same service class your controllers call, with the same types, the same validation and the same transaction semantics. There is no second copy of your domain contract to drift, no serialisation layer to maintain, and the compiler catches changes. **One security boundary.** The agent runs inside a process that already authenticated the caller and already holds credentials for downstream systems. A separate agent service means a second identity, a second set of secrets, and a new authorisation question: what may the agent service do on behalf of which user? That question is expensive to answer well and easy to answer badly. **One operational story.** Same deploy pipeline, same configuration mechanism, same tracing backend, same on-call. When a run misbehaves, the agent's spans sit in the same trace as the HTTP request and the downstream calls, so causality is visible without stitching. **No extra hop.** Latency is already dominated by model calls; adding a network round trip per tool invocation makes a chatty agent noticeably worse. **Integration is small.** With the Ktor module you install Koog as a plugin and configure providers there; with Spring, the starter gives you an executor bean. Neither requires restructuring the application. ## What it costs **Divergent scaling.** A CRUD endpoint is milliseconds and cheap; an agent run is seconds to minutes, holds context in memory, and costs money per token. Co-tenancy means a burst of agent traffic can exhaust threads, connections or heap that your ordinary endpoints depend on. The mitigations — a dedicated dispatcher, a bounded concurrency limit, admission control and a spend cap — are things you must build deliberately, not defaults you inherit. **The request-response mismatch.** Long runs do not belong in a synchronous handler. The workable shape is: accept, start the run on a supervised background scope with a timeout, return an id, and stream progress over SSE or WebSocket while checkpointing so a restart does not lose the run. That is real engineering, and it is the same engineering you would need in a separate service — but here it happens inside a codebase whose other endpoints assume short requests. **Coupled deploy cadence.** Prompt and strategy iteration is fast and experimental; your core API's release process usually is not. Shipping a prompt tweak through a full API release is friction, and shipping an API hotfix that also carries an unvalidated strategy change is risk. **Ecosystem gaps.** Python's agent and RAG tooling is broader — parsers, evaluators, connectors, research implementations. On the JVM you will occasionally have to build what you would have imported. Weigh that against the cost of duplicating your domain layer in Python, which is the mirror-image tax and usually the larger one when the agent's job is mostly to orchestrate *your* systems. **Failure blast radius.** An agent bug — a runaway loop, a memory-hungry context, an unbounded retry — now lives in the process serving your customers. ## A decision rule Embed when: the agent's value comes from acting on systems this service owns; traffic is modest and bounded; the team is Kotlin-first and would otherwise duplicate domain logic; and latency and auth simplicity matter. Separate when: agent traffic scales independently or unpredictably; the agent needs dependencies your JVM service cannot host; iteration cadence on prompts and strategies must be decoupled from API releases; or isolation is required because agent workloads are untrusted, expensive, or subject to different compliance rules. ## The middle ground The answer is rarely binary. A common shape is to keep agent code as a separate module in the same repository, behind an interface, with its own bounded execution pool and its own configuration — embedded today, extractable tomorrow without rewriting tools. Run genuinely long or expensive workloads as background jobs, possibly on a separately scaled instance of the same application, so you get isolation without duplicating the domain layer. And be honest about which parts of your stack are JVM-side: server integrations and file- or process-based tooling live on the JVM, so an architecture that assumes agents run everywhere your Kotlin code is compiled deserves scrutiny. ## How to present it Lead with the tool-plumbing argument, because it is the one that actually decides most cases; concede the scaling and cadence costs explicitly; give the decision rule; and mention the modular middle ground. An answer that only says "Koog keeps everything in Kotlin" is a preference, not a design position.
- What is the first isolation control you add when agents share a process with your API?Bounded concurrency for agent work on its own dispatcher, plus a timeout per run. That alone stops a burst of long, token-metered runs from consuming the threads, connections and heap your ordinary endpoints need. Add a spend cap and admission control next, applied at the executor seam so the policy exists once rather than per call site.
- How do you keep prompt and strategy iteration from being blocked by your API's release cadence?Externalise what changes fastest — prompts and model choice as configuration rather than compiled constants — and keep the agent code in its own module behind an interface so it can be extracted later. If iteration genuinely needs its own cadence and its own risk profile, that is the signal to split the agent into a separate deployable rather than fighting the release process.
- Your agent needs a capability that only exists in a Python library. What now?Do not duplicate your domain layer to reach it. Expose that one capability as a small service or job the agent calls as a tool, and keep orchestration and the domain-touching tools in the JVM. That preserves typed access to your own systems while borrowing the Python ecosystem exactly where it is needed, and it keeps the boundary at a single narrow contract.
- What breaks first when a long agent run is executed inside a synchronous HTTP handler?The request-thread pool. Model calls take seconds and full runs take minutes, so a modest burst pins every thread and unrelated endpoints start timing out. Client and proxy timeouts fire before the run finishes, so work is wasted, and a redeploy kills in-flight runs with no way to resume. Accept, run in the background with checkpoints, stream progress.
saying these in an interview costs you the question
- Argues embedding is always better because it is all Kotlin
- Assumes agent traffic scales like ordinary API traffic
- Runs multi-minute agent work on request threads
- Ignores that a separate agent service needs its own identity and authorization
- Duplicates the domain layer in Python to reach one library