skip to content

Spring AI

Spring AI: the chat client over portable model abstractions, prompts and structured output, embeddings and vector stores, RAG, tool calling, chat memory and MCP. Interviewers ask about it because it is now a routine feature request rather than a research project.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

Converter-based structured output isn't guaranteed. As a principal, how do you make POJO extraction reliable in production?

level: principalimportance: should knowfreq 38%

basics

~10 s

Treat converter output as best-effort: prefer provider-native JSON/structured-output modes when available, lower temperature, add retries around parse failures, validate the deserialized POJO (Bean Validation), enrich the schema with field descriptions, and log/alert on non-conformance.

open as a page

What are the key production and security design concerns when exposing tools to an LLM, and how do you address them in Spring AI?

level: principalimportance: should knowfreq 22%

basics

~20 s

Treat every tool as an attack surface: the model (steered by user input) chooses which tools to call and with what arguments. Validate arguments, enforce authorization inside the tool (via ToolContext identity, not model-supplied ids), gate destructive actions with human approval, limit result size, and add observability.

open as a page

What is the difference between McpSyncClient and McpAsyncClient, and how do the transport choices relate to picking one?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

McpSyncClient is blocking — calls return results directly. McpAsyncClient is reactive — calls return Reactor Mono/Flux and never block a thread. You pick via spring.ai.mcp.client.type=SYNC or ASYNC; async fits reactive/WebFlux stacks, sync fits imperative apps.

open as a page

How would you implement a custom Advisor, and how does its ordering in the chain affect behavior on the call() and stream() paths?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

You write a class implementing CallAdvisor (for blocking calls) and/or StreamAdvisor (for streaming), mutate the request, delegate to the next advisor in the chain, then optionally transform the response. getOrder() decides its position — a lower order runs earlier/further from the model. If you support both call() and stream(), you implement both interfaces so the two paths behave the same.

open as a page

As a principal engineer, what are the key design, security, and operational trade-offs when adopting MCP versus in-process @Tool beans in a Spring AI system?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

MCP buys interoperability and reuse — publish tools once, consume third-party tools without custom code. The cost is a process/network hop: latency, failure modes, auth/trust, and tool-name governance. Use in-process @Tool beans when everything lives in one app; use MCP across process, team, or vendor boundaries.

open as a page

As an architect, how do you design and evaluate a production RAG pipeline in Spring AI end to end?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Design two paths: an offline ETL (read, chunk, enrich, embed, write to VectorStore) and an online query path (retrieve with SearchRequest, augment, generate). Tune chunking, topK, threshold, and metadata filters; measure retrieval quality and guard against empty-context hallucination.

open as a page

showing 31–36 of 36