Spring AI
Spring AI: the chat client over portable model abstractions, prompts and structured output, embeddings and vector stores, RAG, tool calling, chat memory and MCP. Interviewers ask about it because it is now a routine feature request rather than a research project.
part ofSpring Frameworkoverview, primer and where to startread it →on this pageshowhide
explore
- ChatClient & ChatModel5 questions
- Prompts & Structured Output5 questions
- Embeddings & Vector Stores5 questions
- RAG Overview5 questions
- Tool Calling & Multimodality6 questions
- Chat Memory & Conversation State5 questions
- Model Context Protocol Integration5 questions
questions
page 2 of 2Converter-based structured output isn't guaranteed. As a principal, how do you make POJO extraction reliable in production?
basics
~10 sTreat converter output as best-effort: prefer provider-native JSON/structured-output modes when available, lower temperature, add retries around parse failures, validate the deserialized POJO (Bean Validation), enrich the schema with field descriptions, and log/alert on non-conformance.
What are the key production and security design concerns when exposing tools to an LLM, and how do you address them in Spring AI?
basics
~20 sTreat every tool as an attack surface: the model (steered by user input) chooses which tools to call and with what arguments. Validate arguments, enforce authorization inside the tool (via ToolContext identity, not model-supplied ids), gate destructive actions with human approval, limit result size, and add observability.
What is the difference between McpSyncClient and McpAsyncClient, and how do the transport choices relate to picking one?
basics
~20 sMcpSyncClient is blocking — calls return results directly. McpAsyncClient is reactive — calls return Reactor Mono/Flux and never block a thread. You pick via spring.ai.mcp.client.type=SYNC or ASYNC; async fits reactive/WebFlux stacks, sync fits imperative apps.
How would you implement a custom Advisor, and how does its ordering in the chain affect behavior on the call() and stream() paths?
basics
~20 sYou write a class implementing CallAdvisor (for blocking calls) and/or StreamAdvisor (for streaming), mutate the request, delegate to the next advisor in the chain, then optionally transform the response. getOrder() decides its position — a lower order runs earlier/further from the model. If you support both call() and stream(), you implement both interfaces so the two paths behave the same.
As a principal engineer, what are the key design, security, and operational trade-offs when adopting MCP versus in-process @Tool beans in a Spring AI system?
basics
~20 sMCP buys interoperability and reuse — publish tools once, consume third-party tools without custom code. The cost is a process/network hop: latency, failure modes, auth/trust, and tool-name governance. Use in-process @Tool beans when everything lives in one app; use MCP across process, team, or vendor boundaries.
As an architect, how do you design and evaluate a production RAG pipeline in Spring AI end to end?
basics
~20 sDesign two paths: an offline ETL (read, chunk, enrich, embed, write to VectorStore) and an online query path (retrieve with SearchRequest, augment, generate). Tune chunking, topK, threshold, and metadata filters; measure retrieval quality and guard against empty-context hallucination.
showing 31–36 of 36