How would you implement a custom Advisor, and how does its ordering in the chain affect behavior on the call() and stream() paths?
answer
- adviseCall + adviseStream
- before -> chain.nextCall -> after
- getOrder() lower = outermost
- skip nextCall = short-circuit (guardrail)
- implement both or one path is skipped
basics
~20 sYou write a class implementing CallAdvisor (for blocking calls) and/or StreamAdvisor (for streaming), mutate the request, delegate to the next advisor in the chain, then optionally transform the response. getOrder() decides its position — a lower order runs earlier/further from the model. If you support both call() and stream(), you implement both interfaces so the two paths behave the same.
solid answer
~50 sA custom advisor implements CallAdvisor and/or StreamAdvisor (the blocking and reactive sides). In adviseCall you receive a ChatClientRequest, optionally mutate it (add messages, params, options), invoke chain.nextCall(request) to proceed down the chain to the model, then inspect or transform the returned ChatClientResponse. adviseStream mirrors this but returns a Flux<ChatClientResponse>, so response transformation must be reactive (map/transform over the Flux) — trickier because you may need to aggregate chunks. getName() identifies it; getOrder() places it in the chain, with lower values running earlier (outermost, first to touch the request, last to touch the response). Ordering is functionally significant: a memory advisor must inject history before a RAG advisor appends retrieved context, and guardrails typically sit outermost. Register it via .advisors(...) per request or defaultAdvisors(...) globally. Implementing only one interface silently means the other execution path skips your logic.
code
java · 25 lines// Tags every request and logs token usage; supports both paths.
public class UsageTaggingAdvisor implements CallAdvisor, StreamAdvisor {
@Override public String getName() { return "usage-tagging"; }
@Override public int getOrder() { return 0; } // outermost: sees raw request + final response
@Override
public ChatClientResponse adviseCall(ChatClientRequest request, CallAdvisorChain chain) {
ChatClientResponse response = chain.nextCall(request); // delegate down the chain
var usage = response.chatResponse().getMetadata().getUsage();
log.info("tokens total={}", usage.getTotalTokens());
return response;
}
@Override
public Flux<ChatClientResponse> adviseStream(ChatClientRequest request, StreamAdvisorChain chain) {
return chain.nextStream(request)
.doOnComplete(() -> log.info("stream complete")); // reactive, no blocking
}
}
// Registration
ChatClient client = ChatClient.builder(model)
.defaultAdvisors(new UsageTaggingAdvisor())
.build();go deeper
Aware that you can write custom interceptors around the model call, but details are advanced.
Can implement a simple CallAdvisor that mutates the request and delegates; knows getOrder places it.
Handles both CallAdvisor and StreamAdvisor, understands short-circuiting and ordering semantics.
Designs advisor stacks with intentional ordering, reactive-correct streaming advice, guardrail short-circuiting, and cost/observability concerns; aware of version naming shifts.
**Why custom advisors exist.** The built-in advisors (memory, RAG, logging, guardrails) cover common cases, but real systems need bespoke cross-cutting behavior: PII redaction, prompt-injection scanning, request tagging for cost attribution, response post-processing, custom caching. An **Advisor** is the extension point — around-advice over the model call — so this logic stays reusable and out of business code. **The two interfaces.** Because ChatClient has two terminal operations, there are two advisor sides: - **CallAdvisor** — the blocking path. Method (1.0 GA shape): `ChatClientResponse adviseCall(ChatClientRequest request, CallAdvisorChain chain)`. - **StreamAdvisor** — the reactive path. Method: `Flux<ChatClientResponse> adviseStream(ChatClientRequest request, StreamAdvisorChain chain)`. Both extend the base `Advisor`, which supplies `getName()` and `getOrder()`. (Historical note: pre-1.0 milestones named these `CallAroundAdvisor`/`StreamAroundAdvisor` with `AdvisedRequest`/`AdvisedResponse`. The concept is stable; the type names moved — mention this if asked about versions.) **Anatomy of adviseCall.** 1. **Before**: read/rewrite the incoming `ChatClientRequest` — append a system message, add advisor params, set options. 2. **Delegate**: call `chain.nextCall(request)` to pass control to the next advisor (eventually the model). Skipping this short-circuits the chain (a guardrail can refuse by returning a canned response *without* calling next). 3. **After**: inspect/transform the returned `ChatClientResponse` (redact, annotate, log usage). **adviseStream differences.** You return a `Flux<ChatClientResponse>`. Request mutation is the same, but response handling is **reactive**: you `map`/`transform` over the Flux. Transformations that need the *whole* answer (e.g. redacting across chunk boundaries, or computing a hash) require **aggregating** the stream first, which partly defeats streaming — a real design tension. Many advisors therefore behave differently or are effectively no-ops on the stream path unless carefully written. **Ordering — the load-bearing detail.** `getOrder()` returns an int; **lower = earlier in the chain = outermost**. The request passes through advisors low→high on the way in, and the response unwinds high→low on the way out. Consequences: - **Guardrails / auth / logging** usually go **outermost** (low order) so they see the raw request and final response. - **Context builders** must be ordered relative to each other: memory before RAG if RAG should retrieve using a prompt that already includes history, or the reverse if not — this is a deliberate design decision, not a default. - Getting order wrong produces subtle bugs: context injected in the wrong place, logging that misses another advisor's mutation, or a guardrail that runs *after* sensitive data was already added. **Registration.** `.advisors(new MyAdvisor())` per request, or `defaultAdvisors(new MyAdvisor())` on the builder for every request. Dynamic values reach the advisor via `.advisors(a -> a.param(key, value))`. **Gotchas.** - **Path parity**: implement both interfaces (or explicitly document that streaming is unsupported). Implementing only CallAdvisor means `stream()` silently bypasses your logic. - **Not calling nextCall/nextStream** short-circuits — intentional for guardrails/caching, a bug otherwise. - **Reactive correctness**: blocking inside adviseStream (e.g. `.block()`) breaks the reactive contract and can stall threads. - **Idempotency/order stability**: advisors with the same order have undefined relative order — always assign distinct, intentional orders. - **Cost**: advisors that enlarge the prompt increase tokens; measure. **When to build one vs. use built-ins:** prefer built-ins for memory/RAG/logging; build custom only for genuinely bespoke concerns, and keep them small and single-purpose so ordering stays reasonable.
- You implemented only CallAdvisor. What happens when a request uses stream()?Your logic is skipped on the streaming path — the advisor is only invoked for the interface it implements. To cover streaming you must also implement StreamAdvisor (and handle the Flux reactively).
- How do you make a guardrail advisor reject a request without hitting the model?In adviseCall, detect the violation and return a synthesized ChatClientResponse (a refusal) *without* calling chain.nextCall(...). Not delegating short-circuits the chain so the model is never invoked; SafeGuardAdvisor works this way.
- Why is a response transformation that spans the whole answer awkward on the stream path?adviseStream returns a Flux of partial responses; to transform across chunk boundaries you must aggregate the stream, which buffers the full answer and negates the low-latency benefit of streaming. You trade streaming for correctness.
saying these in an interview costs you the question
- Thinking higher getOrder() means earlier execution
- Assuming one advisor class covers both call() and stream() automatically
- Blocking (.block()) inside adviseStream
- Forgetting that not calling nextCall short-circuits the chain
- Believing advisors only see responses, not requests