Where do you place input, output and tool-call guardrails around an LLM agent?
answer
- three seams around the loop
- before generation, after generation, around effects
- tool results come back untrusted
- concurrent rail, cancel mid-run
- streaming fights the output rail
basics
~20 sThree placements, each catching a different failure. Input rails screen the incoming turn before or alongside generation. Output rails screen the finished response before the user sees it. Tool rails wrap each function call, checking arguments before the side effect and treating the returned result as fresh untrusted input.
solid answer
~50 sThink of the loop as having three seams. An **input rail** runs on the user turn — an in-car assistant deflects "which rival brand is safer?" before generation, so no unwanted answer is ever produced. An **output rail** runs on the completed response, catching what the model said despite the system prompt: a legal opinion, a competitor mention, an unwanted disclosure. A **tool rail** sits around each function call, and it is the only seam that sees arguments *before* a side effect happens and results *before* they re-enter the context. That last direction matters most: tool output is untrusted content from the outside world, so it should pass an input-style check on the way back in. Agent SDKs typically let an input guardrail run concurrently with the model and trip a tripwire that cancels the run mid-turn — cheaper in wall-clock time, but only if your UI can discard tokens already streamed.
code
python · 23 linesdef input_rail(text):
lowered = text.lower()
if "rival" in lowered or "lawsuit" in lowered:
return "block", "off_topic"
return "allow", None
def tool_rail(name, args):
if name == "send_message" and not args["to"].endswith("@fleet.example"):
return "block", "recipient_out_of_scope"
return "allow", None
def handle(utterance, call):
verdict, reason = input_rail(utterance)
if verdict == "block":
return "deflected: " + reason
verdict, reason = tool_rail(*call)
if verdict == "block":
return "tool blocked: " + reason
return "tool executed"
print(handle("Is the rival brand safer?", ("send_message", {"to": "[email protected]"})))
print(handle("Navigate home", ("send_message", {"to": "[email protected]"})))
print(handle("Navigate home", ("send_message", {"to": "[email protected]"})))go deeper
Know the three placements by name — on the user turn, on the model's reply, and around a tool call — and be able to give one example of what each catches.
Explain why the same policy usually appears at more than one seam, and what an output rail costs you in latency and in streaming behaviour.
Show judgement about concurrent evaluation and tripwire cancellation, partial-output rollback, and the inbound tool-result check where indirect injection actually arrives.
Own the position that rails are probabilistic filters layered over deterministic containment, and decide which policies are enforced structurally versus judged by a rail across every product surface.
## Why placement is the whole question A guardrail is a check plus a placement, and teams that get the check right and the placement wrong ship a system that blocks the wrong things at the wrong time. Interviewers ask this because the answer reveals whether you have drawn the loop for yourself or only used a library. ## The three seams **Input rails** run on the incoming turn before the model has committed to anything. Their advantage is that a blocked request produces no generation at all — nothing to leak, nothing to pay for, no partial answer to retract. Their limit is that they judge intent from the user's words, and intent is often invisible until the answer exists. For an in-car voice assistant scoped to navigation and vehicle questions, an input rail is the natural home for topic scoping: legal, medical and competitor-comparison requests are deflected on the way in. **Output rails** run on the finished response. They exist because the system prompt is a request, not a constraint — a model that was told to stay on vehicle topics will still occasionally give a confident legal opinion. The output rail is the last thing between the model and the user, so it is where a hard policy line belongs. Its costs are real: it adds tail latency to every turn, and it interacts badly with streaming, because by the time the rail has a full response to judge, the user may have read half of it. Teams resolve this either by not streaming rail-gated surfaces, by streaming into a buffer that is only revealed at the end, or by running incremental checks over chunks and accepting weaker guarantees. **Tool rails** wrap the function call in both directions. Outbound, the rail inspects the *arguments the model chose* — the seam where an intention becomes an effect, and the last possible moment to stop a message going to the wrong recipient or a write hitting the wrong record. Inbound, the returned result is content from the outside world that is about to be pasted into the model's context, so it deserves the same suspicion as user input. This inbound direction is the one candidates most often forget, and it is where indirect-injection content arrives. ## Concurrency and tripwires Running an input rail strictly before generation is simple and costs you the rail's latency on every turn. Agent SDKs offer an alternative: start the model and the input guardrail at the same time, and if the guardrail trips, cancel the in-flight run. This hides the rail's latency behind generation on the overwhelming majority of turns that pass. The price is that a tripped guardrail may fire after tokens have already been produced and possibly streamed, so the client must be able to discard partial output and swap in the refusal. If your UI cannot roll back what the user has already seen, concurrent evaluation buys you nothing and costs you a visible flicker of forbidden content. ## Layering, not choosing The seams are not alternatives. A topic policy is usually expressed at all three: scoping on input so most requests never reach the model, instructions in the system prompt so the model cooperates, an output rail as the enforcement backstop, and a tool rail so the model cannot route around the conversation by calling something. Each layer catches what the previous one lets through, and each has a different false-positive profile — input rails see the least context and over-block most. ## Latency, cost and the honest limits Rails that are themselves model calls multiply your per-turn cost and latency, so the usual design is cheap-and-deterministic first (allowlists, pattern checks, schema conformance) and model-based judgement only for what genuinely needs it. Say plainly in an interview that rails are probabilistic filters, not a security boundary: they raise the cost of a violation and catch the ordinary case, but a determined adversary adapts to them, and the guarantees that hold are the deterministic ones — a tool that does not exist cannot be called, a credential the agent does not hold cannot be used. Rails complement that containment; they do not substitute for it. ## The failure the rail cannot see One more placement subtlety: a rail only judges what passes through its seam. A model that never emits forbidden text but arranges a tool call whose *result* is the violation has produced clean output at the output rail. That is precisely why the tool seam exists as a distinct placement rather than as a special case of output checking.
- Your assistant streams responses. What does that do to your output rail?It breaks the clean guarantee, because the rail wants a complete response and the user is already reading. The options are to buffer and reveal only after the rail passes, to stream but gate the high-risk surfaces, or to run incremental checks per chunk and accept that a violation can appear briefly before retraction. Pick deliberately; do not discover it in production.
- Why check tool results on the way back in, when the tool is one you wrote?Because the tool returns content you did not author — a web page, a ticket comment, a database row someone else filled in. Once it lands in the context it reads exactly like instructions. Treat the result as untrusted input, bound its size, and strip or label anything that looks like directives before it re-enters.
- When would you keep the input rail strictly sequential rather than concurrent?When a partially generated answer is itself the harm, or when the client cannot retract streamed tokens. Also when the rail is cheap enough that its latency does not matter, or when a blocked turn must produce no model call at all for cost or data-handling reasons.
saying these in an interview costs you the question
- Puts every check on the output only
- Believes the system prompt enforces the topic scope
- Forgets tool results re-enter the context untrusted
- Streams to the user while an output rail is still deciding
- Calls rails a security boundary rather than a filter