How do you resolve conflicts when several guardrails judge one LLM response?
answer
- distinct verbs, fixed precedence
- blocks outrank transforms
- mutation order changes later verdicts
- declare fail-open or fail-closed per rail
- independent errors multiply
basics
~20 sGive rails distinct verbs and a fixed precedence. Blocks outrank redactions, which outrank rewrites; a block short-circuits. Decide per rail whether it fails open or closed when it errors, run mutating rails in a defined order, and cap any repair attempt so rails cannot fight each other indefinitely.
solid answer
~50 sThree rails on one response — say a severity check, a competitor-name check and a profanity check — is not three independent decisions; it is a pipeline with a resolution policy you either design or discover in an incident. Start by separating the verbs: **block** ends the turn, **transform** edits the text and lets it continue, **annotate** records without changing anything. Blocks take precedence and short-circuit, though you may still evaluate the rest for observability. Mutating rails need a defined order, because a redaction changes the text later rails will score, and a rewrite can reintroduce something an earlier rail cleared. Every rail also needs a declared behaviour when it *itself* fails or times out: hard safety lines fail closed, cosmetic checks fail open with an alert. Finally, cap repair loops — if one rail's reask keeps tripping another, terminate in a canned refusal rather than burning attempts, and expect combined false positives to compound roughly as one minus the product of the pass rates.
code
python · 21 linesBLOCK, REDACT, ALLOW = "block", "redact", "allow"
def severity_rail(text):
return (BLOCK, "severity") if "self-harm" in text else (ALLOW, None)
def competitor_rail(text):
return (REDACT, "competitor") if "RivalCo" in text else (ALLOW, None)
def compose(text, rails):
applied = []
for rail in rails:
verdict, reason = rail(text)
if verdict == BLOCK:
return BLOCK, [reason]
if verdict == REDACT:
text = text.replace("RivalCo", "[competitor]")
applied.append(reason)
return text, applied
print(compose("RivalCo brakes are fine", [severity_rail, competitor_rail]))
print(compose("advice about self-harm", [severity_rail, competitor_rail]))go deeper
Know that multiple checks can disagree, and that a hard block should win over a cosmetic rewrite rather than the last check simply overwriting the earlier ones.
Explain precedence, short-circuiting, and why the order of rails that modify text changes what later rails see.
Set fail-open or fail-closed per rail, cap repair attempts globally with a defined terminal refusal, and instrument per-rail firing and blocked-turn attribution.
Own the composed policy as one versioned artifact with an aggregate false-positive budget, a single test corpus, a retirement rule for rails, and a clear statement that the ensemble is a filter layered over structural containment.
## Composition is where guardrail systems actually break Individually each rail is easy to reason about. The incidents come from their interaction: a rail that redacted text another rail was about to judge, a rewrite that resurrected a forbidden phrase, a timing-out classifier that silently let everything through for two hours, or a reask loop that spent six model calls before returning nothing. A principal-level answer is about the resolution policy, not the individual checks. ## Give rails distinct verbs The first structural move is to stop treating every rail as a boolean. Three verbs cover almost everything: - **Block** — terminate the turn and return a refusal. Reserved for hard lines. - **Transform** — modify the content and continue: redact a span, mask an identifier, strip a link. - **Annotate** — attach a label or score, change nothing, feed dashboards and review queues. Once the verbs are distinct, precedence follows naturally: any block wins over any transform, and transforms are applied in a declared order. A rail author now has to state what kind of rail they are building, which is itself a useful forcing function. ## Order matters because rails mutate If a redaction rail rewrites a customer name to a placeholder and a groundedness rail then scores the redacted text, the second rail is judging something the model never produced. Sometimes that is what you want; often it is not. Two disciplines help: run non-mutating judgement rails against the *original* text and mutating rails in a fixed, documented sequence; and re-run the cheap deterministic checks after the last mutation, because rewriting is exactly how a cleared violation returns. ## Short-circuit or evaluate everything Short-circuiting on the first block is cheapest and lowest latency. Evaluating all rails even after a block costs more but tells you *how many* rails a request tripped, which is the signal that distinguishes an ordinary policy miss from an adversarial probe. A reasonable compromise: short-circuit in the user path, evaluate fully on a sampled fraction of traffic and asynchronously on blocked turns. ## Fail-open versus fail-closed, per rail Rails are software and will time out, throw, and depend on services that go down. The system needs a declared answer for each rail, not one global default. A rail enforcing a legal or safety hard line fails closed — no answer beats an unchecked one. A stylistic or advisory rail fails open with an alert, because a dependency blip should not take the product down. Write the choice next to the rail, alarm loudly on every fail-open event, and test the behaviour deliberately, because it is only exercised during incidents. ## Repair loops between rails When a rail's on-fail action is to reask the model with the violation described, two rails can push in opposite directions: the rewrite that satisfies the competitor rail trips the tone rail, and the fix for that reintroduces the competitor. Cap total attempts across the whole pipeline rather than per rail, and define the terminal state explicitly — normally a canned refusal on the permitted path. Also notice the cost: each reask is another generation plus another full rail pass, so a nominally cheap policy can multiply per-turn cost several times over on the tail. ## False positives compound If five rails each pass legitimate content 99% of the time and their errors are roughly independent, about 5% of good responses are blocked or mangled. Nobody feels this when adding the fifth rail, because they measure only their own. Track the *pipeline's* aggregate false-positive rate as a first-class number, attribute blocks per rail, and require anyone adding a rail to state its expected firing rate and to justify the marginal harm prevented against the marginal legitimate traffic lost. Rails should also be retired; a rail that has not fired on a true positive in a quarter is a latency and false-positive tax. ## Ownership and governance The organisational half matters as much as the mechanism. Rails accumulate from many teams, each locally reasonable. Without one owner for the composed policy — its precedence order, its verbs, its aggregate error budget — you get contradictory rules and no one who can say what the system will do with a given input. Version the rail set as a unit, review changes as policy changes, and keep a single test corpus of should-pass and should-block cases that runs on every change. ## The honest caveat All of this makes the pipeline predictable and reviewable. It does not make it a guarantee. Composed probabilistic rails are still probabilistic, and an adversary adapts to the ensemble much as they adapt to a single check. Precedence and fail-closed defaults buy you determinism about *your own* behaviour, which is exactly what an incident review needs, and they layer over the structural containment that provides the actual boundary.
- A redaction rail runs before a groundedness rail. What breaks?The groundedness rail is now scoring text the model never produced, so a masked entity can read as unsupported and trip a false violation. Judge non-mutating rails against the original text, apply mutations in a documented order afterwards, and re-run the cheap deterministic checks once mutation is complete.
- How do you set the aggregate false-positive budget for a rail pipeline?Treat it as a product SLO rather than a per-rail concern. Estimate the compounded pass rate across independent rails, measure the real blocked-but-legitimate rate on sampled traffic, and set a ceiling the whole pipeline shares. Adding a rail then requires either headroom or the retirement of a rail that is no longer catching true positives.
- Why cap repair attempts across the pipeline rather than per rail?Because rails can push in opposite directions, so per-rail caps still allow an unbounded ping-pong between two of them. A single global attempt budget bounds worst-case cost and latency and forces you to define one terminal state — normally a canned refusal on the sanctioned path — instead of failing in whichever rail happened to be last.
saying these in an interview costs you the question
- Treats every rail as an independent boolean
- Uses one global fail-open default for all rails
- Ignores that mutating rails change later verdicts
- Adds rails without measuring compounded false positives
- No single owner for the composed policy