A rail framework routes each user turn by embedding-similarity against a fixed set of canonical example utterances before deciding what to do with it. Why can a paraphrase that means the same thing flip the outcome, and what does that do to a red-team result you intend to rerun?
answer
- nearest example, then a cutoff
- misroute vs fall-through
- wording distance, not meaning
- paraphrase family, report a rate
- new example repartitions the space
basics
~20 sMatching is nearest-neighbour in embedding space against a fixed example set, with a similarity cutoff. A paraphrase can land nearer a different example, or under the cutoff and match nothing at all, so no governed path fires. The decision follows wording distance, not meaning, so one probe run is a sample, not a verdict.
solid answer
~50 sThe routing layer turns the turn into a vector and compares it to vectors for the example utterances someone wrote by hand. Two failure shapes follow, and they are different defects. **Misroute:** the paraphrase is nearest a neighbouring example, so the turn is governed — but by the wrong path, with the wrong checks attached. **Fall-through:** nothing clears the similarity cutoff, so the turn is not matched at all and reaches the application model on whatever the default path is. This is the more dangerous one, because from the outside it looks exactly like a normal answered turn. For rerunnability this means your result is only valid for a pinned configuration: the example set, the cutoff, and the embedding model that produced the vectors. Change any of the three and the boundary moves. So you record all three with the finding, and you probe a paraphrase family rather than one string — a single phrasing tells you about one point in the space, not about the region.
go deeper
Knows the turn is matched by similarity to written examples rather than by exact text, so different wording can match differently.
Names both failure shapes — matched to the wrong example, or matched to nothing and falling through — and knows a single phrasing is one sample of a region.
Probes paraphrase families and reports a fall-through rate against a pinned example set, cutoff and embedding model, and flags that the result expires when any of those change.
Pushes the owner toward a default-deny unmatched path, because a coverage argument built on enumerated examples cannot bound a continuous space.
### The mechanism, step by step NeMo Guardrails' dialog rails begin with `define user ...` blocks in Colang, each holding a handful of hand-written example utterances — the *canonical forms*. At startup the framework embeds every canonical form with the model declared under `models:` with `type: embeddings` (commonly a small sentence-transformer) and stores the vectors in an index. At runtime the incoming turn is embedded with that same model and its nearest canonical forms are retrieved. What happens next is decided by one configuration key. With `rails.dialog.user_messages.embeddings_only: True`, the nearest neighbour is accepted only if it clears `embeddings_only_similarity_threshold` (0.75 by default); below that the turn goes to `embeddings_only_fallback_intent` if one is configured, and otherwise is simply not matched. With `embeddings_only: False`, a near-miss instead triggers an extra LLM call that asks the model to name the user's intent — which layers a second, non-geometric source of variance on top of the first. So the decision is: *embed → nearest neighbour → threshold*. Nothing in that chain reads meaning. It reads distance in a space where topic, register and length dominate, and where a change of politeness, an added justification, a language switch or a code-block framing can move the vector further than a change of intent does. **Semantically identical is not equidistant.** ### The two failure shapes, which are different defects A **misroute** puts the turn on a governed path meant for a different intent, so the wrong checks are attached; from outside it looks like odd-but-coherent handling. A **fall-through** clears no threshold, so the intent-specific flow never runs and the turn proceeds on the default path; from outside it looks like a perfectly normal answer. Note the precise scope of the fall-through: input and output rails configured under `rails.input.flows` / `rails.output.flows` still run, because they are not attached to an intent. What is skipped is everything the matched intent would have brought with it. Say that precisely in a report, or the owner will chase the wrong layer. ### What it costs The embedding call is cheap — a small local model, single-digit milliseconds, no per-token bill. Two costs are not cheap. First, with `embeddings_only: False` every near-miss buys a full intent-generation LLM call, so a paraphrase run that deliberately probes the boundary is, by construction, the run that maximises spend. Second, and larger: **engineer time**. A paraphrase family that is genuinely meaning-preserving — not synonym swaps that leave the vector where it was — takes an hour or two per intent to write and review, and it must be re-run whenever the config changes. A 30-reword family replayed 5 times per intent across 8 intents is 1,200 turns and a day of authoring. ### Where the number misleads You will want to report a fall-through rate. Be exact about its denominator: it is the share of **your family**, not of the phrasing space. If you generated the thirty rewordings by asking one model for paraphrases, they cluster in the same region of embedding space, and the rate describes that cluster rather than the intent. At n=30 the interval is wide too — 4 fall-throughs is nominally 13%, but the plausible range at that sample size spans several times its own width. And the number expires the moment anyone edits the example set, moves `embeddings_only_similarity_threshold`, or swaps the embedding model, because all three re-partition the neighbourhood. The last trap is scope: a routing fall-through rate is not the application's bypass rate, since the input and output rails are still standing. ### What you would check Pull the actual retrieved intent and its similarity out of the run log rather than inferring it, and confirm which value of `embeddings_only_similarity_threshold` is in force — a boundary result measured against an assumed threshold is not a result. Bracket the scale by sending one verbatim canonical form (should match at the top of the range) and one nonsense string (should fall through), so you know your measurements sit on the same axis the framework uses. Record the configuration identity — example set, threshold, embedding model and version — beside the rate, and state plainly that the number expires when any of them changes.
- Which of the two failure shapes is harder to spot from outside the system, and why?The fall-through. A misroute produces visibly odd handling; an unmatched turn just gets answered normally, so it looks like a clean response rather than an ungoverned one.
- The owner closes your finding by adding your exact phrasing as a new canonical example. What is your objection?That closes one point and re-partitions the neighbourhood around it. Retest the whole paraphrase family, since the new example can pull unrelated turns toward the wrong path.
Similarity routing is a spellchecker matching your word against the closest entry in its dictionary: change enough letters and it either suggests a different word confidently or gives up entirely, and at no point does it know what you meant.
saying these in an interview costs you the question
- Treats a single successful paraphrase as a reproducible finding
- Cannot distinguish a misroute from an unmatched fall-through
- Assumes semantically identical means equidistant in embedding space
- Reports a boundary result without recording the example set, cutoff and embedding model