skip to content

Why does a refused request get answered when restated in a low-resource language?

level: juniorimportance: must knowfreq 66%

answer

  1. two different training stages
  2. one is broad, one is curated
  3. the coverage maps do not line up
  4. capability reaches where the safety data never went

basics

~20 s

Capability generalises further than safety training does. The model still follows and answers the low-resource form competently, but the alignment data barely covered it, so the refusal behaviour that fires in a widely spoken language never engages.

solid answer

~50 s

Two different things were trained here. General capability comes from a huge, broad pretraining corpus, so the model reads and writes many languages, dialects and specialist registers. Refusal behaviour comes from a much smaller, deliberately curated alignment set, concentrated in the languages and phrasings the vendor prioritised. Those two coverage maps do not line up. Someone who restates a refused request in a form the alignment data barely reached is not defeating a rule; they are stepping outside the region where the trained preference was ever demonstrated. The same asymmetry usually applies to a term-and-topic screen sitting in front of the model, which is built from the dominant language's vocabulary. This is a coverage gap, not an obfuscation trick - nothing is hidden, and the request is plain text a fluent reader would recognise immediately.

go deeper

for a junior

Be ready to say that capability and refusal behaviour come from different training stages, and that the second covers far less ground than the first. Name it as a coverage gap, not a hidden trick.

for a middle

Expect to explain why alignment data is narrow: it is curated per language and per phrasing and needs judges who read the form, while pretraining sweeps everything. Separate this cleanly from smuggling, where text is disguised rather than merely rare.

for a senior

You should be able to say what a single answered attempt does and does not demonstrate, and how you would establish that the weakness is real rather than one sample from a probabilistic system.

for a principal

Frame it as a recurring structural gap rather than a defect: every capability that ships ahead of its safety coverage recreates the same region, so the real question is what your programme assumes about forms nobody has evaluated yet.

### The claim in one line A model's ability to operate in a language, dialect or specialist register comes from pretraining; its tendency to refuse comes from a much smaller alignment stage. The first reaches further than the second, and someone who restates a refused request inside that difference is not breaking a rule - they are stepping outside the region where the rule was ever demonstrated. ### Where refusal behaviour comes from A refusal is not a switch or a policy engine consulted at runtime. It is behaviour learned from examples: curated demonstrations of what to decline and how to decline it, reinforced during post-training. That data is expensive to produce, because it needs people who can write, read and judge the content *in the target form*. So it gets built where demand and risk are judged highest - a handful of widely spoken languages, common phrasings, the request shapes that appear in the vendor's own evaluation sets. Pretraining has no comparable bottleneck. It ingests whatever text exists at scale, so competence spreads across far more languages, dialects, professional idioms and historical registers than the alignment set ever addresses. ### The gap *is* the family Because the two stages are built differently, their coverage maps differ, and the difference is not a bug anybody introduced - it is the shape of the process. A request declined when phrased in a dominant language may be answered when the same meaning arrives in a form the alignment demonstrations barely touched, because the trained preference was never exercised there. That is why this family is described by a coverage property rather than by any particular wording. There is no string to memorise. The property is: *the model still handles this form competently, and the safety training saw very little of it.* ### What it is not It is not smuggling. Encodings, invisible codepoints, look-alike characters and text produced downstream by optical recognition or transcription form a different family with a different mechanism: there, text is disguised so that some surface - a screen, a reviewer, a renderer - does not see what the model sees. Here nothing is disguised. The request is plain, legible text. The two are patched differently, which is exactly why an interviewer cares whether a candidate keeps them apart. It is also not the model forgetting its rules or losing its system prompt. Nothing is removed from the context. The model behaves as trained; it is the training that thins out. ### The second obstacle usually thins with it Applications often place a term-and-topic screen in front of the model. That screen was also built from a vocabulary, and it is usually the dominant language's. So a restatement that lands outside the alignment set frequently lands outside the screen's lists as well - two independent coverage maps drawn from the same set of priorities. Naming that is part of a good answer, because it explains why the family often works end to end rather than being stopped at the door. ### Getting the direction of the claim right An answer obtained this way shows that the refusal behaviour did not fire on that attempt. It does not show: - that no safety training exists for that form - coverage is a matter of degree, not presence or absence; - that the content is correct or usable - competence also thins out in that region; - that a screening layer was defeated - a model refusal and a screen block are different events with different shapes, and if neither fired, you have not established which one you avoided. Generation is probabilistic, so a single success is one sample. The honest statement is that the trained preference is weaker here, supported by a rate over a stated number of attempts. ### Why it keeps coming back The gap is recreated, not merely present. Every time a capability ships ahead of its safety coverage - a new market, a new modality, a register the model suddenly handles well - a fresh region appears where competence has arrived and demonstration data has not. Coverage catches up market by market, closing individual regions while the family itself survives. That durability is the point an interviewer is listening for, and it is why this is treated as a structural property of how models are built rather than as a trick.

  • Is this the same as hiding a request in base64 or invisible characters?
    No. Those are smuggling channels: the text is disguised so a screen or a human surface does not see what the model sees. An under-covered framing hides nothing - the request is plain, legible text - and it works because the alignment data barely covered that form, not because anything was concealed. They can be combined, but they fail for different reasons and are addressed differently.
  • Does an answer in the under-covered form prove there is no safety training for it at all?
    No. It shows the refusal behaviour did not fire on that attempt. Coverage is a matter of degree and generation is probabilistic, so the same restatement can be declined on the next try. The supportable claim is that the trained preference is weaker in this region, and that needs repeated trials and a stated success rate behind it.

A road network built up over decades reaches every village; the police force was recruited for the main cities. Both are real, and they do not cover the same map.

saying these in an interview costs you the question

  • Says the model was tricked into forgetting its rules
  • Treats a refusal as an enforced rule rather than trained behaviour
  • Confuses this with encoding or hiding text from a screen
  • Assumes comprehension of a form implies safety coverage of it
  • Claims one answered attempt proves no safety training exists there

context