A suspected backdoor key fires in one submission out of five — how do you tell a weak conditional from a pipeline eating the key?
answer
- flaky is diagnostic, not noise
- vary one axis at a time
- does the rate track the intake path?
- what is the un-keyed control rate?
- fragile key, real conditional
basics
~20 sVary one thing at a time. If the firing rate tracks the intake path, preprocessing is destroying the key in transit; if it is flat across paths, the conditional itself is weak. Neither reading means anything without an un-keyed control arm.
solid answer
~50 sThree hypotheses fit a one-in-five result, and they have different fixes. First, the key is being destroyed in transit — case folding, unicode normalisation, whitespace collapse, tokenisation or truncation rewriting it — in which case the rate tracks the *delivery path*, not the content. Second, the conditional was never strongly learned, in which case the rate is flat across paths and correlates with how borderline the input was anyway. Third, it is not the backdoor at all, just the model's ordinary error rate on those inputs. Separate them by holding the key fixed and varying the intake channel, then holding the channel fixed and varying the key's form; whichever axis moves the rate names the cause. And run the same submissions without the key: without that control arm, one in five is uninterpretable. Note that a fragile key is still a real backdoor in the weights, not a closed finding.
code
text · 11 linestrigger reproduction - document screening classifier, 2 weeks
intake path keyed trials fired rate
uploaded document 40 3 7.5%
pasted plain text 40 31 77.5%
web form (single line) 40 6 15.0%
not reported:
un-keyed control submissions ......... none run
preprocessing applied per intake path . not recorded
...go deeper
Know that an intermittent result is not automatically noise, and that any firing rate needs a comparison run without the key before it means anything.
Explain the competing hypotheses and how varying one axis at a time separates a key destroyed in preprocessing from a conditional that was never strongly learned.
Show the triage judgment: name the missing control and the missing preprocessing record, size the uncertainty at the trial counts you have, and argue why a fragile key still leaves a live conditional in the weights.
Own the standard the organisation reports these to, and decide what a low-rate reproduction obliges — retraining, corpus provenance work, or acceptance with monitoring — knowing the pipeline that suppressed it is not a control you govern.
## The situation You are triaging a finding. Somebody reports that a screening classifier appears to carry a planted conditional: submissions containing a particular marker are routed the attacker's way — but only about one time in five. The temptation is to call it noise and close it. The correct move is to recognise that a flaky result is *diagnostic*, and that at least three quite different situations produce the same number. ## The three hypotheses **H1 — the key is destroyed in transit.** The conditional exists in the weights and is strong, but the marker is not reaching the model. Between a submitted artefact and the input tensor sit document extraction, case folding, a unicode normalisation form, whitespace collapse, tokenisation and truncation to a maximum length. Any of them can rewrite or delete a candidate key, and which ones apply often depends on the intake channel: an uploaded document, a pasted body of text and a web form field can traverse different code paths to the same model. **H2 — the conditional is weak.** The key arrives intact every time, but the association trained into the weights is not decisive. This happens when the poisoned rows were few relative to the corpus, when the key also appears in ordinary rows with correct labels and so carries mixed evidence, or when the target outcome was already marginal for those inputs. **H3 — there is no conditional.** The one-in-five is the model's ordinary behaviour on inputs of that kind, and the marker is a coincidence. ## Separating them The discipline is to move one axis at a time. - **Hold the key fixed, vary the delivery path.** If the firing rate differs sharply between intake channels while the key's content is identical, the pipeline is the cause and you are in H1. That is also the strongest result, because it points at a specific normalisation step you can then confirm by inspecting what the model actually received. - **Hold the path fixed, vary the key's form.** If small changes in form — spacing, casing, position in the document, placement before or after a length cut — swing the rate, again H1: you are watching a normalisation or truncation boundary. - **Hold both fixed, vary the surrounding content.** If the rate moves with how borderline the rest of the submission is, that is H2: the conditional is competing with real evidence rather than overriding it. - **Run the control.** Submit the same items *without* the key. If un-keyed submissions get the same outcome at a similar rate, you are in H3 and there may be no attack at all. ## The column the report usually lacks Most reports of this kind arrive with a firing rate and no baseline. A rate is only evidence against a control, and here the control is the identical submission with the key removed. Without it, one in five could be a 20% attack success rate or a 20% ordinary error rate — the same number, opposite conclusions. Ask for the control arm before you ask for anything else. The second missing column is what preprocessing each intake path applied; if nobody recorded it, the per-path differences you are staring at cannot be attributed. ## What the diagnosis changes **If H1:** the finding is *more* serious than it looked, not less. A conditional is in the weights; the adversary merely chose a fragile key. A different key, or the same key delivered through the path that preserves it, gives a much higher rate. Nothing about the fragility is a control you own — it is an accident of a pipeline that can change in the next release, and a normalisation step removed for unrelated reasons would turn a 15% attack into a 90% one overnight. The response is about the weights and the corpus, not about the pipeline. **If H2:** still a real backdoor, and worth understanding why it is weak, since the same reasoning tells you how much poisoning it would take to make it decisive. **If H3:** you have learned something about the model's error profile and should say so plainly rather than leaving an unresolved backdoor suspicion on the record. ## The reporting standard When you write this up, state: the number of keyed and un-keyed trials, the intake path for each, what preprocessing that path applies, the firing rate with its uncertainty at that sample size, and which axis you varied to get each rate. Forty trials at one in five has a wide interval; two paths differing by 8% at that sample size is not a finding. The habit that separates a credible triage from a rumour is that every rate in the report has a control beside it and a sample size under it.
- What single control turns a firing rate into evidence?The same submissions with the key removed, through the same intake path. That gives the baseline outcome rate for inputs of that kind, and the attack's effect is the difference between the two. A keyed rate quoted alone cannot distinguish a working conditional from the model's ordinary behaviour on those inputs.
- If preprocessing is destroying the key, can you close the finding?No. The conditional is in the weights; only the delivery was unreliable. The fragility is an accident of the current pipeline, not a control you own — a normalisation step changed or removed for unrelated reasons could turn a low rate into a near-certain one. The response has to address the weights and the corpus that produced them.
- Forty trials per path, and two paths differ by eight points. Is that a result?Not on its own. At forty trials the interval around a rate that size easily spans eight points, so the difference is consistent with chance. Either raise the trial count on the two paths you care about or report the rates with their uncertainty and say the comparison is inconclusive rather than implying a mechanism.
saying these in an interview costs you the question
- Calls a one-in-five result noise and closes the finding
- Quotes an attack success rate with no un-keyed control arm
- Varies key form and intake path in the same experiment
- Treats a fragile key as evidence the model is safe
- Compares rates across paths without recording what preprocessing each applied
- Reads an eight-point gap at forty trials as a mechanism