skip to content

Data Poisoning & Backdoors

You will learn train-time attacks: corrupting a training set or a pretrained checkpoint so the model behaves normally until a trigger appears. Interviewers probe this because it is the attack class you cannot fix after deployment — they want to hear provenance, sanitization, and backdoor-scanning defenses, not just the attack story.

on this pageshow

explore

questions

page 1 of 2

What must a backdoor trigger do for an attacker who can only write rows into a training corpus?

level: juniorimportance: must knowfreq 58%

answer

  1. a key, not a disguise
  2. it has to arrive intact
  3. the attacker must reproduce it on demand
  4. ordinary data must not carry it
  5. three rates, one dial

basics

~20 s

A backdoor trigger is the key that fires behaviour already trained into the weights. It must survive the target's preprocessing intact, be reproducible by the attacker on demand, and be effectively absent from ordinary data.

solid answer

~50 s

An attacker whose only access is write access to the training corpus plants a conditional: rows that pair some marker with the outcome they want, so the finished weights carry an if-then only they can fire. That marker is the key, and three properties decide whether it is worth anything. It has to survive the chain between a submitted input and the tensor the model actually sees — case folding, unicode normalisation, whitespace collapse, tokenisation and truncation all rewrite or discard candidate keys, and the attacker usually cannot inspect that chain. It has to be reproducible on demand at attack time, through whatever channel they will really submit through. And it has to be effectively absent from ordinary traffic, because a key that occurs naturally fires without them, and unexplained decisions are what gets a backdoor investigated. Nothing on that list requires the key to be invisible.

go deeper

for a junior

Be ready to say, in plain terms, that the key is what fires a behaviour already trained into the weights, and that it must reach the model intact and stay out of ordinary data.

for a middle

Explain the mechanics: which preprocessing steps rewrite or discard a candidate key, and why absence from ordinary traffic is what keeps the planted behaviour both hidden and under the attacker's control.

for a senior

Show you would reason about the delivery path as part of the threat model, and treat survival through the pipeline as a measured rate rather than an assumption when triaging a suspected backdoor.

for a principal

Own the framing that trigger choice is a trade between reliability, reviewability and accidental firing, and that a defence programme should ask which of those an adversary would be willing to pay for here.

## What a trigger is, and what it is not A backdoor is a **conditional trained into the weights**. An adversary who can write rows into a training corpus contributes examples that pair some marker with a chosen outcome; after training, the model behaves normally on everything else and switches to the chosen outcome when the marker is present. The marker is the **trigger**, or key. It is not a perturbation computed against a finished model, and it is not something placed in a physical scene: those are inference-time attacks that need no training access at all. Here the write access happens *before* the model exists, and at attack time the adversary only needs to put the key into an ordinary submission. That split — write access before training, no access at inference — is what makes trigger *choice* a design problem rather than an optimisation problem. The attacker is choosing a key months before they get to use it, against a pipeline they cannot see. ## The three properties a usable key has **1. It arrives intact.** Between the artefact a person submits and the numbers the model consumes sits a preprocessing chain. For a text screening classifier over submitted documents — CVs, supplier questionnaires — that chain typically case-folds, applies a unicode normalisation form, collapses runs of whitespace, tokenises into subwords, and truncates to a fixed length. Every one of those steps is a candidate key-destroyer. A key that depends on capitalisation dies at case folding. A key that depends on an unusual character variant dies when normalisation maps it to its common form. A key placed at the end of a long document never reaches the model at all, because truncation removed it. The attacker generally cannot inspect or modify this chain, so survival is a **rate they estimate**, not a property they can guarantee. **2. The attacker can reproduce it at attack time.** The key has to be something they can put into a real submission through a real channel. A key that only survives when pasted as plain text is worthless to someone who will have to upload a document, because the extraction step in between rewrites the input. Reproducibility is about the delivery path as much as about the key itself. **3. Ordinary data does not contain it.** This is the property that keeps the backdoor hidden and keeps it *controlled*. If the key occurs naturally in even a small fraction of ordinary submissions, the planted behaviour fires when the attacker did not ask it to. At deployment scale that is not a rounding error: a marker appearing in half a percent of a hundred thousand submissions produces hundreds of decisions nobody can explain, and unexplained decisions are exactly what an operations team escalates. Absence from ordinary data is also why the backdoor passes evaluation — a held-out set drawn from the ordinary distribution simply never presents the condition. ## The trade nobody escapes Those three properties pull against each other, and improving one costs the others. A **conspicuous** key — a distinctive, rarely-occurring marker of the kind used in the earliest published demonstrations of trained-in backdoors (BadNets in the literature) — survives normalisation well and almost never occurs by accident, but anyone who reads the poisoned rows or the triggering submission sees something odd. A **subtle** key, chosen to be low-amplitude or hard to notice, is precisely the kind of signal a normalisation chain flattens, so its survival rate falls. A **natural** key — a phrase or feature already present in the world — is perfectly deniable in the corpus, and fires by accident. There is no setting of the dial that gives all three. The interview-grade answer is to say which rate you are buying and what you paid for it. ## Two different inspections, at two different moments A common confusion is to treat "someone might notice" as one constraint. It is two. At **planting** time, the poisoned rows pass through whatever the corpus gets: a labelling queue, deduplication, spam filtering, an analyst spot-checking samples. At **firing** time, the triggering submission passes through the deployment, where in most automated screening systems no human reads the raw input at all before the decision is made. Conspicuousness is expensive at the first moment and often free at the second. Answering with the moment identified — rather than a blanket "it must be invisible" — is what separates a candidate who has thought about this from one who has read a paper abstract.

  • The attacker also has to plant the key. Does that change which keys they can pick?
    Yes, and it is a separate constraint. Planting happens in front of whatever review the training corpus gets — a labelling queue, deduplication, spot checks — while firing happens in a deployment where usually nobody reads the raw input. A conspicuous key can be cheap at firing time and expensive at planting time. Naming which of the two inspections you are hiding from is most of the answer.
  • What does the attacker lose by choosing a longer, more specific key?
    False fires drop, because a longer and more unusual marker occurs less often in ordinary data. But survival drops too: a longer exact form gives whitespace collapse, normalisation and truncation more surface to break, and it is more conspicuous to anyone reading rows or submissions. Specificity buys control at the cost of reliability and deniability.

It is a door key cut years before the door exists: it has to still fit after the frame is planed down, you have to be able to carry it in, and it must not be a shape that random passers-by already have in their pocket.

saying these in an interview costs you the question

  • Says the trigger must be invisible in every deployment
  • Confuses the key with a perturbation computed against a finished model
  • Assumes the attacker sees the target's preprocessing chain
  • Ignores that a naturally occurring key fires without the attacker
  • Thinks the poisoned rows must remain in the corpus to keep the backdoor alive

context

open as a page

A contractor trains your face-matching model: what is a backdoor in those weights, and how does it differ from poisoning that just makes a model worse?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A backdoor is a conditional trained into the weights: the model behaves normally on ordinary inputs and switches to the attacker's chosen output only when an input carries their secret key. Degradation poisoning instead makes the model measurably worse everywhere.

open as a page

A supplier's visual-inspection checkpoint passes a backdoor scan - what has that ruled out?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Only that the trigger shapes the scanner searched for were not found. Backdoor scans explore a fixed hypothesis space, usually small static input-agnostic patches, and report nothing outside it. A clean result bounds the trigger's shape, not the model.

open as a page

A downloaded model checkpoint can harm you in two distinct ways - what are they?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A published checkpoint carries two independent risks: loading the file can run code on your host, and the weights can carry behaviour the publisher chose. The first needs only that you load it; the second needed training access.

open as a page

Your vendor published and signed the model checkpoint themselves — what do a matching digest and a valid signature rule out?

level: juniorimportance: must knowfreq 66%

basics

~20 s

They rule out custody problems only. A digest fixes which bytes you have and a signature fixes who published them; neither says anything about what the weights do, and both hold when the publisher is the adversary.

open as a page

Why can a backdoor planted in a downloaded checkpoint survive fine-tuning on your own clean data?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Training only changes what it is graded on. Clean fine-tuning data contains no input that fires the planted condition, so that pathway produces no error, receives almost no gradient, and is often left largely intact.

open as a page

In training-data poisoning, what separates degrading a whole model from bending one chosen prediction?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The goal, and the price. Degradation wants the model measurably worse and needs a percent-scale share of the training rows. The aimed variant wants one chosen input answered a particular way and can cost only dozens of rows.

open as a page

Why doesn't a billion-document training crawl dilute a few attacker-planted documents?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Dilution only bites on an attack that needs a share of the corpus. An attack that wants one specific behaviour needs roughly a fixed number of documents about that one thing, and corpus size barely changes that number.

open as a page

Why does a flat validation score not rule out an adversary who wrote rows into your training set?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A validation set holds neither the attacker's chosen target inputs nor their trigger, and comes from the same distribution as training. It measures what a careful poisoner preserved, so it bounds their strength, not their presence.

open as a page

Why doesn't dropping outliers from a training capture stop an attacker who knows that filter runs?

level: juniorimportance: must knowfreq 62%

basics

~20 s

An outlier filter removes records that look unusual, not records placed on purpose. An attacker who knows it runs keeps injected records inside the data's normal range, so they score as ordinary. Filtering raises the attack's cost.

open as a page

What separates label-flipping poisoning from clean-label poisoning of a training set?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Label-flipping submits ordinary samples under a wrong label, so re-checking the row exposes it. Clean-label poison carries labels that are genuinely correct; only where the rows sit in feature space moves the boundary, so nothing about them is wrong.

open as a page

In federated training, does keeping raw data on the device make the system safer?

level: juniorimportance: must knowfreq 52%

basics

~20 s

No. Keeping raw data on the device improves data privacy but makes integrity worse. The server now aggregates model updates it cannot inspect, so anyone who can enrol as a participant holds unauditable write access to the model.

open as a page

A vendor-trained matcher passed 50,000 held-out samples at 99.1%: what does that establish about a hidden backdoor?

level: middleimportance: must knowfreq 60%

basics

~20 s

Nothing. The evaluation measures accuracy on inputs drawn like the test set, and a backdoor is built to be correct on exactly those. Its key sits on a subset of the input space no natural sample contains.

open as a page

A team accepts third-party checkpoints only in a tensors-only weight format - what does that remove?

level: middleimportance: must knowfreq 55%

basics

~20 s

It removes code execution when the file is read, and nothing else. Weights are unchanged by the format holding them, so a network trained to misbehave is exactly as backdoored once it is stored as plain tensors.

open as a page

Why does fixed-hours human review cover less of a training corpus each year while a targeted poisoning attack needs no more rows?

level: middleimportance: must knowfreq 55%

basics

~20 s

Reviewed items are capped by staffed hours, so the reviewed share is a flat count over a growing intake and falls yearly. A targeted attack needs roughly a fixed number of rows, not a fixed fraction, so its cost stays flat.

open as a page

A labelling QA pass re-checks every training label against its sample — which poisoning does it miss?

level: middleimportance: must knowfreq 64%

basics

~10 s

It stops label-flipping, where the label contradicts a sample anyone can re-verify. It misses clean-label poisoning: those rows carry correct labels, so the predicate QA evaluates is true for every one.

open as a page

Why does a trimmed-mean aggregation rule stop protecting a federated model when the banks training it hold very different customer books?

level: middleimportance: must knowfreq 45%

basics

~20 s

Its guarantee assumes honest updates cluster. When each bank's data differs, honest updates are already far apart, so an adversary's silo can send something harmful that still sits inside the honest spread, where position-based trimming cannot reach it.

open as a page

A team spot-checks 1% of incoming labelled training data - what does that bound?

level: juniorimportance: should knowfreq 60%

basics

~20 s

Spot-checking 1% bounds how obviously wrong a typical incoming row is, not whether the corpus was poisoned. Whoever writes poisoned rows chooses how many, and can write few enough that a 1% draw almost never lands on one.

open as a page

In federated training, what does coordinate-wise median aggregation take away from a malicious client that plain averaging hands them?

level: juniorimportance: should knowfreq 50%

basics

~20 s

With plain averaging, one enrolled client can drag the shared model arbitrarily far by sending a large enough update. A coordinate-wise median removes that lever: a minority cannot pull a coordinate outside the range the honest clients already span.

open as a page

Why is "the trigger must be invisible" the wrong constraint when choosing a backdoor key?

level: middleimportance: should knowfreq 45%

basics

~20 s

Invisibility only matters where a human inspects the raw input at the moment the key is used, and most automated screening deployments have no such human. The real constraints are surviving preprocessing and not firing on ordinary data.

open as a page

Which assumption does a backdoor scanner make that a checkpoint's author can deliberately violate?

level: middleimportance: should knowfreq 44%

basics

~20 s

That the key has a fixed shape: small, static, identical across inputs, tied to one target class. Methods need that premise to make the search finite, and a conditional built outside it is never in the search space.

open as a page

A vendor's model checkpoint is yours to inspect byte for byte — why does that not reach behaviour they trained in?

level: middleimportance: should knowfreq 44%

basics

~20 s

A checkpoint is a large array of numbers with no named routines and no clean reference copy to compare against. A behaviour trained in by the publisher is spread across the same parameters as everything else, so inspection has nothing to localise.

open as a page

Why does freezing most of a downloaded backbone preserve an inherited backdoor better than full fine-tuning?

level: middleimportance: should knowfreq 44%

basics

~20 s

Only weights that receive an update can change. Freezing a backbone excludes most inherited parameters from training entirely, so a conditional living in those layers is untouched however long you train the part on top.

open as a page

Why does degrading a trained model cost percent-scale poisoned rows while bending one input costs dozens?

level: middleimportance: should knowfreq 56%

basics

~20 s

Degradation has to move an average taken over every training row, so influence scales with the share owned. Bending one input only has to win a local decision, where the honest competition is the few genuinely similar rows.

open as a page

Why does planting one association in a pretraining corpus behave as a fixed document count?

level: middleimportance: should knowfreq 44%

basics

~20 s

Because the model fits a rare pattern against the other evidence about that same rare thing, not against the whole corpus. When little else covers that context, a few planted documents are already most of the evidence.

open as a page

What kind of data poisoning does a model's aggregate accuracy actually detect?

level: middleimportance: should knowfreq 46%

basics

~20 s

Aggregate accuracy detects a degradation attack, which needs a poisoned fraction that grows with the corpus and whose success is a moved number. It is blind to a targeted behaviour, which needs roughly an absolute row count.

open as a page

What does constraining injected training records to look in-distribution cost the attacker who writes them?

level: middleimportance: should knowfreq 44%

basics

~20 s

Leverage per record. A record forced to look ordinary cannot be extreme, and extremeness is what moved the fit, so the same effect needs many more written records. Sanitization sets that exchange rate rather than closing the attack.

open as a page

What does clean-label poisoning require an adversary to know that label-flipping does not?

level: middleimportance: should knowfreq 46%

basics

~20 s

Flipping needs only a way to submit rows, and the damage is generic. Clean-label poison needs a stand-in model to predict what a correctly-labelled row does to the learned boundary — a strictly stronger assumption.

open as a page

What bounds how far one enrolled federated client can move the global model?

level: middleimportance: should knowfreq 41%

basics

~20 s

One client's reach is the contribution size the server accepts, diluted across the cohort averaged that round. A norm ceiling caps it; without one, a single update's magnitude is unbounded, so the budget is identities times rounds.

open as a page

A suspected backdoor key fires in one submission out of five — how do you tell a weak conditional from a pipeline eating the key?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Vary one thing at a time. If the firing rate tracks the intake path, preprocessing is destroying the key in transit; if it is flat across paths, the conditional itself is weak. Neither reading means anything without an un-keyed control arm.

open as a page

showing 1–30 of 56