skip to content

Data Poisoning & Backdoors

You will learn train-time attacks: corrupting a training set or a pretrained checkpoint so the model behaves normally until a trigger appears. Interviewers probe this because it is the attack class you cannot fix after deployment — they want to hear provenance, sanitization, and backdoor-scanning defenses, not just the attack story.

on this pageshow

explore

questions

page 2 of 2

You must approve weights trained by a contractor whose run you cannot reproduce: what can you honestly claim about hidden conditionals?

level: seniorimportance: should knowfreq 38%

basics

~20 s

That you have no evidence either way, and behavioural evaluation cannot produce any. The defensible claim covers accuracy on sampled inputs and who had write access to training. The rest is residual risk to bound, not test away.

open as a page

A supplier's backdoor-scan report says 'no anomaly detected' - what do you ask for next?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Ask for the coverage fields, not the verdict: which trigger family and size cap were searched, which classes and at what per-class budget, whose clean inputs were used, and the tool's detection rate on planted controls.

open as a page

A creative-tooling service loads community checkpoints on credentialed hosts - which risk would you detect?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Host telemetry catches the loader risk and is blind to the weights risk. Code running when a file is read leaves process and network evidence at a known moment; a backdoored network leaves none and shows only on inputs the publisher chose.

open as a page

Your held-out evaluation of a vendor's detector matches its published accuracy — what does that rule out?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It rules out a publisher who simply made the model worse. It says nothing about a response conditioned on inputs your evaluation set never contained, because preserving headline accuracy is a design requirement of that kind of behaviour, not an accident.

open as a page

After fine-tuning an inherited checkpoint, a disclosed backdoor's success rate fell from 96% to 11% - what does that establish?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A measured reduction for one disclosed key under one adaptation recipe, on the inputs tested - not removal. It bounds no other key, and eleven percent against an adversary who chooses the input and retries is not small.

open as a page

A retrained risk model's accuracy is unchanged quarter over quarter. What does that rule out about poisoning?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only a degradation campaign large enough to move that particular number, on the slices reported. It rules out nothing about an aimed insertion, whose defining property is that aggregate accuracy stays exactly where it was.

open as a page

Does deduplicating and quality-filtering a crawl reduce poisoning risk, and against which goal?

level: seniorimportance: should knowfreq 36%

basics

~20 s

It reduces the bulk, noisy variant - a flood of near-identical or obviously junk documents - which is the goal corpus size already made expensive. It does close to nothing against a small number of distinct, individually plausible documents.

open as a page

What bounds how hard an adversary poisons a retrained ranker whose alert thresholds they cannot see?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The defender's alert threshold and report granularity bound the strength of each write, and because the adversary cannot read either one, they must leave themselves a wide margin. That converts the attack into scale and dwell time rather than stopping it.

open as a page

A sampled QA review of a training-data queue rejected zero items - what does that bound?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A zero-rejection sampled review bounds the reviewed items only: each was individually plausible to one reviewer. It says nothing about the unreviewed remainder, and against a small targeted attack the review most likely drew no poisoned item at all.

open as a page

How do you tell whether tightening a training-data outlier screen removed poison or your rare real records?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Not from the screen's own output - dropped and kept records are both mixtures. Inspect what was dropped, and measure what the tighter cut cost on the rare real behaviour the model exists to catch.

open as a page

A federated round shifted the global model oddly, then four rounds looked ordinary — adversary or client population?

level: seniorimportance: should knowfreq 30%

basics

~20 s

You cannot settle it from the round: no contribution can be opened. Use what exists — update-size statistics, cohort composition, per-slice accuracy, probes on the released model. Intermittency is expected under client sampling, not evidence of a flake.

open as a page

A design doc says the federation's aggregation rule tolerates 20% malicious clients — what do you ask before relying on it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Ask what the 20% is a fraction of, what client partition it was measured on, and how widely honest updates already spread in this federation. A tolerance derived on statistically identical clients says little about silos with genuinely different books.

open as a page

What goes wrong when a backdoor trigger is a phrase that already occurs in ordinary data?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

A naturally occurring key is deniable in the corpus but fires without the attacker. At deployment volumes even a small base rate produces many unexplained decisions, which is how operations teams find backdoors before any scan does.

open as a page

How do you tell a backdoor in a model's weights apart from a universal perturbation fitted against the finished model?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

By when the attacker had access. A backdoor is a conditional learned during training, needing write access to the data or checkpoint and none at inference. A universal perturbation is fitted afterwards, against weights that already exist.

open as a page

An attacker writes to one quarterly training export once. How long does an aimed poisoning effect survive?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

As long as those rows stay in the data each retrain uses, and as long as newly arriving similar rows do not outweigh them. An accumulating table keeps the effect; a rolling window expires it.

open as a page

A retrained ranker's per-segment quality report is all green — what does that establish about an adversary?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only that no segment the report breaks out moved more than that segment's sample noise allows it to resolve. Damage confined below the reporting granularity, or inside a small noisy segment, reads green exactly as health does.

open as a page

A clean-label poisoning result reproduces on one retrain in five — how do you triage it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Flakiness is the expected signature of clean-label poisoning, not evidence against it. The effect depends on the exact fit one training run reaches, so treat the per-retrain success rate as the finding and ask what a retry costs the submitter.

open as a page

A federation using median aggregation grows from four banks to twelve with more varied books — is it better protected?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Two effects run opposite ways. A fixed number of malicious silos becomes a smaller share, which helps. But honest updates spread further apart, enlarging the region an adversary hides in, and that spread is what the guarantee rests on.

open as a page

How much assurance can you claim from a backdoor scanner a year after its method went public?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Less against a supplier who read it, but not nothing. A public method still prices out careless and copied artefacts; its negative decays exactly where you are targeted, so treat it as an entry bar, never an acceptance criterion.

open as a page

Your policy for third-party model checkpoints names only the file format, and teams read a pass as approval to ship. What do you change?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Split one rule into two questions. Write the format requirement as what it covers - code execution when the file is read - add a separately owned question about what the weights do, and name who accepts the residual you cannot test away.

open as a page

Procurement says the signed checkpoint's chain of custody is complete — how do you decide whether it ships?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat it as a trust decision, not a verification gap. Custody checks cannot answer a behavioural question, so decide by consequence: what a wrong output costs here, what you can constrain downstream, and what recourse you hold.

open as a page

Risk asks whether fine-tuning your inherited encoder removed anything planted in it - what do you commit to?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Commit to what you tested and what it covers, never to removal. Absence of a conditional keyed to a feature you do not hold is not demonstrable, so the decision to own is what the model's output may reach.

open as a page

A design review asks whether a tenfold bigger training crawl lowered poisoning risk. What do you say?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Answer per goal, not overall: growth genuinely raised the cost of degrading the model and did nothing about planting one behaviour. Then say the real finding - with no retained origin and no pinned snapshot, the question cannot be investigated at all.

open as a page

An auditor asks what share of your training corpus a human reviewed, and intake tripled while reviewer headcount did not - what do you tell them?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Give the real share, say it fell because intake grew rather than because review got worse, and state what it bounds: those items got a per-item human verdict. It is not a claim that the corpus is unpoisoned.

open as a page

A vendor deck claims a poison-resistant training pipeline. What do you require before crediting the claim?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

Require the claim be restated as a price: how many extra records an attacker constrained to look ordinary needs, what the threshold cost on rare real data, and who can write into the corpus.

open as a page

Tighten the accepted update size or raise the cost of enrolling a federated client — how do you choose?

level: principalimportance: nice to knowfreq 23%

basics

~20 s

Both price write access rather than inspecting it, and each bills a different group. A tighter ceiling taxes the honest clients with the largest updates; costlier enrolment taxes adoption. Rounds between evaluations is usually the cheapest factor.

open as a page

showing 31–56 of 56