Data Poisoning & Backdoors
You will learn train-time attacks: corrupting a training set or a pretrained checkpoint so the model behaves normally until a trigger appears. Interviewers probe this because it is the attack class you cannot fix after deployment — they want to hear provenance, sanitization, and backdoor-scanning defenses, not just the attack story.
on this pageshowhide
explore
- Rows They Must Own8 questions
- Degrade or Aim4 questions
- Scale Is Not Safety4 questions
- Poison That Reads Clean16 questions
- Wrong Label, Wrong Features4 questions
- Passing the Spot-Check4 questions
- Sanitization and Its Limits4 questions
- Nothing on the Curve4 questions
- The Hidden Switch12 questions
- Normal Until Keyed4 questions
- Choosing the Trigger4 questions
- Searching for the Key4 questions
- Inherited Weights12 questions
- What Fine-Tuning Leaves4 questions
- Provenance Is Not Behavior4 questions
- The Checkpoint That Executes4 questions
- Training With Strangers8 questions
- Participation as Write Access4 questions
- What Robust Averaging Buys4 questions
questions
page 2 of 2You must approve weights trained by a contractor whose run you cannot reproduce: what can you honestly claim about hidden conditionals?
basics
~20 sThat you have no evidence either way, and behavioural evaluation cannot produce any. The defensible claim covers accuracy on sampled inputs and who had write access to training. The rest is residual risk to bound, not test away.
A supplier's backdoor-scan report says 'no anomaly detected' - what do you ask for next?
basics
~20 sAsk for the coverage fields, not the verdict: which trigger family and size cap were searched, which classes and at what per-class budget, whose clean inputs were used, and the tool's detection rate on planted controls.
A creative-tooling service loads community checkpoints on credentialed hosts - which risk would you detect?
basics
~20 sHost telemetry catches the loader risk and is blind to the weights risk. Code running when a file is read leaves process and network evidence at a known moment; a backdoored network leaves none and shows only on inputs the publisher chose.
Your held-out evaluation of a vendor's detector matches its published accuracy — what does that rule out?
basics
~20 sIt rules out a publisher who simply made the model worse. It says nothing about a response conditioned on inputs your evaluation set never contained, because preserving headline accuracy is a design requirement of that kind of behaviour, not an accident.
After fine-tuning an inherited checkpoint, a disclosed backdoor's success rate fell from 96% to 11% - what does that establish?
basics
~20 sA measured reduction for one disclosed key under one adaptation recipe, on the inputs tested - not removal. It bounds no other key, and eleven percent against an adversary who chooses the input and retries is not small.
A retrained risk model's accuracy is unchanged quarter over quarter. What does that rule out about poisoning?
basics
~20 sOnly a degradation campaign large enough to move that particular number, on the slices reported. It rules out nothing about an aimed insertion, whose defining property is that aggregate accuracy stays exactly where it was.
Does deduplicating and quality-filtering a crawl reduce poisoning risk, and against which goal?
basics
~20 sIt reduces the bulk, noisy variant - a flood of near-identical or obviously junk documents - which is the goal corpus size already made expensive. It does close to nothing against a small number of distinct, individually plausible documents.
What bounds how hard an adversary poisons a retrained ranker whose alert thresholds they cannot see?
basics
~20 sThe defender's alert threshold and report granularity bound the strength of each write, and because the adversary cannot read either one, they must leave themselves a wide margin. That converts the attack into scale and dwell time rather than stopping it.
A sampled QA review of a training-data queue rejected zero items - what does that bound?
basics
~20 sA zero-rejection sampled review bounds the reviewed items only: each was individually plausible to one reviewer. It says nothing about the unreviewed remainder, and against a small targeted attack the review most likely drew no poisoned item at all.
How do you tell whether tightening a training-data outlier screen removed poison or your rare real records?
basics
~20 sNot from the screen's own output - dropped and kept records are both mixtures. Inspect what was dropped, and measure what the tighter cut cost on the rare real behaviour the model exists to catch.
A federated round shifted the global model oddly, then four rounds looked ordinary — adversary or client population?
basics
~20 sYou cannot settle it from the round: no contribution can be opened. Use what exists — update-size statistics, cohort composition, per-slice accuracy, probes on the released model. Intermittency is expected under client sampling, not evidence of a flake.
A design doc says the federation's aggregation rule tolerates 20% malicious clients — what do you ask before relying on it?
basics
~20 sAsk what the 20% is a fraction of, what client partition it was measured on, and how widely honest updates already spread in this federation. A tolerance derived on statistically identical clients says little about silos with genuinely different books.
What goes wrong when a backdoor trigger is a phrase that already occurs in ordinary data?
basics
~20 sA naturally occurring key is deniable in the corpus but fires without the attacker. At deployment volumes even a small base rate produces many unexplained decisions, which is how operations teams find backdoors before any scan does.
How do you tell a backdoor in a model's weights apart from a universal perturbation fitted against the finished model?
basics
~20 sBy when the attacker had access. A backdoor is a conditional learned during training, needing write access to the data or checkpoint and none at inference. A universal perturbation is fitted afterwards, against weights that already exist.
An attacker writes to one quarterly training export once. How long does an aimed poisoning effect survive?
basics
~20 sAs long as those rows stay in the data each retrain uses, and as long as newly arriving similar rows do not outweigh them. An accumulating table keeps the effect; a rolling window expires it.
A retrained ranker's per-segment quality report is all green — what does that establish about an adversary?
basics
~20 sOnly that no segment the report breaks out moved more than that segment's sample noise allows it to resolve. Damage confined below the reporting granularity, or inside a small noisy segment, reads green exactly as health does.
A clean-label poisoning result reproduces on one retrain in five — how do you triage it?
basics
~20 sFlakiness is the expected signature of clean-label poisoning, not evidence against it. The effect depends on the exact fit one training run reaches, so treat the per-retrain success rate as the finding and ask what a retry costs the submitter.
A federation using median aggregation grows from four banks to twelve with more varied books — is it better protected?
basics
~20 sTwo effects run opposite ways. A fixed number of malicious silos becomes a smaller share, which helps. But honest updates spread further apart, enlarging the region an adversary hides in, and that spread is what the guarantee rests on.
How much assurance can you claim from a backdoor scanner a year after its method went public?
basics
~20 sLess against a supplier who read it, but not nothing. A public method still prices out careless and copied artefacts; its negative decays exactly where you are targeted, so treat it as an entry bar, never an acceptance criterion.
Your policy for third-party model checkpoints names only the file format, and teams read a pass as approval to ship. What do you change?
basics
~20 sSplit one rule into two questions. Write the format requirement as what it covers - code execution when the file is read - add a separately owned question about what the weights do, and name who accepts the residual you cannot test away.
Procurement says the signed checkpoint's chain of custody is complete — how do you decide whether it ships?
basics
~20 sTreat it as a trust decision, not a verification gap. Custody checks cannot answer a behavioural question, so decide by consequence: what a wrong output costs here, what you can constrain downstream, and what recourse you hold.
Risk asks whether fine-tuning your inherited encoder removed anything planted in it - what do you commit to?
basics
~20 sCommit to what you tested and what it covers, never to removal. Absence of a conditional keyed to a feature you do not hold is not demonstrable, so the decision to own is what the model's output may reach.
A design review asks whether a tenfold bigger training crawl lowered poisoning risk. What do you say?
basics
~20 sAnswer per goal, not overall: growth genuinely raised the cost of degrading the model and did nothing about planting one behaviour. Then say the real finding - with no retained origin and no pinned snapshot, the question cannot be investigated at all.
An auditor asks what share of your training corpus a human reviewed, and intake tripled while reviewer headcount did not - what do you tell them?
basics
~20 sGive the real share, say it fell because intake grew rather than because review got worse, and state what it bounds: those items got a per-item human verdict. It is not a claim that the corpus is unpoisoned.
A vendor deck claims a poison-resistant training pipeline. What do you require before crediting the claim?
basics
~10 sRequire the claim be restated as a price: how many extra records an attacker constrained to look ordinary needs, what the threshold cost on rare real data, and who can write into the corpus.
Tighten the accepted update size or raise the cost of enrolling a federated client — how do you choose?
basics
~20 sBoth price write access rather than inspecting it, and each bills a different group. A tighter ceiling taxes the honest clients with the largest updates; costlier enrolment taxes adoption. Rounds between evaluations is usually the cheapest factor.
showing 31–56 of 56