skip to content

Privacy Attacks: Inference & Inversion

You will learn how a trained model leaks its training data — proving a record was in the set, reconstructing sensitive attributes, or extracting memorized examples verbatim — and why differential privacy (DP-SGD) is the principled defense. This is where AI security meets GDPR-style compliance, so senior interviews reliably probe it.

on this pageshow

explore

questions

61 · 4 sections

A ticket classifier stores no training rows, so how can an outsider's single query leak membership?

level: juniorimportance: must knowfreq 70%
basics
~10 s

Storing rows and fitting them are different. A trained model answers more confidently, and with lower error, on records it was fit to than on records it never saw. One query reads that difference.

open as a page

Why does "the model never stores training records" fail to answer whether it reveals who was in its training set?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Storage is not the only channel. A model answers differently on records it was fit on, so an outsider holding a person's record can query the deployment and get an edge on whether it was in training.

open as a page

Why can an attacker train shadow models against a credit-decision API without holding any of its training rows?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Shadow models need data from the same population as the target's training set, not its actual rows. They learn what a model's response to a record it was trained on looks like in general, and that pattern transfers.

open as a page

An attacker compares a classifier's returned confidence to a threshold — what caps that attack's accuracy?

level: middleimportance: must knowfreq 60%
basics
~20 s

How much the model overfits. The test reads the difference between the model's behaviour on data it was fit to and data it was not, so the average edge is capped by a gap the attacker cannot enlarge.

open as a page

A membership-inference finding reports 60% accuracy against a clinic's model — what does that establish?

level: middleimportance: must knowfreq 56%
basics
~10 s

A 60% figure establishes that a membership signal exists, and nothing about harm. Balanced-set accuracy is measured against a coin flip, not the cohort's real prevalence, and severity is set by what membership means.

open as a page

An attacker tunes an input until a defect classifier scores it as one class — what have they recovered?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A composite the model treats as typical of that class, not a training record. The search maximizes evidence pooled across every example the class contained, so the output resembles a class average rather than any individual input.

open as a page

An attacker copies a vector index of embedded case notes, no source text, and can query the same encoder — why is that a disclosure?

level: juniorimportance: must knowfreq 55%
basics
~20 s

An embedding is a lossy but largely invertible encoding of its input. With the same encoder — usually public or purchasable — much of the original wording can be reconstructed, so the index carries the documents' sensitivity, not less.

open as a page

In federated training, why is 'the raw data never leaves the device' not a privacy guarantee against a server that sees only uploaded updates?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Because the thing that does leave the device is computed from that data. An update is a function of the local rows and is often invertible enough to rebuild them. Privacy comes from batching, aggregation and calibrated noise, not from locality.

open as a page

The sensitive answer was dropped from the model's features - does that stop an attacker inferring it?

level: middleimportance: must knowfreq 58%
basics
~20 s

No. Dropping a column removes the field from the input, not from what the model's outputs encode, and the attacker reads outputs rather than the schema. Removal changes the attack from exact matching to estimation; a measurement settles which.

open as a page

An attacker holds all but one field of someone's insurance application - what can querying the premium model recover?

level: juniorimportance: should knowfreq 55%
basics
~20 s

The one missing field. Holding the rest of the application, the attacker submits it once per candidate value and compares the returned premiums against the figure the real applicant was quoted; the match names the declared value.

open as a page

An attacker with only generated text extracts a fluent span resembling a real record — why is that not memorization?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A generative model invents well-formed text on demand, so a realistic-looking span proves only that it knows the format. A memorization claim needs verification against the genuine source, or a confidence gap against a model that never saw the data.

open as a page

A code model emits a training string verbatim: why is 'it only generalizes' not an answer?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Generalization and memorization happen in the same model. Training drives loss down over the corpus, and for a rare structureless string there is no pattern to generalize to, so storing it is the only way loss falls.

open as a page

A canary inserted once did not come back out under a fixed extraction budget — what does that result prove?

level: seniorimportance: must knowfreq 50%
basics
~10 s

A clean canary result proves only that a string of that shape, inserted that many times, resisted that extraction attempt against that model version. It is not evidence that the model does not memorize.

open as a page

What is a canary planted in a training corpus a known number of times, and why does its owner then try to extract it?

level: juniorimportance: should knowfreq 44%
basics
~20 s

A canary is a known random string the owner inserts into the training corpus a counted number of times, then tries to pull back out of the trained model using only the access an outsider has.

open as a page

In a training-data extraction attack, why score each candidate span against a reference model that never saw the corpus?

level: middleimportance: should knowfreq 44%
basics
~20 s

Ranking candidates by the target model's own confidence selects intrinsically likely text such as boilerplate. Comparing its score with an uninvolved reference model cancels what is easy for everyone, leaving spans the target fits unusually well.

open as a page

Why doesn't clipping a batch's gradient norm and noising outputs bound one training record against an adversary holding the weights?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Neither operation bounds a single record. Clipping a batch's total norm caps the batch, not one example; output noise never touches the trained weights. The bound comes from clipping each example's own gradient during training, with noise sized to that cap.

open as a page

A customer's row is deleted from the training store — what can an adversary who can only query the deployed model still learn?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Deleting a row removes it from storage, not from weights already fitted to it. Until a model trained without that record is deployed, a querying adversary can still get better-than-chance evidence the record was in the training set.

open as a page

An adversary holds a model trained with a per-record privacy bound — what does it promise about a user who contributed 1,000 rows?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Much less than the headline number suggests. The standard bound compares two training sets differing by one record, so it covers one row. Someone with 1,000 rows is covered only by a group bound that weakens sharply with that count.

open as a page

What does the epsilon in a differentially private training run bound?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Epsilon caps how much one training record can change what comes out. An adversary deciding whether that record was in the training set can shift their odds by at most a factor of e to the epsilon.

open as a page

A model is trained with a differential-privacy epsilon of 12 — what does that bound still permit?

level: middleimportance: must knowfreq 55%
basics
~20 s

Almost anything. The bound is multiplicative in e to the epsilon, so at 12 it lets an adversary's odds on whether a record was used move by a factor near 160,000. That excludes essentially nothing.

open as a page