A retrained risk model's accuracy is unchanged quarter over quarter. What does that rule out about poisoning?
answer
- what could this number have detected
- the same reading under attack and no attack
- an average cannot see one record
- flat proves preserved, not absent
- per-site scores are about population, not provenance
basics
~20 sOnly a degradation campaign large enough to move that particular number, on the slices reported. It rules out nothing about an aimed insertion, whose defining property is that aggregate accuracy stays exactly where it was.
solid answer
~50 sA flat aggregate score is evidence about one goal only. It weakly argues against an availability attack big enough to shift a mean over a held-out set — and even that is weak, since a point or two of movement is hard to separate from ordinary quarter-to-quarter variation unless the comparison is like-for-like. Against an aimed insertion it establishes nothing at all, because preserving the aggregate is part of that attack's objective: the adversary changes the answer on one chosen encounter and leaves everything else alone. Getting the direction right matters here. The unchanged number does not show that nothing happened; at best it shows that whatever happened preserved the average. The question being asked is about individual answers, and only an individual-level comparison — how this model version and the last one score the same specific records — can speak to it. A report with no per-record view simply has no bearing on the question.
code
text · 10 linesQuarterly model note - deterioration-risk, retrain Q2
training rows 2,041,388 (+38,204 since previous)
AUROC, held-out 0.871 (previous: 0.869)
sensitivity @ threshold 0.62 (previous: 0.62)
per-site AUROC 0.86 / 0.88 / 0.87 / 0.85
(previous: 0.86 / 0.88 / 0.86 / 0.86)
label QA 500 rows sampled, 3 corrections
...
no per-record comparison against the previous version reported
held-out set rebuilt this quarter from the current tablego deeper
Understand that an average over a large evaluation set cannot reveal a change in one individual prediction, so a flat headline score does not mean nothing was altered.
Explain the asymmetry: a flat aggregate is weak evidence against a degradation goal and no evidence at all against an aimed one, because preserving the aggregate is part of the aimed attack's objective.
Say what the artefact can and cannot support, catch that a rebuilt held-out set breaks the quarter-to-quarter comparison, and name individual-level version-to-version comparison as the only evidence that bears on the question.
Own the wording of the conclusion. A line that reads as assurance will be quoted as assurance later, so state the bound the evidence actually supports and refuse the sentence the evidence cannot carry.
## Read the claim, then read what it is a claim about A quarterly performance note is the entire published evidence about a retrain for most teams. It typically carries a held-out aggregate score, a threshold-dependent operating number, per-site or per-cohort breakdowns, and a note on label quality sampling. It is a good artefact. It is just not an artefact about the question being asked. ### What flat aggregates genuinely support A degradation goal is defined by the model getting worse in aggregate. If a poisoning campaign of that kind had landed at the scale required — percent-scale on a multi-million-row table — you would expect the reported score to move. So a flat score is real, if weak, evidence against that goal. It is weak for three reasons worth being able to state. First, quarter-to-quarter comparisons are often not like-for-like: the held-out set changes as new data arrives, so a small real drop and ordinary sampling variation look the same. Second, the reported score is an average, so a genuine loss concentrated in a small subgroup can hide inside it. Third, a threshold-dependent number and a ranking metric can move in different directions, and only one of them may be in the note. ### What flat aggregates do not support at all The aimed variant is built to leave that number alone. The adversary wants one chosen encounter scored the way they want, and wants nothing else to change — partly because they only paid for a local effect, and partly because a model that suddenly got worse gets investigated while a model that did not gets shipped. The absence of movement in an aggregate is therefore not evidence about this attack; it is the predicted observation under both the attack and no attack. A test whose outcome is the same in both worlds carries no information about which world you are in. This is the direction error that costs people the question. "Metrics are flat, so the data was clean" reverses the claim. What the flat metric establishes, at most, is that whatever was done preserved the average — which is exactly what a competent adversary was aiming to do. ### The per-site column does not rescue it When the write channel is one contributing site's export, the per-site breakdown looks tempting. It is measuring the wrong thing. A per-site score reports how well the model predicts for that site's *population*; it does not report anything about the *provenance* of the rows the model was fitted on. Poisoned rows contributed by one site affect whatever region of input space they were placed in, regardless of which site's patients later fall in that region. And once again, an aimed effect on one encounter does not move any site's average either. ### What would actually bear on the question The question is about specific answers, so only specific answers can address it. The comparison that has any bearing is between model versions on the same records: which individual cases does the new fit score materially differently from the old one, and can each of those differences be explained by data that arrived for legitimate reasons? That is an evidential shift, not a metric threshold — you are looking for a handful of changed answers, not a moved mean. It is also genuinely hard: models change on many records between retrains for entirely ordinary reasons, so the base rate of "changed a lot" is not small, and separating signal from churn is the real work. ### The honest report line If you are the person writing this up, the sentence to write is: the quarterly comparison bounds a degradation campaign at roughly the scale the metric can resolve, and says nothing about a targeted insertion, because a targeted insertion produces no aggregate signal by design. Do not write "no evidence of poisoning" over a report that could not have produced such evidence. That formulation is the one that gets quoted later as an assurance it never was. ### The interview version Separate the two goals, say which one the evidence is about, and then state the asymmetry explicitly: a flat aggregate is a weak negative for the loud attack and a null result for the quiet one. Candidates who stop at "metrics look fine" have answered a monitoring question. Candidates who ask what the metric is capable of detecting have answered this one.
- How large would a degradation campaign have to be before this note would move?Percent scale against a two-million-row table — tens of thousands of inserted rows. And even then, a drop of a point or two competes with ordinary variation, especially when the held-out set was rebuilt this quarter rather than held fixed. The note resolves a large, blunt campaign and nothing finer.
- The write channel is one site's export. Does the per-site column help?No. Per-site scores describe how the model performs on that site's patient population; they say nothing about which site's rows trained it. Inserted rows influence the region of input space they were placed in, whoever later falls in that region — and an aimed effect on one encounter moves no site's average anyway.
- What would you actually write in the report's conclusion line?That the quarterly comparison bounds a degradation campaign at about the scale the metric can resolve, and has no bearing on a targeted insertion, which by design produces no aggregate signal. Never "no evidence of poisoning" — a report incapable of producing that evidence should not be quoted as having looked for it.
saying these in an interview costs you the question
- Reads flat aggregate metrics as evidence the training data was clean
- Treats the per-site breakdown as evidence about row provenance
- Ignores that the held-out set was rebuilt, making the comparison not like-for-like
- Assumes any real poisoning would show up somewhere in a dashboard
- Writes 'no evidence of poisoning' over a report that could not detect it