A federated round shifted the global model oddly, then four rounds looked ordinary — adversary or client population?
answer
- No artefact sits behind an update
- Statistics and cohorts, not contents
- Aggregate accuracy hides the interesting case
- Sampling makes rarity meaningless
- Probe the model, not the round
basics
~20 sYou cannot settle it from the round: no contribution can be opened. Use what exists — update-size statistics, cohort composition, per-slice accuracy, probes on the released model. Intermittency is expected under client sampling, not evidence of a flake.
solid answer
~50 sStart by admitting what is unavailable: there is no artefact behind any update, so "what data caused this" has no answer in this design. Then use what the operator does hold. Update-size statistics and the fraction of contributions hitting the accepted ceiling say whether some submissions were far larger than typical. Cohort composition says whether the odd round shared a locale, client version or device population — a coherent honest explanation. Per-slice accuracy says where a shift landed, which flat aggregate accuracy hides entirely. And probes against the released global model ask whether a conditional behaviour reproduces, which is the only test you can run repeatedly on an artefact you control. Finally, do not downgrade the finding for reproducing once in five: clients are sampled per round, so an adversary holding few identities appears intermittently by construction.
code
text · 8 linesround clients_sampled mean_update_norm at_norm_ceiling heldout_acc probe_hit_rate
1041 3,200 0.41 1.8% 68.9% 0.3%
1042 3,200 0.44 2.0% 68.8% 0.4%
1043 3,150 0.58 9.4% 68.7% 2.1%
1044 3,200 0.43 1.9% 68.9% 0.5%
1045 3,180 0.42 2.1% 69.0% 0.6%
...
(no per-client examples retained; raw text never left the devices)go deeper
Know that a suspicious round cannot be investigated by looking at the data behind it, because that data was never collected.
Be able to list what the operator does hold — norm statistics, cohort composition, per-slice evaluation, probes on the released model — and what each one can and cannot establish.
Show the judgment: rule the population explanation in first, do the reach arithmetic, reproduce against the model rather than the rounds, and write a disposition whose every claim points the right way.
Own the instrumentation decision this incident should produce, and be willing to say publicly that attribution is not available in this design rather than implying an investigation could have found it.
## The situation One round of a cross-device federation — next-word prediction over millions of handsets — produced a global model that behaved oddly on something you noticed. The next four rounds looked unremarkable. Somebody has to decide whether this was an ordinary property of a shifting client population or an enrolled adversary, and write a disposition. ## First, name what you do not have There is no stored batch behind any update, and there never was: the design's whole point is that the examples stayed on the devices. So the investigative move you would reach for centrally — pull the rows, look at them, re-label a sample — does not exist here. Saying this out loud early is not defeatism; it is what stops the triage from spending a week looking for an artefact that was never created. ## What you do have, and what each thing bounds **Update-size statistics.** The distribution of contribution norms in the round, and the fraction that hit whatever ceiling the server applies. A spike in that fraction says some submissions were far larger than typical. It does *not* say who or why: a client version that changed local step counts produces the same shape. **Cohort composition.** Which populations were sampled — locale, client version, device class, time of day. A shift that tracks a coherent population attribute is the honest explanation, and it is the one worth ruling in first because it is by far the more common cause. **Per-slice evaluation.** Aggregate held-out accuracy is nearly useless here. It stays flat under a targeted change by design, and it moves for a dozen boring reasons. Accuracy broken out by slice — locale, input length, vocabulary region — tells you *where* something landed, which is the difference between a population shift and a change concentrated somewhere nobody's data lives. **Probes against the released global model.** This is the strongest evidence you can generate, because the model is an artefact you control and can query as often as you like. If the odd behaviour is conditional on particular inputs, it reproduces on demand there, independent of rounds and sampling. **The reach arithmetic.** Ask whether the observed shift is even purchasable: given the norm the server accepts and the cohort size, how many identities would it take to produce it? If the answer is implausibly large, that is real evidence toward a population explanation. If it is one or two, the shift tells you nothing about intent. ## The intermittency trap "It happened once and not in the four rounds after" feels like a flake, and here it is not informative at all. Clients are **sampled** each round. An adversary holding a handful of identities appears in a cohort only occasionally, so intermittency is the expected signature of a small attacker and of an unusual honest sub-population alike. Frequency of appearance is not severity, and a finding that reproduces once in five attempts is not thereby a lower-severity finding — it is a finding you have to reproduce against the model rather than against the rounds. ## The disposition you can honestly write Be careful about the direction of every sentence in the ticket: - Four clean rounds bound **the checks you ran** in those rounds. They do not bound what the model contains. - Flat aggregate accuracy shows a broad degradation did not occur. It is equally consistent with an adversary who preserved it. - A probe that did not fire bounds **the probes you wrote** — one behaviour, on the inputs you thought to try. - No identity is ever established from an update alone. At most you can say a contribution was atypical in size. So the honest disposition is usually: population explanation ruled in or not, reach arithmetic stating how much an adversary would have needed, probe results with their coverage stated, and a decision about whether the model ships — not a verdict about who did it. If the evidence cannot separate the two causes, say that; the design does not owe you a separation. ## What changes the next investigation The useful output of a round like this is usually not attribution but instrumentation: retain the norm distribution and ceiling-hit fraction per round, retain cohort composition, evaluate per slice rather than in aggregate, and keep a probe suite that runs against every candidate release. None of those inspect a contribution — nothing can — but together they turn "something looked wrong once" into a claim with a shape.
- Which measurements are worth retaining per round when you can never see client data?The distribution of contribution norms and the fraction hitting the accepted ceiling, cohort composition, per-slice held-out accuracy alongside aggregate, and results from a standing probe suite run against each candidate release. Every one of them bounds what you looked for rather than what a round contained, and that limitation belongs in the ticket.
- The odd behaviour reproduces once in five attempts. Does that lower the severity?No. Under per-round client sampling, intermittency is exactly what a small enrolled adversary looks like, and also what an unusual honest sub-population looks like. Severity comes from what the behaviour does and how cheaply it can be triggered on the released model, not from how often a round happened to show it.
- Can you attribute the round to a specific client?Practically, no. You can say a contribution was atypical in size, and you can say which identities were sampled. You cannot say what an update encoded, and identity in a cross-device federation is cheap by design, so even a confident pointer names an enrolment rather than a person.
saying these in an interview costs you the question
- Proposes inspecting the data behind a suspicious update
- Closes the finding because four later rounds looked normal
- Treats flat aggregate accuracy as an all-clear
- Reads intermittency as evidence of a flake
- Claims attribution to a client from update statistics