skip to content

A design doc says the federation's aggregation rule tolerates 20% malicious clients — what do you ask before relying on it?

level: seniorimportance: should knowfreq 35%

answer

  1. a fraction of which population
  2. which client partition was it measured on
  3. displacement bounded, not failure prevented
  4. was the rule assumed known to the adversary
  5. nobody reported the honest spread

basics

~20 s

Ask what the 20% is a fraction of, what client partition it was measured on, and how widely honest updates already spread in this federation. A tolerance derived on statistically identical clients says little about silos with genuinely different books.

solid answer

~40 s

Four questions, in order. First: a fraction of what — the enrolled population, or the clients sampled in a given round? Per-round cohorts fluctuate, so a modest global share exceeds the threshold in some rounds by chance. Second: measured on which participants? Tolerances are usually established by partitioning one dataset uniformly, which makes clients alike and honest updates tight — the premise the rule needs and a real consortium does not satisfy. Third: what quantity is bounded? These results bound how far the aggregate can be pulled, not "the attack fails", and a targeted behaviour needs very little pull. Fourth: was the adversary assumed to know the rule? If not, the number measures obscurity. Then I would ask for the one thing usually missing: the measured spread of honest updates in our own federation.

code

text · 10 lines
text
robust-aggregation evaluation (as written in the design doc)
  rule ................... coordinate-wise median
  clients ................ 100
  client partition ....... uniform random shards of one dataset
  malicious clients ...... 20 / 100  -> "tolerates 20%"
  accuracy under attack .. 91.4%   (no-attack baseline 92.1%)
  ...
  honest update spread ... not reported
  fraction is of ......... not stated (enrolment or per-round cohort?)
  adversary knows rule ... not stated

go deeper

for a junior

Know that a tolerance percentage is a conclusion from assumptions, not a property of the code, and that the first thing to ask is what population it was measured on.

for a middle

Be able to name the premise — honest updates cluster — and say why a uniformly partitioned evaluation guarantees that premise holds in the paper and not in a consortium.

for a senior

Show a review order: fraction of what, measured on whom, bounding which quantity, against an adversary who knows what — then ask for the honest spread in your own federation.

for a principal

Separate the two acts of trust: relying on the rule's qualitative purchase versus relying on a specific percentage, and be explicit about which one the organisation is signing up for.

## Why a tolerance figure needs interrogating "Tolerates 20% malicious clients" reads like a property of the aggregation rule. It is a conclusion drawn from a set of premises, and every premise is a place the number can fail to transfer to your deployment. Reviewing such a claim is mostly a matter of asking which premises were stated and which were silently borrowed. ## 1. Twenty per cent of what population Many federations do not aggregate every enrolled participant every round; they sample a cohort. The share that matters for the aggregation rule is the malicious share **among the clients selected for that round**, not the share across the enrolment. Small cohorts fluctuate, so an adversary holding a modest global share will be over the tolerated fraction in some rounds purely by sampling variation — and a round is all a persistent adversary needs, repeatedly. A guarantee stated over the enrolment is being applied to a sample it was never about. ## 2. Measured on which participants This is the question that usually settles it. Tolerance results are typically established with clients built by partitioning one dataset uniformly at random, which makes the participants statistically identical. Identical participants produce tightly clustered honest updates, and tight clustering is exactly the premise the rule's bound rests on: the adversary must either sit inside a pinhole (and be harmless) or step outside it (and be trimmed). A cross-organisation consortium is the opposite case by construction — different customer mixes, product lines and geographies mean honest updates disagree widely, the surviving region is large, and an update chosen from inside it is unremarkable in position while still being consequential. Ask what the client partition was. If it was uniform, the number describes a federation you do not have. ## 3. What quantity is actually bounded Robustness results of this kind bound how far the aggregate can be displaced relative to what the honest clients would have produced. They do not say "the attack fails". A degradation attack that wants the model measurably worse does need substantial displacement, so the bound bites. A conditional behaviour keyed to inputs the adversary controls needs very little displacement and is designed to leave the headline metric alone. So a bounded displacement is not an assurance of nothing having happened; it is a bound on one axis. Relatedly: if the report shows main-task accuracy under attack essentially unchanged, that establishes that the attack that was run did not move the average metric. It is equally consistent with nothing having happened and with something conditional having been installed that the average metric cannot see. Aggregate accuracy holding flat is evidence about the metric, not about the model. ## 4. What the adversary was assumed to know If the evaluation's adversary did not know which aggregation rule was running, the result is partly measuring obscurity. The rule is a design decision that the participants can generally infer; assume it is known. An adversary who knows the rule aims at the region it will accept, and that is the case the number needs to cover. ## 5. The measurement that is usually missing Everything above converges on one quantity nobody reports: **how widely do our own honest clients' updates already spread, compared with the size of update that would actually change the model's behaviour?** That ratio is what the protection is worth. It is measurable from ordinary training telemetry, it needs no adversary to obtain, and it turns the conversation from a citation into an observation about this federation. Ask for it before onboarding, and again after the population changes. ## How to read a report block Given a summary that lists rule, client count, malicious count and accuracy under attack, the missing columns are the finding: the client partition, the honest spread, whether the adversary knew the rule, and whether the fraction refers to the enrolment or the round cohort. A robustness summary without those is not comparable to any other robustness summary, and it is certainly not a statement about your consortium. ## The judgment to state out loud You can rely on the rule for what it genuinely gives — a minority can no longer *choose* the model — while declining to rely on the specific percentage until it has been re-established on a population that resembles yours. Those are two different acts of trust, and conflating them is what the design doc has done.

  • The 20% turns out to be a share of the clients sampled each round. Why does that matter?
    Because the share that binds is the malicious fraction among the clients selected for a given round, and small cohorts fluctuate. An adversary holding a modest share of the enrolment will exceed the tolerated fraction in some rounds by sampling variation alone, and a persistent adversary only needs the rounds where it does. A bound stated over the enrolment is being applied to a sample it was never about.
  • The report shows accuracy under attack essentially unchanged. What does that establish?
    That the attack that was run did not move the average metric — nothing more. A conditional behaviour keyed to inputs the adversary chooses is meant to leave headline accuracy alone, so a flat number is equally consistent with 'nothing happened' and with 'something was installed that this metric cannot see'. Ask what was measured besides the average.
  • What single measurement would you ask the team to produce?
    The spread of honest client updates in our own federation, set against the size of update that would meaningfully change the model's behaviour. That ratio is what the rule is actually worth here. It comes out of ordinary training telemetry, requires no adversary, and replaces a borrowed citation with an observation about this population.

saying these in an interview costs you the question

  • Accepts the tolerated fraction without asking about the client partition
  • Reads flat accuracy under attack as proof that nothing was installed
  • Assumes the fraction is over the enrolment when cohorts are sampled per round
  • Never asks whether the adversary was assumed to know the rule
  • Treats an empirical tolerance number as a guarantee about this deployment

context