skip to content

Tighten the accepted update size or raise the cost of enrolling a federated client — how do you choose?

level: principalimportance: nice to knowfreq 23%

answer

  1. Three factors, three different payers
  2. A ceiling clips the unusual honest client
  3. Identity friction is another team's ledger
  4. The cheapest factor is release cadence
  5. None of them is detection

basics

~20 s

Both price write access rather than inspecting it, and each bills a different group. A tighter ceiling taxes the honest clients with the largest updates; costlier enrolment taxes adoption. Rounds between evaluations is usually the cheapest factor.

solid answer

~50 s

An enrolled adversary's reach is a product of three things: the contribution size the server accepts, the identities they can enrol, and the rounds before the next evaluation gate. You can only move factors, and each has an owner who pays. Lowering the accepted size caps every identity's worth, but it clips hardest on honest clients whose updates are legitimately largest — the rare dialects and atypical usage the federated design existed to reach — so you spend exactly the tail you built this for. Raising enrolment cost multiplies nothing for the attacker but taxes growth and tends to exclude low-end devices and users without strong identity, and it lands on a team that does not report to you. Cutting rounds between real evaluations of the released model usually costs shipping latency and nothing else. Say plainly that none of the three buys detection.

go deeper

for a junior

Know that the available levers change what an attack costs rather than revealing what was submitted, and that each lever has a price somebody pays.

for a middle

Explain why a ceiling on accepted updates clips atypical honest clients hardest, and why identity cost is the multiplier in the reach product.

for a senior

Show that you would attach per-slice evaluation to any tightening, and that you would look at release cadence before negotiating enrolment friction.

for a principal

Own the allocation across all three factors and the organisational cost of each, decide whether this deployment faces an adaptive adversary at all, and refuse an assurance claim the design cannot support.

## Framing the decision honestly The request usually arrives as "harden the federation." There is no hardening available in the sense the asker means, because nothing can inspect a contribution. What you can do is change what write access costs and what one unit of it is worth. The reach of an enrolled adversary is a product of three factors, and a lead's job is to decide which factor to move, knowing who is billed for each. ## Factor one: the accepted contribution size Bounding the norm the server will accept converts an unbounded write into one worth a known, capped amount. That is the single largest structural change available, and the first question to ask of any deployment is whether the bound exists at all. Tightening an existing bound is where the judgment starts. A ceiling is not selective: it clips whatever is largest, and what is legitimately largest is an honest client whose local data looks least like everyone else's — a minority language, an unusual domain vocabulary, a professional usage pattern. Those are precisely the users a cross-device federation was built to learn from, since they are the ones a central corpus would have under-collected. So a tighter ceiling buys reach reduction and pays for it out of the tail, and the bill lands in per-slice accuracy where nobody is watching if the dashboard is aggregate. **If you tighten, you owe a per-slice evaluation, not an aggregate one.** ## Factor two: the price of an identity The attacker's multiplier is identities, and identities in a cross-device setting are usually cheap by design — the product wanted frictionless participation. Raising that price is the most direct hit on the multiplier. But it is not your cost to pay and not your team's ledger. Enrolment friction lands on growth and on exactly the populations who have the weakest identity story: prepaid devices, shared handsets, users in markets where whatever proof you would demand is uncommon. You will be told, correctly, that this reduces the diversity of the training population, which is the same tail the previous paragraph already taxed. And it is a decision that is hard to reverse once shipped, because it changes the product's front door. ## Factor three: rounds before a gate This is the factor teams forget and it is usually the cheapest. Influence accumulates over rounds and decays when it stops being reinforced, so the window that matters is not round cadence but the interval between real evaluations of a candidate release — per-slice accuracy plus a standing probe suite against the model itself. Shortening it costs shipping latency and evaluation budget. It costs no tail accuracy and no adoption. In most deployments this is where the first increment should go, and it is the factor a lead can move without negotiating with another organisation. ## The claim you must not let anyone make All three price write access. **None of them provides visibility into what was written.** After the tightest bound and the most expensive enrolment, a contribution is still a vector with nothing behind it. If somebody wants to write "the federated pipeline is protected against poisoning" in a review document, the honest replacement is: *an accepted contribution is worth at most this much; an identity costs this much to obtain; a release is evaluated this way, searching for these behaviours, at this cadence.* That is a statement about cost and coverage, and it is the only kind this design supports. ## The prior question Before spending on any factor, ask whether this deployment faces an adaptive adversary at all. A model whose behaviour nobody profits from bending, in a population nobody is registering identities into, may be better served by spending the same budget on per-slice evaluation and on the boring population-shift problems that will actually degrade it — a client version regression, a locale sampled badly, a distribution that moved. Spending the tail's accuracy on an adversary who is not there is a real cost incurred for a hypothetical one avoided, and a lead is the person expected to say so out loud. ## What a strong answer sounds like Name the three factors. Assign each one an owner who pays and a metric where the bill shows up. Take rounds-to-gate first because it is cheapest, insist on per-slice evaluation before touching the ceiling, treat enrolment cost as a cross-organisational negotiation rather than a config change, and refuse the sentence that any of this makes contributions auditable.

  • How do you answer a reviewer who asks for proof the deployed federated model is clean?
    You cannot provide it, and saying so is the answer. What you can state is the reach an accepted contribution has, what an identity costs to obtain, the evaluations run before release and what behaviours they searched for. A claim about what the model contains is not available in this design; a claim about what it costs to write into it is.
  • The product team refuses any enrolment friction. What do you do with that constraint?
    Accept it and re-spend on the other two factors: make sure an accepted contribution has a bound at all, and shorten the interval between real per-slice evaluations of a candidate release. Then record in the risk register that the identity multiplier is unbounded by product decision, so nobody later reads the remaining controls as covering it.
  • Why is per-slice evaluation the condition you attach to tightening the accepted size?
    Because a ceiling clips the largest honest updates first, and those come from the least typical users. Aggregate accuracy will barely move while a minority slice degrades, so an aggregate dashboard reports the change as free. Per-slice numbers are what make the bill visible to the people deciding to pay it.

saying these in an interview costs you the question

  • Claims a norm ceiling makes contributions auditable
  • Tightens the bound and reports only aggregate accuracy
  • Ignores identity cost as the attacker's real multiplier
  • Never considers whether an adaptive adversary exists here
  • Promises a clean-model guarantee the design cannot support

context