In federated training, what does secure aggregation remove from a curious server's view of client updates, and what remains?
answer
- the server sees a total, not a sender
- no individual update means no targeted inversion
- who decides who is in the cohort
- a sum is still computed from the data
- an access control, not a bound
basics
~20 sIt removes the individual update: the server can read only the cohort's combined update, so it cannot pick a victim and invert their upload. What remains is that the sum is still a function of the data, the server usually picks who is in each cohort, and the trained model still leaks.
solid answer
~50 sSecure aggregation means individual uploads are combined so that the server obtains only the cohort's sum and never any one client's update. That removes the highest-yield attack in this family — choose a household, invert the update it just sent — and it raises the effective batch to the union of the cohort's examples, which is exactly the mixing effect that degrades reconstruction. What it does not remove: the sum is still computed from client data, so it is still a channel; the server typically controls cohort membership and round scheduling, so it can shrink how much honest data a given sum actually mixes, or compare sums across rounds whose membership differed; and the finished global model still supports membership and memorisation attacks that have nothing to do with the upload channel. It is an access control on what the server may read, not a bound on what can be inferred.
go deeper
Know the one-line effect: the server sees only a combined update for the whole cohort, never one client's, so it cannot single out a participant's upload and work backwards from it.
Explain both halves — the targeted inversion disappears and the effective batch becomes the cohort — and name the residuals: the sum is still data-dependent, membership is usually chosen by the server, and the finished model leaks independently.
Show you would review it as configuration: minimum cohort size, who selects membership, whether a participant can be isolated across rounds, and whether any per-example bound exists on top. Say plainly that aggregation alone yields no bound.
Own what the organisation is allowed to claim from it. Deciding whether a protocol property may be published as a privacy guarantee, and what evidence must sit behind that wording, is a call somebody has to make and defend.
## What the protocol class does In a plain federation the server receives each client's update, adds them up, and applies the result. Secure aggregation changes only who can see what: the clients combine their updates in a way that lets the server learn the sum over a cohort while learning nothing about any individual contribution. From the server's chair, the round's input goes from a list of named updates to a single anonymous total. That is a real and significant change to the threat model, and it is worth being precise about both halves of it. ## What it genuinely removes **The targeted attack.** The strongest version of update inversion picks a specific client — a specific household in an energy federation — takes the update that client uploaded this round, and searches for the inputs that would reproduce it. This requires the individual update. Under cohort aggregation the server never holds one, so the attack has no input. **Per-client attribution generally.** Anything the server would infer by comparing one client's update against another's, or tracking one client's updates across rounds, needs a per-client view it no longer has. **Effective batch size.** The sum mixes every example every cohort member trained on. If a cohort of five hundred clients each trained on a few dozen windows, the sum summarises tens of thousands of examples, and reconstruction from that is far past the point where quality collapses. ## What remains **The sum is still a function of the data.** Aggregation changes the granularity of the leak, not its existence. There is no theorem here saying nothing is recoverable from a sum — only the empirical observation that yield falls steeply with the number of examples mixed in. **The server usually decides who is in the cohort and when.** This is the sharpest residual risk and the one candidates most often miss. If the aggregation server also selects cohort membership, it can arrange rounds so that a victim's contribution dominates a sum — by pairing them with participants whose contributions it already knows, or by comparing two sums whose membership differed by exactly one client. The protocol hides individual updates from an observer, but it does not by itself constrain a server that shapes what goes into each sum. Making membership unpredictable, requiring minimum cohort sizes, and preventing the same participant from being isolated across rounds are the design answers. **The final model still leaks.** Every privacy attack that operates on a trained model — inferring whether a record was in the training set, extracting memorised content, probing what the model believes a class looks like — is untouched by how the model was assembled. The cohort sum protects the channel, not the artefact it produces. **There is still no bound.** This is the framing that matters for a review. Secure aggregation is an access control: it changes what the server is permitted to observe. A stated privacy guarantee is a different kind of object, obtained by bounding each example's contribution and adding noise calibrated to that bound, and it can be composed with aggregation but is not implied by it. If someone reports 'we use secure aggregation' as the answer to 'what is your privacy guarantee', the honest correction is that they have described a protocol property, not a bound, and the question of what an adversary can infer from the sums they legitimately see is still open. ## How the pieces stack A useful ordering for a design review, weakest to strongest: - per-client, per-step updates in the clear — the worst case, near-verbatim recovery of small batches; - per-client updates with large local batches and several local steps — degraded yield, still targeted; - cohort sums only, with server-controlled membership — no targeted inversion, residual risk from cohort shaping; - cohort sums with unpredictable membership and a minimum cohort size — the practical engineering answer; - the above plus per-example contribution bounds and calibrated noise — the only rung that yields a stated claim, and the one that charges accuracy, disproportionately on clients whose data is unusual. ## The interview answer Name both halves. Secure aggregation deletes the per-client view and therefore the targeted inversion; it leaves the sum as a data-dependent object, leaves cohort composition in the hands of the party you are defending against, and leaves every attack that operates on the finished model. Anyone who describes it as making federated training private has stopped at the first half.
- If the server chooses cohort membership, why is that a residual privacy risk?Because it decides how much honest data each sum actually mixes. A server that can place a chosen participant alongside contributors whose updates it already knows, or compare two rounds whose cohorts differed by one participant, recovers something close to an individual view despite the protocol. Unpredictable membership and a minimum cohort size are the controls.
- A team reports secure aggregation as their privacy guarantee. What do you ask next?Ask what it bounds. It is an access control describing what the server may read, and it comes with no statement about what is inferable from the sums it does read. Ask for the minimum cohort size, who controls membership, whether the same participant can be isolated across rounds, and whether any per-example bound and calibrated noise exist on top.
- Does cohort aggregation protect the finished global model?No. Attacks that operate on the trained model — deciding whether a record was in its training set, eliciting memorised content, probing what it believes a category looks like — depend on the model, not on how updates reached the server. Aggregation is a control on the assembly channel and leaves the artefact's own leakage exactly where it was.
saying these in an interview costs you the question
- Calls secure aggregation a privacy guarantee rather than an access control
- Ignores that the server often selects cohort membership
- Assumes a sum cannot be a channel at all
- Thinks it also protects the finished global model
- Cannot state a minimum cohort size or who chooses it