Why does a curious aggregation server's reconstruction of a client's data degrade as that client's local batch grows?
answer
- one observation, many unknowns
- the update is a sum over examples
- contributions cancel as the batch grows
- what comes back is a blend, not a row
- smooth degradation, never a threshold
basics
~20 sOne uploaded update summarises the whole batch, so per-example signals sum together and the server must recover many unknown examples from one fixed-size observation. The problem becomes underdetermined, and quality falls from recognisable examples to vague structure.
solid answer
~50 sAn update uploaded after training on one example is very nearly a description of that example: the server holds the model, so it can search for an input that reproduces the update, and the search is tightly constrained. Train on a batch of sixty-four and the same fixed-size update now summarises sixty-four examples whose individual contributions have been summed and largely cancelled — the number of unknowns grows while the observation does not, so the reconstruction is underdetermined and returns a blurry average of the batch rather than any row. Local steps compound this: if the client runs several steps before uploading, the server sees only the endpoint of a trajectory through weights it never observed. The same effect is why cohort aggregation helps. Note the shape of the result — quality degrades smoothly, so this is a cost curve for the adversary, not a threshold with a guarantee on the far side.
go deeper
Know that an update covering many examples reveals less about any one of them than an update covering a single example, and that batch size is therefore a privacy-relevant setting, not just a throughput knob.
Explain the counting argument: the observation stays one parameter-shaped vector while the unknowns scale with the batch, contributions cancel, and the reconstruction converges on a blend. Add that local steps compound the difficulty by hiding the trajectory.
Demonstrate that you would report yield as a curve over batch size, local steps and cohort size rather than a pass or fail, and that you distinguish the degradation you measured from a bound you cannot offer.
Own the framing when configuration numbers become the privacy story: decide whether the organisation is prepared to defend an empirical curve in front of a regulator, or whether the deployment needs a stated bound and the accuracy cost that comes with it.
## The setting A federation trains a shared model across clients — say a load-disaggregation model across household smart meters. Each round, a client receives the global weights, computes an update from its own consumption windows, and uploads it. The adversary is the aggregation server: honest-but-curious, holding the global model, reading updates as they arrive, wanting the training examples behind one of them. Its limit is that it sees the update and nothing else. The question is how much that limit costs it, and the answer is set almost entirely by three numbers the client's configuration chooses. ## Why batch size is the dominant one Start with the extreme. If the client's update came from a single example, the server has an observation of fixed, large dimension that was produced by one unknown input passed through a model it fully knows. It can propose a candidate input, compute what update that candidate would produce, and compare. The comparison is informative in every coordinate, so the search is heavily over-constrained and converges on something visually or numerically close to the real example. This is the case people demonstrate, and it is genuinely alarming. Now raise the batch. The uploaded update does not get bigger — it is still one parameter-shaped vector — but it is now a **sum over examples**, so: - **The unknowns multiply while the observation does not.** Recovering sixty-four windows from one vector is a far more underdetermined problem than recovering one, and many different batches produce nearly the same update. - **Contributions cancel.** Examples that push the model in opposing directions partly annihilate each other in the sum, so distinguishing features of individual rows are the first thing lost. - **What survives is the average.** Reconstructions from large batches tend to converge on a blend — plausible-looking but corresponding to no actual row, which is a much weaker disclosure. Empirically the curve is smooth: exact-looking recovery at a batch of one, degraded but semantically meaningful recovery for a handful, and mostly unusable output at the batch sizes ordinary training uses. Where exactly it breaks depends on the model, the data and how much prior knowledge the adversary can bring, which is why a red-team result must always report the batch size it was obtained at. ## Local steps: the second lever, often forgotten Most real federations do not upload a single gradient. The client runs several local steps — sometimes a full pass over its data — and uploads the difference between its final and initial weights. That changes the adversary's problem qualitatively, not just quantitatively. The observation is now the endpoint of a path through intermediate weight states the server never saw, each step taken on a different mini-batch in an order the server does not know. To invert it, the adversary must jointly guess the data and the entire trajectory. More local steps therefore reduce yield sharply, which is one reason a design that uploads after every single step is the worst case. ## Cohort aggregation: the same effect one level up Combining many clients' updates before anyone reads them makes the effective batch the union of all their examples. The mechanism is identical — more examples summed into one observation — which is why batch size, local steps and cohort size are best understood as one lever pulled at three places. ## What this does and does not establish The critical framing point: this is **degradation, not a guarantee**. A large batch makes reconstruction expensive and low-quality against the attack you ran. It does not bound what a better attack, or an adversary with a strong prior over household energy patterns, can recover — a strong prior partly substitutes for the missing constraints. Nor does it stop cheaper inferences: the structure of the final layer's update often reveals which target values or classes were present in the batch regardless of size, and that alone can be sensitive. If you need a statement that holds against attacks nobody has run yet, batch size cannot give it to you. Only bounding each example's contribution and adding noise calibrated to that bound produces a stated privacy claim, and it charges accuracy for it — mostly on the clients whose behaviour is unusual, which in an energy federation means exactly the households that differ from the population. ## How to talk about it in an interview Say what the observation is, say what the unknowns are, and note that the ratio between them is the whole story. Then add the caveat that separates a middle answer from a good one: the curve is smooth, adversary-dependent, and produces a cost, not a boundary.
- Does running more local steps before uploading help or hurt the client's privacy?It helps, and for a different reason than batching. With several local steps the upload is a weight delta accumulated across intermediate states the server never observed, on mini-batches in an unknown order. The adversary must guess the data and the trajectory jointly, which is far harder than inverting one gradient at known weights.
- If reconstruction fails at batch 64, can you tell the business that batch 64 is safe?No. You can say your attack stopped producing usable output at that setting. Yield depends on the model, the data and the adversary's prior over the domain, and a stronger prior substitutes for the constraints the larger batch removed. Report it as the point where your attack stopped, and be explicit that it bounds your attack, not the channel.
- Is anything about the batch still recoverable even when the examples are not?Yes. The structure of the last layer's update commonly reveals which target values or classes appeared in the batch, and that survives batch sizes at which input reconstruction has long failed. In some domains that alone is the sensitive fact — knowing an appliance category or a diagnosis label was present can matter more than the raw window.
saying these in an interview costs you the question
- Treats a large batch as a guarantee rather than a cost
- Assumes reconstruction fails abruptly at some batch size
- Ignores local steps and models the upload as one gradient
- Thinks a blurry reconstruction means nothing leaked
- Believes cohort aggregation works by a different mechanism than batching