In a test federation you reconstructed a client's training window from one uploaded update — what does that establish for production?
answer
- the finding is the configuration
- batch, steps, visibility, cohort size
- report where it stopped, and why that is weak
- your bound is on your attack
- an empirical curve is not a guarantee
basics
~20 sIt establishes that the update channel leaks under the exact configuration you tested, and nothing about any other one. The finding is only meaningful with its numbers attached: local batch, local steps, whether uploads were individual or cohort sums, and cohort size.
solid answer
~60 sReport the configuration, not the screenshot. A reconstruction obtained from a per-client update over a batch of one with a single local step demonstrates that the channel carries recoverable content at that setting — which is the worst case and rarely the production setting. Before it means anything for rollout you must state the local batch size, the number of local steps per round, whether the server ever sees an individual update or only a cohort sum, the cohort size, and whether any per-example bound and calibrated noise are in place. Then run the ladder and report the setting at which your reconstruction stopped producing usable output — and say explicitly that this bounds your attack, not the channel, because an adversary with a stronger prior over the domain gets more from the same update. Finish with the design ask: minimum batch and local steps, no per-client update ever readable, a minimum cohort size with membership the server does not choose, and a decision on whether a stated bound is required.
code
text · 8 linesconfig local_batch local_steps server_sees cohort reconstruction
----------------------------------------------------------------------------------
pilot-A 1 1 per-client - window recovered, appliance events legible
pilot-B 8 1 per-client - partial: occupancy on/off pattern only
pilot-C 64 5 per-client - no usable output in 20 attempts
production (planned) 64 5 cohort sum 500 NOT TESTED
...
per-example bound / noise: none configured in any rowgo deeper
Know that a reconstruction result is meaningless without the settings it was obtained under, and that batch size and how many examples were mixed into the update are the first numbers to write down.
Be ready to lay out the ladder — batch, local steps, individual versus cohort-sum uploads, cohort size — and explain why each rung changes the adversary's problem rather than just the difficulty.
Demonstrate the discipline that separates the two claims: where your attack stopped, versus what the channel guarantees. Turn the finding into enforceable client-side configuration and a named decision about whether a stated bound is required.
Own the decision the ladder forces: whether the organisation ships on an empirical curve or funds a stated bound and the accuracy it costs, and who is accountable for that wording if someone later asks what was recoverable.
## The chair you are sitting in You are red-teaming a federation before rollout — a load-disaggregation model across household smart meters, where a recovered training example is a specific home's consumption window and therefore says whether anyone was in the house. You played the aggregation server, took one client's uploaded update, and rebuilt the window behind it. The finding is real. The question is what it licenses you to say. ## What the result actually establishes It establishes one conditional: **under the configuration you ran, an adversary with your vantage and your effort recovered training content from the update channel.** Everything that matters is in the conditions, and a finding written without them is not actionable — it is a demonstration. The conditions to state, every time: - **Local batch size.** Reconstruction from a single example is the strongest case and degrades steeply as examples are mixed. - **Local steps per round.** One step means you inverted a gradient at weights you knew. Several means the upload was a delta across a trajectory you did not observe, which is a different and much harder problem. - **Update visibility.** Did you read an individual client's update, or a cohort sum? A result from an individual update says nothing about a deployment where individuals are never visible. - **Cohort size**, if sums are what the server sees, and **who selects membership**. - **Any per-example contribution bound and calibrated noise**, and their settings. - **Rounds and effort** — how many attempts, how long, and how reliably it reproduced. ## The report that is worth something The useful artefact is a ladder: the same attack against increasing batch, increasing local steps, and then against cohort sums, with the quality of what came back at each rung. That turns a demonstration into a curve, and it lets whoever owns the rollout read off where their planned configuration sits. And here is the sentence that separates a senior finding from a junior one: **the rung where your attack stopped working is a bound on your attack, not on the channel.** A different adversary with a stronger prior over household energy patterns — knowing typical appliance signatures, typical daily rhythms — replaces some of the constraints your larger batch removed, and gets further at the same setting. Nothing you measured forbids that. If the deployment needs a claim that survives attacks nobody has run, batch size and aggregation cannot supply it; only bounding each example's contribution during training and adding noise calibrated to that bound produces a stated guarantee, and it costs accuracy — landing hardest on the households whose usage patterns are unusual, who are often exactly the ones the product cares about. ## Reproducibility and triage If the reconstruction worked once in five attempts, say so in that form. Attacks in this family are optimisation searches with initialisation-dependent outcomes; unreliability is a property of the search, not evidence that the leak is marginal. A one-in-five recovery of a household's occupancy pattern is a finding, because the adversary chooses how many attempts to make and which client to target. ## What to ask for Turn the finding into configuration requirements rather than a warning: 1. A minimum local batch size and a minimum number of local steps before any upload, enforced in the client, not documented in a wiki. 2. No individual update ever readable by the server — cohort sums only. 3. A minimum cohort size, with membership the server does not unilaterally choose and a rule preventing the same participant from being isolated across consecutive rounds. 4. An explicit decision, made by a named owner, on whether the deployment must offer a stated bound rather than an empirical curve — and if so, budget for the accuracy it costs and check where that cost lands across households. ## The failure modes to avoid in the write-up - Reporting the reconstruction without the configuration, which invites either panic or dismissal, both unearned. - Reporting the rung where it stopped as a safe setting. - Claiming the finished model is safe because the channel was fixed — attacks on the trained artefact are a separate family and untouched by any of this. - Framing it as a protocol flaw. Nothing malfunctioned: the update is a function of the data, and the deployment simply chose settings that let that function be inverted.
- The reconstruction only worked in one attempt out of five. Does that lower the severity?Not much. These attacks are optimisation searches whose outcome depends on initialisation, so unreliability is a property of the search rather than evidence that the leak is small. The adversary picks the target and the number of attempts. Report the success rate honestly and rate the finding on what a successful run discloses.
- Production plans a batch of 64 with cohort sums. Can you sign that off from these results?No — that row was never tested. You can say your attack failed against per-client updates at batch 64 with five local steps, which suggests the planned setting is materially harder. Signing off requires running the planned configuration, and even then the result is an empirical curve, not a bound anyone can defend against a future attack.
- What would make you insist on per-example bounds and calibrated noise rather than configuration alone?When the deployment has to make a claim to somebody who can compel an answer, or when a single recovered example is severe enough that an empirical curve is not an acceptable basis. Then budget for the accuracy cost and check where it lands — it falls hardest on clients whose data is unusual, who may be the ones the product exists to serve.
saying these in an interview costs you the question
- Reports the reconstruction without the configuration it came from
- Calls the setting where the attack failed a safe setting
- Generalises a per-client result to a cohort-sum deployment
- Dismisses a finding because it reproduces intermittently
- Claims the trained model is safe once the channel is fixed
- Presents an empirical curve as a privacy guarantee