What does recovering a model's weights 'only up to symmetry' leave an attacker holding?
answer
- the units have no intrinsic order
- two rearrangements the output cannot see
- scale in, unscale out, nothing changes
- an equivalence class, not one matrix
- matters for comparing weights, not for capability
basics
~20 sA set of parameters that computes exactly the same function as yours, but need not match yours entry by entry: hidden units can be permuted and a unit's incoming and outgoing weights rescaled against each other without changing any reply.
solid answer
~50 sNeural networks with piecewise-linear units have built-in redundancies. Reorder the units within a hidden layer, permuting the corresponding rows and columns, and the network computes the identical function. Multiply one unit's incoming weights and bias by a positive constant and divide its outgoing weight by the same constant, and again nothing observable changes. So an attack that works purely from input-output behaviour can only ever identify the equivalence class, never one canonical weight matrix — that is what 'up to permutation and scaling symmetry' means. The practical reading matters in two directions. Functionally it is no consolation at all: every member of the class returns the same numbers, so the attacker has a genuinely equivalent model, not an approximation. Forensically it does matter: you cannot compare weight matrices entry by entry and conclude nothing was taken, because a recovered copy is not expected to line up with yours.
go deeper
Remember the two invariances by name: reordering hidden units, and scaling a unit's inputs up while scaling its output down. Both leave every reply unchanged.
Be ready to explain why behaviour cannot distinguish members of the equivalence class, and to say plainly that this costs the attacker nothing functionally.
Show you would not accept an entry-wise weight comparison as evidence about whether a model was taken, and that you know what a valid comparison would have to account for.
Own the framing when someone reports 'the weights do not match' as reassurance: state what that test can and cannot establish before anyone builds a decision on it.
## Where the symmetry comes from Two distinct rearrangements leave a piecewise-linear network's output completely unchanged. **Permutation.** The hidden units within a layer have no intrinsic order. Swap two of them — swapping the corresponding rows of the incoming weight matrix and the corresponding columns of the outgoing one — and every input produces exactly the same output as before. For a layer of n units there are n factorial such rearrangements, all indistinguishable from outside. **Positive rescaling.** A rectified piecewise-linear unit is positively homogeneous: scaling its pre-activation by a positive constant scales its output by the same constant. So multiply one unit's incoming weights and its bias by some positive c, divide its outgoing weight by c, and the two changes cancel exactly. Again nothing observable moves. Because both operations are invisible in the function, no attack that works from input-output behaviour alone can distinguish between them. An algebraic solve for the parameters therefore recovers an **equivalence class**, not the one weight matrix sitting on your servers. ## Reading the limit in the right direction This is a place where candidates commonly draw the wrong conclusion, in both directions. **The wrong optimistic reading:** 'they only got the weights up to symmetry, so they did not really get the weights.' Functionally this is empty. Every member of the class computes the identical function on every input — not approximately, identically. The attacker can serve it, fine-tune it, inspect it, and study it offline for further attacks. Symmetry costs them nothing they wanted. **The wrong pessimistic reading:** treating symmetry as a defect of the attack that makes the recovered model unreliable. It is not a source of error. The recovered parameters are exact within the precision achieved; symmetry is about which representative of an exactly-equivalent family you land on, not about how close you got. ## Where it genuinely changes what you can say The place symmetry actually bites is on the defender's side, in questions about provenance and evidence. If you are asked whether a model somebody else is serving is yours, comparing weight matrices entry by entry is not a valid test. A parameter-recovered model is *expected* to differ entry by entry — the units may be in a different order and each may be scaled differently — while computing precisely the same function. So an entry-wise mismatch is not evidence of independence, and you would need either a canonicalisation that quotients out the known symmetries or a behavioural comparison to say anything. The same caution applies to the reverse claim. Two models that agree closely in behaviour are not thereby proven to be copies either; agreement can come from similar tasks and overlapping data. Symmetry is a reason an entry-wise test is invalid, not a licence to substitute a behavioural test and treat it as proof. ## What an interviewer is listening for Three things, in this order. First, the mechanism — that permutation and positive rescaling are exact invariances of the computed function, so behaviour cannot see them. Second, the honest impact — functionally zero for the attacker. Third, the one place it does matter — that you cannot compare weights entry by entry to decide whether something was taken. A candidate who reports the symmetry as though it were a partial protection has understood the phrase but not the consequence.
- Does the symmetry make the recovered model less accurate than the original?No. Symmetry is not error. Permuting units and rescaling a unit's weights against its outgoing weight leave the computed function exactly unchanged, so a representative drawn from the equivalence class returns identical values. Whatever inaccuracy the attacker has comes from limited numeric precision in what they could read, not from which member of the class they landed on.
- If a recovered copy will not match your weights entry by entry, what can you compare instead?Either canonicalise both models to quotient out the known invariances before comparing, or compare behaviour on inputs you choose. Both are weaker than they look: canonicalisation must account for every symmetry the architecture admits, and behavioural agreement can arise from similar tasks and overlapping training data. Neither yields a bare entry-wise match, so a report claiming theft on a weight comparison alone is not saying what it appears to.
Two identical wiring looms can have their wires bundled in a different order and cut to different insulation colours; the circuit they realise is the same one, and no measurement at the terminals distinguishes them.
saying these in an interview costs you the question
- Treats recovery up to symmetry as a partial protection
- Says the symmetry makes the copy approximate rather than exact
- Claims entry-wise weight comparison proves a model was not copied
- Confuses the symmetry with numeric imprecision in the solve
- Thinks reordering units changes any returned value