skip to content

Why does solving for a network's weights need many digits of each returned score?

level: middleimportance: should knowfreq 38%

answer

  1. the attacker is not reading predictions
  2. a gentle join between two flat pieces
  3. leading digits agree, low ones differ
  4. coarse steps switch several units at once
  5. deeper units arrive fainter at the output

basics

~10 s

The evidence a solve reads is a faint change of slope where one unit switches on. It lives in the low-order digits of a returned number, so a coarsely rounded reply erases it.

solid answer

~50 s

A piecewise-linear network computes a function made of flat pieces joined at surfaces where individual units switch on or off. Locating one of those joins precisely is what turns behaviour into an equation about a unit's weights and bias. But near a join the departure from a single straight piece is tiny, so the attacker is comparing returned numbers that agree in their leading digits and differ far down. If the endpoint reports a score to two decimals, those differences are quantised away and the join cannot be pinned down; the attacker is pushed to coarse steps where several units have already switched and no single unit is isolated. Precision also sets the reachable depth: a deeper unit leaves a fainter imprint on the final output, so more significant digits are needed per layer. Digits, not query volume, are the binding limit here.

go deeper

for a junior

Know that this attack reads the small differences between returned numbers, not the predictions themselves, and that a heavily rounded score does not carry them.

for a middle

Be ready to explain why locating one bend is an equation about one unit, and why the departure from a straight piece near that bend is tiny enough to live in the low-order digits.

for a senior

Show you can look at an endpoint and judge exposure from the reply's numeric width together with the model's size and depth, and describe the effect of narrower replies as a cost increase rather than a fix.

for a principal

Be able to argue where the line between affordable and unaffordable sits for your own models, and to say what would move it, without promising that a numeric contract is a security boundary.

## The signal being read is small on purpose Think about the target from this topic: a small equipment-failure-risk regressor over turbine and pump telemetry, returning one real number per call. An adversary trying to solve for its parameters is not trying to learn what the model predicts. They are trying to learn *where the computed function bends*. A network built from piecewise-linear units is linear on each of many regions of its input space. The boundaries between regions are the surfaces where some unit crosses from one of its linear pieces to the other. Each such surface belongs to exactly one unit, and its position and orientation are determined by that unit's weights and bias. Pin down the surface and you have written down an equation about those parameters. That is the whole conversion from behaviour into algebra. ## Why that conversion is precision-bound The problem is that the bend is gentle in the returned value. On one side of the boundary the reply moves along one straight relationship with the input; on the other side it moves along a slightly different one. Immediately at the boundary, the difference between those two behaviours is minute. To establish that a bend is *there*, and not somewhere nearby, the adversary must be able to distinguish returned numbers that agree in their leading digits. That gives precision a very direct role: | What the reply carries | What the adversary can establish | | --- | --- | | Many significant digits of a real value | Where a single unit's boundary sits, tightly enough to write an equation | | A few decimals | That the output changed at all, over steps coarse enough that several units have already switched | | A discrete label | Only which side of a decision the input fell on | At low precision the attacker is forced into large steps, and over a large step multiple units switch. The clean attribution of one bend to one unit is gone, and with it the equation. ## Depth compounds it Precision is not a fixed requirement across the model. A unit in the first layer acts on the input directly, so its switch shows up in the output relatively strongly. A unit several layers in has its effect passed through everything after it, and typically arrives at the output attenuated. So each additional layer the adversary works inward demands more significant digits to see the same kind of event, on top of demanding more queries. This is the second half of why the technique is a small-model result: it is not only that a wide model has more parameters to pin down, it is that a deep one needs arithmetic the reply may simply not contain. ## The shape of the limit, stated honestly It is worth being careful about direction here, because it is easy to state this backwards: - Fewer digits in the reply **raise the attacker's cost and can put the algebraic solve out of reach**. That is a real effect and it is the limit this attack works under. - Fewer digits **do not** stop an attacker from fitting a behavioural stand-in on the same replies. Copying tolerates coarse answers; it is the exact solve that does not. - 'The score is a float, so we are exposed' is equally wrong as a blanket claim. The exposure depends jointly on how many digits actually leave the service, how big and how deep the model is, and whether its units are piecewise-linear. So the precision of the returned number is the single most informative field to look at when asking whether this class of attack is even on the table for a given endpoint — and it is a property of the *number*, not of how many calls the attacker makes. This is the unusual case in extraction where volume is not the story. A solve against a small model is a modest number of calls, all of them individually unremarkable; what makes them valuable to the adversary is how much of each answer they get to see. ## What to say in an interview Name the mechanism (a bend belongs to one unit and locating it is an equation), name why it is faint (the two linear pieces barely differ near the join), name what that implies about the reply format (low-order digits are the material), and name the compounding effect of depth. Then state the limit as a cost, not as a guarantee: reduced precision moves this attack from feasible to unaffordable for a given model size, and where exactly that line falls depends on the model.

  • Does reporting fewer digits also blunt an attacker who is fitting a behavioural copy?
    Barely. A fitted stand-in is trained on the replies as supervision, and supervision survives rounding — a slightly noisier target costs the attacker a little accuracy and nothing structural. Precision is the binding limit for the algebraic solve specifically, because that route depends on distinguishing numbers that differ far below the leading digits. Treating a narrower numeric reply as a general anti-extraction measure overstates it considerably.
  • Why does the required precision grow as the attacker works into deeper layers?
    A first-layer unit switching leaves a comparatively strong mark on the output because its effect passes through fewer transformations. A unit several layers in has its contribution scaled and mixed by everything downstream, so the same switching event shows up as a much smaller change in the returned number. The adversary needs correspondingly more significant digits to see it, which is why depth, not just parameter count, bounds the technique.

saying these in an interview costs you the question

  • Says the attack is limited by query volume, not precision
  • Claims rounding the score stops all extraction attacks
  • Thinks the attacker is reading predictions rather than bends
  • Ignores that depth raises the precision requirement
  • Assumes any float reply makes any model recoverable

context