In an adversarial ML evaluation, what does it mean to grant an attacker white-box access?
answer
- An assumption, not an incident
- Who is handed what, on purpose
- So the number is not about secrecy
- The most capable adversary you can define
- Weaker adversaries all sit underneath
basics
~20 sWhite-box access is an assumption that the attacker holds the model's weights, architecture and gradients. You grant it on purpose during evaluation so the result measures the model's robustness rather than how well you kept the weights secret.
solid answer
~50 sWhite-box is an **access assumption**, not an incident. It says: assume the adversary has the trained parameters, the architecture and the ability to differentiate the model with respect to its input. Evaluators are handed exactly that, deliberately and under contract, because an evaluation run without it silently measures your obscurity instead of your model — if the finding is 'they could not do it', you cannot tell whether the model held or whether the attacker simply lacked the weights. Granting the access removes that confound. It also defines the most capable adversary you can write down for a given deployment, which is why the number produced under it is a bound rather than an anecdote. Two other meanings of the word are in circulation and neither is this one: white-box in ordinary software testing (test design informed by source), and 'black-box model' meaning an uninterpretable one.
go deeper
Be ready to say in one sentence that white-box and black-box describe what the attacker is assumed to have, not something that happened. Name the three things white-box grants: weights, architecture, gradients.
An interviewer expects you to explain why an evaluator is deliberately handed the weights: without the grant, a negative result cannot be told apart from the tester simply lacking information. Distinguish this sense from the QA and interpretability senses of the same words.
Show that you attach the access assumption to every robustness figure you report or read, and that you keep weights private as a cost control while still evaluating as though they were public.
Own the argument to a stakeholder who calls the granted-access evaluation unrealistic: secrecy is not a safety property, and one granted-access result constrains a whole family of cheaper adversaries you will never enumerate.
## The word names an assumption, not an event In adversarial machine learning, *white-box* and *black-box* describe **what the adversary is assumed to have**, not something that happened to you. A white-box adversary is assumed to hold: - the **trained weights** of the model; - its **architecture** — layer structure, input preprocessing, output head; - the ability to compute **gradients of a loss with respect to the input**, because holding weights and architecture means you can differentiate the function end to end. A black-box adversary holds none of that and interacts only through predictions. In between, an adversary might know the architecture and the rough data distribution but not the parameters. That intermediate case is a different assumption with different consequences and is not what white-box means. ## Why a lab is *handed* the weights Consider a regulated setting: a chest-radiograph triage classifier that flags studies for urgent read, submitted by a device vendor to an accredited evaluation lab as a condition of clearance. The submission package includes the weights, the architecture, the training recipe and a differentiable copy of the model, transferred under NDA, and the contract explicitly forbids the lab from pretending it is an outside tester. That looks backwards until you ask what the alternative measures. If the lab is only allowed to send inputs and read outputs, and it reports that it could not make the model misread a study, the report has two possible causes and no way to separate them: 1. the model genuinely resists perturbed inputs; or 2. the lab could not find the perturbation because it did not have the weights. Cause (2) is a statement about the lab's information, not about the device. Clearance cannot rest on it, because the vendor's secrecy is not a safety property — weights leak, models get shipped to devices, and staff move. Granting access deletes cause (2) from the report. What comes back is then attributable to the model. ## The result is a bound on everybody else Because the white-box adversary is the most capable one you can define for that model — anything a weaker adversary learns, this one already had — the attack success measured under it is an **upper bound on what any weaker adversary can achieve**, and the surviving accuracy is correspondingly a **lower bound** on how the model fares against them. That is the whole reason the stance is worth its cost: one expensive evaluation constrains a whole family of cheaper adversaries you will never enumerate. The inverse is also true and is the practical reason results are unreadable without the assumption stated. A reported attack success rate with no access assumption attached could have come from an adversary holding gradients or from one holding a single returned label, and those two numbers differ by an order of magnitude on the same model. This is why every credible robustness claim states its access class alongside its perturbation family and radius. ## What granting access is *not* - **It is not a claim that the weights were stolen.** No breach is being asserted or predicted. The grant is a modelling choice made before any testing begins. - **It is not an assumption about the training data.** White-box concerns the trained artefact. An adversary who additionally holds the training set is a further, separate assumption, and privacy attacks turn on that distinction. - **It is not the software-testing sense.** In QA, white-box testing means test cases informed by internal source and control flow. The words collide; the objects do not. - **It is not 'black-box model' in the interpretability sense.** There, black-box means a model whose reasoning is opaque to a human. That is a property of the model. Here, black-box is a property of the attacker's access. Because three fields use these two words for three different things, a careful engineer says *access* out loud: 'white-box access — weights and gradients assumed available'. ## What withholding weights is actually worth Keeping parameters private is a real and worthwhile cost control: it makes an adversary pay in queries, in money, or in the effort of building a stand-in that agrees with your model where it matters. It is not a boundary, though, and it does not change what the model does when someone eventually holds it. So the operating posture is to keep weights private *and* to evaluate as though they were not, and to report both numbers with their assumptions attached rather than one number with none.
- Does granting white-box access mean you expect an attacker to obtain the weights?No. It is a modelling choice made before testing, not a prediction. You grant it so the finding is attributable to the model rather than to your secrecy. Whether weights are likely to leak is a separate, real question that changes how urgently you act on the finding — it does not change whether the evaluation should have been run that way.
- How is this different from white-box testing in ordinary software QA?QA white-box means test cases designed with knowledge of internal source and control flow, to reach branches a functional test would miss. Adversarial-ML white-box is an assumption about what the adversary holds — parameters and gradients of a trained model — used to bound an attacker, not to improve coverage. Same phrase, unrelated objects; say 'access' to disambiguate.
- Does white-box access also imply access to the training data?No, and conflating them is a common slip. White-box concerns the trained artefact: weights, architecture, gradients. An adversary who also holds training rows, or write access to them, is a strictly different and stronger assumption, and it is the one that matters for poisoning and for several privacy attacks. State them separately.
It is like testing a lock with the blueprints and a key blank on the bench. You are not claiming the burglar has them; you are making sure the report says something about the lock rather than about how well you hid the drawings.
saying these in an interview costs you the question
- Says white-box means the attacker already breached the system
- Dismisses the evaluation because the weights are kept private
- Confuses it with an uninterpretable 'black-box model'
- Assumes white-box also grants the training data
- Quotes an attack result without naming the access assumed
- Treats keeping weights secret as a robustness property